{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Titanic - Machine Learning From Disaster","metadata":{}},{"cell_type":"markdown","source":"### 1. Importing the libraries","metadata":{}},{"cell_type":"markdown","source":"A Python library is a collection of related modules. It contains bundles of code that can be used repeatedly in different programs. It makes Python Programming simpler and convenient for the programmer. As we don't need to write the same code again and again for different programs.\n\nIn this notebook, we will be using the following libraries.","metadata":{}},{"cell_type":"code","source":"### Data Wrangling \n\nimport numpy as np\nimport pandas as pd\nimport missingno\nfrom collections import Counter\nfrom collections import OrderedDict\n\n### Data Visualization\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n### Modelling \n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import confusion_matrix, accuracy_score\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.naive_bayes import GaussianNB\nfrom sklearn.svm import SVC\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.ensemble import RandomForestClassifier\n\n### Tabulating the results\n\nfrom tabulate import tabulate\n\n### Model Validation\n\nfrom sklearn.model_selection import cross_val_score\n\n### Remove unnecessary warnings\n\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:33:59.487088Z","iopub.execute_input":"2022-07-11T01:33:59.488476Z","iopub.status.idle":"2022-07-11T01:34:00.126709Z","shell.execute_reply.started":"2022-07-11T01:33:59.488383Z","shell.execute_reply":"2022-07-11T01:34:00.125421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2. Importing the data","metadata":{}},{"cell_type":"markdown","source":"In this section, I will fetch the training and test datasets that are available in the Kaggle's project description in the Data section.\n\nThe training set should be used to build your machine learning models. For the training set, we provide the outcome (also known as the “ground truth”) for each passenger. Your model will be based on “features” like passengers’ gender and class. You can also use feature engineering to create new features.\n\nThe test set should be used to see how well your model performs on unseen data. For the test set, we do not provide the ground truth for each passenger. It is your job to predict these outcomes. For each passenger in the test set, use the model you trained to predict whether or not they survived the sinking of the Titanic.","metadata":{}},{"cell_type":"code","source":"### Fetching the train and test datasets\n\ntrain_dataset = pd.read_csv(\"/kaggle/input/titanic/train.csv\")\ntest_dataset = pd.read_csv(\"/kaggle/input/titanic/test.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:00.128204Z","iopub.execute_input":"2022-07-11T01:34:00.128559Z","iopub.status.idle":"2022-07-11T01:34:00.1464Z","shell.execute_reply.started":"2022-07-11T01:34:00.128528Z","shell.execute_reply":"2022-07-11T01:34:00.145539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Looking at the sample data in the training dataset\n\ntrain_dataset.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:00.147685Z","iopub.execute_input":"2022-07-11T01:34:00.148434Z","iopub.status.idle":"2022-07-11T01:34:00.171035Z","shell.execute_reply.started":"2022-07-11T01:34:00.148401Z","shell.execute_reply":"2022-07-11T01:34:00.17004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Shape of the training set\n\ntrain_dataset.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:00.173538Z","iopub.execute_input":"2022-07-11T01:34:00.174102Z","iopub.status.idle":"2022-07-11T01:34:00.180866Z","shell.execute_reply.started":"2022-07-11T01:34:00.174069Z","shell.execute_reply":"2022-07-11T01:34:00.179926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The training dataset consists of 12 columns and 891 rows.","metadata":{}},{"cell_type":"code","source":"### Looking at the sample data in the test set\n\ntest_dataset.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:00.182594Z","iopub.execute_input":"2022-07-11T01:34:00.183284Z","iopub.status.idle":"2022-07-11T01:34:00.202961Z","shell.execute_reply.started":"2022-07-11T01:34:00.183248Z","shell.execute_reply":"2022-07-11T01:34:00.201204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Shape of the test dataset\n\ntest_dataset.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:00.204411Z","iopub.execute_input":"2022-07-11T01:34:00.206012Z","iopub.status.idle":"2022-07-11T01:34:00.215902Z","shell.execute_reply.started":"2022-07-11T01:34:00.205968Z","shell.execute_reply":"2022-07-11T01:34:00.214827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The test dataset consists of 11 columns and 418 rows.","metadata":{}},{"cell_type":"markdown","source":"### 3. Exploratory Data Analysis (EDA)","metadata":{}},{"cell_type":"markdown","source":"Exploratory Data Analysis refers to the critical process of performing initial investigations on data so as to discover patterns,to spot anomalies,to test hypothesis and to check assumptions with the help of summary statistics and graphical representations.\n\nHere, we will perform EDA on the categorical columns of the dataset - Sex, Pclass, Embarked and the numerical columns of the dataset - Age, Fare, Sibsp, Parch.","metadata":{}},{"cell_type":"markdown","source":"#### 3.1 Datatypes, Missing Data, and Summary Statistics","metadata":{}},{"cell_type":"code","source":"### Looking at the datatypes of the training and test data\n\ntrain_dataset.info()\nprint('-' * 50)\ntest_dataset.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:00.217779Z","iopub.execute_input":"2022-07-11T01:34:00.218167Z","iopub.status.idle":"2022-07-11T01:34:00.242588Z","shell.execute_reply.started":"2022-07-11T01:34:00.218134Z","shell.execute_reply":"2022-07-11T01:34:00.240992Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here, the columns - Pclass, Name, Ticket, Cabin, and Embarked are categorical. Hence, we modify the datatype of these columns to category.","metadata":{}},{"cell_type":"code","source":"### Changing the datatype of the columns - Pclass, Name, Ticket, Cabin, Embarked to category in the training data\n\ntrain_dataset.Pclass = train_dataset.Pclass.astype('category')\ntrain_dataset.Name = train_dataset.Name.astype('category')\ntrain_dataset.Sex = train_dataset.Sex.astype('category')\ntrain_dataset.Ticket = train_dataset.Ticket.astype('category')\ntrain_dataset.Cabin = train_dataset.Cabin.astype('category')\ntrain_dataset.Embarked = train_dataset.Embarked.astype('category')","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:00.244308Z","iopub.execute_input":"2022-07-11T01:34:00.24469Z","iopub.status.idle":"2022-07-11T01:34:00.262339Z","shell.execute_reply.started":"2022-07-11T01:34:00.244655Z","shell.execute_reply":"2022-07-11T01:34:00.261176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Changing the datatype of the columns - Pclass, Name, Sex, Ticket, Cabin, Embarked to category in the test data\n\ntest_dataset.Pclass = test_dataset.Pclass.astype('category')\ntest_dataset.Name = test_dataset.Name.astype('category')\ntest_dataset.Sex = test_dataset.Sex.astype('category')\ntest_dataset.Ticket = test_dataset.Ticket.astype('category')\ntest_dataset.Cabin = test_dataset.Cabin.astype('category')\ntest_dataset.Embarked = test_dataset.Embarked.astype('category')","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:00.263942Z","iopub.execute_input":"2022-07-11T01:34:00.264392Z","iopub.status.idle":"2022-07-11T01:34:00.278926Z","shell.execute_reply.started":"2022-07-11T01:34:00.26436Z","shell.execute_reply":"2022-07-11T01:34:00.277629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looking at the modified datatypes of the columns in both the training and test datasets.","metadata":{}},{"cell_type":"code","source":"### Looking at the modified datatypes of the training and test data\n\ntrain_dataset.info()\nprint('-' * 50)\ntest_dataset.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:00.280243Z","iopub.execute_input":"2022-07-11T01:34:00.281099Z","iopub.status.idle":"2022-07-11T01:34:00.310577Z","shell.execute_reply.started":"2022-07-11T01:34:00.281067Z","shell.execute_reply":"2022-07-11T01:34:00.309383Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above data, it is evident that there are missing values in the dataset.","metadata":{}},{"cell_type":"code","source":"### Missing data by columns in the training set\n\ntrain_dataset.isnull().sum().sort_values(ascending = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:00.31197Z","iopub.execute_input":"2022-07-11T01:34:00.312851Z","iopub.status.idle":"2022-07-11T01:34:00.323464Z","shell.execute_reply.started":"2022-07-11T01:34:00.312803Z","shell.execute_reply":"2022-07-11T01:34:00.322214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here, we can see that there are missing values in the columns - Cabin, Age, Embarked in the training dataset.","metadata":{}},{"cell_type":"code","source":"### Missing data by columns in the test set\n\ntest_dataset.isnull().sum().sort_values(ascending = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:00.324928Z","iopub.execute_input":"2022-07-11T01:34:00.325796Z","iopub.status.idle":"2022-07-11T01:34:00.336245Z","shell.execute_reply.started":"2022-07-11T01:34:00.325761Z","shell.execute_reply":"2022-07-11T01:34:00.335344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here, we can see that there are missing values in the columns - Cabin, Age, Fare in the test dataset.","metadata":{}},{"cell_type":"code","source":"### Visual representation of the missing data in the training set\n\nmissingno.matrix(train_dataset)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:00.341865Z","iopub.execute_input":"2022-07-11T01:34:00.34298Z","iopub.status.idle":"2022-07-11T01:34:00.843255Z","shell.execute_reply.started":"2022-07-11T01:34:00.342944Z","shell.execute_reply":"2022-07-11T01:34:00.842405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Visual representation of the missing data in the test set\n\nmissingno.matrix(test_dataset)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:00.844335Z","iopub.execute_input":"2022-07-11T01:34:00.844859Z","iopub.status.idle":"2022-07-11T01:34:01.311915Z","shell.execute_reply.started":"2022-07-11T01:34:00.844822Z","shell.execute_reply":"2022-07-11T01:34:01.31096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Summary statistics of the numerical columns in the training dataset\n\ntrain_dataset.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:01.312987Z","iopub.execute_input":"2022-07-11T01:34:01.313656Z","iopub.status.idle":"2022-07-11T01:34:01.346204Z","shell.execute_reply.started":"2022-07-11T01:34:01.313623Z","shell.execute_reply":"2022-07-11T01:34:01.344986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Summary statistics of the numerical columns in the test dataset\n\ntest_dataset.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:01.347906Z","iopub.execute_input":"2022-07-11T01:34:01.348225Z","iopub.status.idle":"2022-07-11T01:34:01.375691Z","shell.execute_reply.started":"2022-07-11T01:34:01.348197Z","shell.execute_reply":"2022-07-11T01:34:01.374536Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 3.2 Feature Analysis","metadata":{}},{"cell_type":"markdown","source":"##### 3.2.1 Categorical variable - Survived","metadata":{}},{"cell_type":"code","source":"### Value counts of the column - Survived\n\nsurvived_count = train_dataset['Survived'].value_counts(dropna = False)\nsurvived_count","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:01.37721Z","iopub.execute_input":"2022-07-11T01:34:01.377559Z","iopub.status.idle":"2022-07-11T01:34:01.386005Z","shell.execute_reply.started":"2022-07-11T01:34:01.377512Z","shell.execute_reply":"2022-07-11T01:34:01.384797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Bar graph showing the value counts of the column - Survived\n\nsns.barplot(survived_count.index, survived_count.values, alpha = 0.8)\nplt.title('Bar graph showing the value counts of the column - Survived')\nplt.ylabel('Number of Occurrences', fontsize = 12)\nplt.xlabel('Survived', fontsize = 12)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:01.387922Z","iopub.execute_input":"2022-07-11T01:34:01.38917Z","iopub.status.idle":"2022-07-11T01:34:01.558437Z","shell.execute_reply.started":"2022-07-11T01:34:01.389128Z","shell.execute_reply":"2022-07-11T01:34:01.557177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above graph, we can see that most of the people in the Titanic died during the infamous incident.","metadata":{}},{"cell_type":"markdown","source":"##### 3.2.2 Categorical variable - Sex","metadata":{}},{"cell_type":"code","source":"### Value counts of the column - Sex\n\nsex_count = train_dataset['Sex'].value_counts(dropna = False)\nsex_count","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:01.559943Z","iopub.execute_input":"2022-07-11T01:34:01.5603Z","iopub.status.idle":"2022-07-11T01:34:01.571022Z","shell.execute_reply.started":"2022-07-11T01:34:01.560267Z","shell.execute_reply":"2022-07-11T01:34:01.569645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Bar graph showing the value counts of the column - Sex\n\nsns.barplot(sex_count.index, sex_count.values, alpha = 0.8)\nplt.title('Bar graph showing the value counts of the column - Sex')\nplt.ylabel('Number of Occurrences', fontsize = 12)\nplt.xlabel('Sex', fontsize = 12)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:01.573208Z","iopub.execute_input":"2022-07-11T01:34:01.573659Z","iopub.status.idle":"2022-07-11T01:34:01.740952Z","shell.execute_reply.started":"2022-07-11T01:34:01.573625Z","shell.execute_reply":"2022-07-11T01:34:01.740156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above garph, we can see that most of the passengers in the Titanic are Male.","metadata":{}},{"cell_type":"code","source":"### Mean of survival by Sex\n\nsex_survived = train_dataset[['Sex', 'Survived']].groupby('Sex', as_index = False).mean()\nsex_survived","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:01.742335Z","iopub.execute_input":"2022-07-11T01:34:01.742979Z","iopub.status.idle":"2022-07-11T01:34:01.759352Z","shell.execute_reply.started":"2022-07-11T01:34:01.742944Z","shell.execute_reply":"2022-07-11T01:34:01.758175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Survival Probability by Gender\n\nsns.barplot(sex_survived['Sex'], sex_survived['Survived'], alpha = 0.8)\nplt.title('Survival Probability by Gender')\nplt.ylabel('Probability', fontsize = 12)\nplt.xlabel('Sex', fontsize = 12)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:01.762838Z","iopub.execute_input":"2022-07-11T01:34:01.763274Z","iopub.status.idle":"2022-07-11T01:34:01.934639Z","shell.execute_reply.started":"2022-07-11T01:34:01.763232Z","shell.execute_reply":"2022-07-11T01:34:01.933539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above graph, we can see that the probability of survival is higher in females than males.","metadata":{}},{"cell_type":"markdown","source":"##### 3.2.3 Categorical variable - Pclass","metadata":{}},{"cell_type":"code","source":"### Value counts of the column - Pclass\n\npclass_count = train_dataset['Pclass'].value_counts(dropna = False)\npclass_count","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:01.936068Z","iopub.execute_input":"2022-07-11T01:34:01.936409Z","iopub.status.idle":"2022-07-11T01:34:01.946632Z","shell.execute_reply.started":"2022-07-11T01:34:01.936378Z","shell.execute_reply":"2022-07-11T01:34:01.945449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Bar graph showing the value counts of the column - Pclass\n\nsns.barplot(pclass_count.index, pclass_count.values, alpha = 0.8)\nplt.title('Bar graph showing the value counts of the column - Pclass')\nplt.ylabel('Number of Occurrences', fontsize = 12)\nplt.xlabel('Passenger class', fontsize = 12)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:01.948077Z","iopub.execute_input":"2022-07-11T01:34:01.948421Z","iopub.status.idle":"2022-07-11T01:34:02.124466Z","shell.execute_reply.started":"2022-07-11T01:34:01.948377Z","shell.execute_reply":"2022-07-11T01:34:02.123331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above graph, we can see that most of the passengers belong to Pclass 3.","metadata":{}},{"cell_type":"code","source":"### Mean of survival by Pclass\n\npclass_survived = train_dataset[['Pclass', 'Survived']].groupby('Pclass', as_index = False).mean()\npclass_survived","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:02.126384Z","iopub.execute_input":"2022-07-11T01:34:02.126856Z","iopub.status.idle":"2022-07-11T01:34:02.141093Z","shell.execute_reply.started":"2022-07-11T01:34:02.126802Z","shell.execute_reply":"2022-07-11T01:34:02.14025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Survival Probability by Pclass\n\nsns.barplot(pclass_survived['Pclass'], pclass_survived['Survived'], alpha = 0.8)\nplt.title('Survival Probability by Passenger Class')\nplt.ylabel('Probability', fontsize = 12)\nplt.xlabel('Passenger class', fontsize = 12)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:02.142158Z","iopub.execute_input":"2022-07-11T01:34:02.142934Z","iopub.status.idle":"2022-07-11T01:34:02.320946Z","shell.execute_reply.started":"2022-07-11T01:34:02.142899Z","shell.execute_reply":"2022-07-11T01:34:02.319811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above graph, we can see that the survival rate is higher in the passengers who belong to Pclass 1 than compared to the passengers in Pclass 2 and 3. One reason might be that most of the passengers in Pclass 1 might be female and that's the reason they have a high probability of survival as females have a high probability of survival than compared to male.\n\nNow, let's this hypothesis.","metadata":{}},{"cell_type":"code","source":"### Survival Probability by Sex and Passenger Class\n\nsns.factorplot(x = 'Pclass', y = 'Survived', hue = 'Sex', data = train_dataset, kind = 'bar')\nplt.ylabel('Survival Probability')\nplt.title('Survival Probability by Sex and Passenger Class')","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:02.322567Z","iopub.execute_input":"2022-07-11T01:34:02.322899Z","iopub.status.idle":"2022-07-11T01:34:02.858123Z","shell.execute_reply.started":"2022-07-11T01:34:02.322862Z","shell.execute_reply":"2022-07-11T01:34:02.856986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above graph, it is clearly evident that most of the passengers travelling in Pclass 1 are female. Hence, the Pclass 1 passengers have a higher probability of survival than the passengers in Pclass 2 and 3. Hence our hypothesis is verified correctly.","metadata":{}},{"cell_type":"markdown","source":"##### 3.2.4 Categorical variable - Embarked","metadata":{}},{"cell_type":"code","source":"### Value counts of the column - Embarked\n\nembarked_count = train_dataset['Embarked'].value_counts(dropna = False)\nembarked_count","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:02.859972Z","iopub.execute_input":"2022-07-11T01:34:02.860315Z","iopub.status.idle":"2022-07-11T01:34:02.870751Z","shell.execute_reply.started":"2022-07-11T01:34:02.860284Z","shell.execute_reply":"2022-07-11T01:34:02.869552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Bar graph showing the value counts of the column - Embarked\n\nsns.barplot(embarked_count.index, embarked_count.values, alpha = 0.8)\nplt.title('Bar graph showing the value counts of the column - Embarked')\nplt.ylabel('Number of Occurrences', fontsize = 12)\nplt.xlabel('Port of Embarkation', fontsize = 12)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:02.87215Z","iopub.execute_input":"2022-07-11T01:34:02.872544Z","iopub.status.idle":"2022-07-11T01:34:03.004163Z","shell.execute_reply.started":"2022-07-11T01:34:02.872484Z","shell.execute_reply":"2022-07-11T01:34:03.003296Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above graph, we can see that most of the passengers who travelled in Titanic had their Port of Embarkation - S (Southampton).","metadata":{}},{"cell_type":"code","source":"### Mean of survival by Embarked\n\nembarked_survived = train_dataset[['Embarked', 'Survived']].groupby('Embarked', as_index = False).mean()\nembarked_survived","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:03.00548Z","iopub.execute_input":"2022-07-11T01:34:03.006643Z","iopub.status.idle":"2022-07-11T01:34:03.024451Z","shell.execute_reply.started":"2022-07-11T01:34:03.006596Z","shell.execute_reply":"2022-07-11T01:34:03.023193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Survival Probability by Embarked\n\nsns.barplot(embarked_survived['Embarked'], embarked_survived['Survived'], alpha = 0.8)\nplt.title('Survival Probability by Port of Embarkation')\nplt.ylabel('Probability', fontsize = 12)\nplt.xlabel('Port of Embarkation', fontsize = 12)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:03.026156Z","iopub.execute_input":"2022-07-11T01:34:03.027466Z","iopub.status.idle":"2022-07-11T01:34:03.151967Z","shell.execute_reply.started":"2022-07-11T01:34:03.027429Z","shell.execute_reply":"2022-07-11T01:34:03.150841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above graph, we can see that passengers embarked at Cherbourg have a higher probability of survival than the passengers who embarked at Queenstown and Southampton. This could be because most of the passengers who embarked at Cherbourg might be first class passengers, hence it has a higher probability of survival.\n\nNow let's this hypothesis.","metadata":{}},{"cell_type":"code","source":"### Distribution of Pclass for each Port of Embarkation\n\nsns.factorplot('Pclass', col = 'Embarked', data = train_dataset, kind = 'count')","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:03.153727Z","iopub.execute_input":"2022-07-11T01:34:03.154435Z","iopub.status.idle":"2022-07-11T01:34:03.70412Z","shell.execute_reply.started":"2022-07-11T01:34:03.154392Z","shell.execute_reply":"2022-07-11T01:34:03.702941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above graph, we can see that most of the passengers who embarked at Cherbourg belong to Pclass 1, hence they have higher probability of survival. Similarly, for the passengers who embarked at Queenstown and Southampton most of them belong to Pclass 3, hence the lower probability of survival. Hence, our hypothesis is true.","metadata":{}},{"cell_type":"markdown","source":"##### 3.2.5 Numerical variable - Age","metadata":{}},{"cell_type":"code","source":"### Understanding the distribution of the column - Age\n\nsns.distplot(train_dataset['Age'], label = 'Skewness: %.2f'%(train_dataset['Age'].skew()))\nplt.legend(loc = 'best')\nplt.title('Passenger Age Distribution')","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:03.705895Z","iopub.execute_input":"2022-07-11T01:34:03.706629Z","iopub.status.idle":"2022-07-11T01:34:03.919718Z","shell.execute_reply.started":"2022-07-11T01:34:03.706585Z","shell.execute_reply":"2022-07-11T01:34:03.918574Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above graph, we can see that the distribution is close to normal distribution with a slight skewness.","metadata":{}},{"cell_type":"code","source":"### Age distribution by survival\n\ngrid = sns.FacetGrid(train_dataset, col = 'Survived')\ngrid.map(sns.distplot, 'Age')","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:03.921454Z","iopub.execute_input":"2022-07-11T01:34:03.921938Z","iopub.status.idle":"2022-07-11T01:34:04.330755Z","shell.execute_reply.started":"2022-07-11T01:34:03.921893Z","shell.execute_reply":"2022-07-11T01:34:04.329611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Merging both the above graphs into one\n\nsns.kdeplot(train_dataset['Age'][train_dataset['Survived'] == 0], label = 'Did not survive')\nsns.kdeplot(train_dataset['Age'][train_dataset['Survived'] == 1], label = 'Survived')\nplt.xlabel('Age')\nplt.legend()\nplt.title('Passenger Age Distribution by Survival')","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:04.332531Z","iopub.execute_input":"2022-07-11T01:34:04.332907Z","iopub.status.idle":"2022-07-11T01:34:04.571351Z","shell.execute_reply.started":"2022-07-11T01:34:04.332875Z","shell.execute_reply":"2022-07-11T01:34:04.568681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 3.2.6 Numerical variable - Fare","metadata":{}},{"cell_type":"code","source":"### Understanding the distribution of the column - Fare\n\nsns.distplot(train_dataset['Fare'], label = 'Skewness: %.2f'%(train_dataset['Fare'].skew()))\nplt.legend(loc = 'best')\nplt.title('Passenger Fare Distribution')","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:04.573229Z","iopub.execute_input":"2022-07-11T01:34:04.573636Z","iopub.status.idle":"2022-07-11T01:34:04.915144Z","shell.execute_reply.started":"2022-07-11T01:34:04.573604Z","shell.execute_reply":"2022-07-11T01:34:04.913683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above graph, we can see that there is a high degree of skewness in the above distribution.","metadata":{}},{"cell_type":"code","source":"### Fare distribution by Passenger class\n\ngrid = sns.FacetGrid(train_dataset, col = 'Pclass')\ngrid.map(sns.distplot, 'Fare')","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:04.916718Z","iopub.execute_input":"2022-07-11T01:34:04.917046Z","iopub.status.idle":"2022-07-11T01:34:05.601352Z","shell.execute_reply.started":"2022-07-11T01:34:04.917017Z","shell.execute_reply":"2022-07-11T01:34:05.599683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above graphs, we can see that the distribution of Fare for the passengers in Pclass 1 is wider than compared to the passengers in Pclass 2 and 3. That means, the passengers in Pclass 1 paid a higher fare than compared to the passengers in Pclass 2 and 3.","metadata":{}},{"cell_type":"markdown","source":"### 4. Data preprocessing","metadata":{}},{"cell_type":"markdown","source":"Data preprocessing is the process of getting our dataset ready for model training. In this section, we will perform the following preprocessing steps:\n\n1. Detect and remove outliers in numerical variables\n2. Drop and fill missing values\n3. Feature engineering\n4. Data trasformation\n5. Feature encoding","metadata":{}},{"cell_type":"markdown","source":"#### 4.1 Detect and remove outliers in numerical variables","metadata":{}},{"cell_type":"markdown","source":"Outliers are data points that have extreme values and they do not conform with the majority of the data. It is important to address this because outliers tend to skew our data towards extremes and can cause inaccurate model predictions. I will use the Tukey method to remove these outliers.\n\nHere, we will write a function that will loop through a list of features and detect outliers in each one of those features. In each loop, a data point is deemed an outlier if it is less than the first quartile minus the outlier step or exceeds third quartile plus the outlier step. The outlier step is defined as 1.5 times the interquartile range. Once the outliers have been determined for one feature, their indices will be stored in a list before proceeding to the next feature and the process repeats until the very last feature is completed. Finally, using the list with outlier indices, we will count the frequencies of the index numbers and return them if their frequency exceeds n times.","metadata":{}},{"cell_type":"code","source":"def detect_outliers(df, n, features_list):\n    outlier_indices = [] \n    for feature in features_list: \n        Q1 = np.percentile(df[feature], 25)\n        Q3 = np.percentile(df[feature], 75)\n        IQR = Q3 - Q1\n        outlier_step = 1.5 * IQR \n        outlier_list_col = df[(df[feature] < Q1 - outlier_step) | (df[feature] > Q3 + outlier_step)].index\n        outlier_indices.extend(outlier_list_col) \n    outlier_indices = Counter(outlier_indices)\n    multiple_outliers = list(key for key, value in outlier_indices.items() if value > n) \n    return multiple_outliers\n\noutliers_to_drop = detect_outliers(train_dataset, 2, ['Age', 'SibSp', 'Parch', 'Fare'])\nprint(\"We will drop these {} indices: \".format(len(outliers_to_drop)), outliers_to_drop)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.602583Z","iopub.execute_input":"2022-07-11T01:34:05.602879Z","iopub.status.idle":"2022-07-11T01:34:05.621527Z","shell.execute_reply.started":"2022-07-11T01:34:05.602853Z","shell.execute_reply":"2022-07-11T01:34:05.620445Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's look at the data present in the rows.","metadata":{}},{"cell_type":"code","source":"train_dataset.iloc[outliers_to_drop, :]","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.623415Z","iopub.execute_input":"2022-07-11T01:34:05.624442Z","iopub.status.idle":"2022-07-11T01:34:05.646977Z","shell.execute_reply.started":"2022-07-11T01:34:05.624394Z","shell.execute_reply":"2022-07-11T01:34:05.645711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will drop these rows from the training dataset.","metadata":{}},{"cell_type":"code","source":"### Drop outliers and reset index\n\nprint(\"Before: {} rows\".format(len(train_dataset)))\ntrain_dataset = train_dataset.drop(outliers_to_drop, axis = 0).reset_index(drop = True)\nprint(\"After: {} rows\".format(len(train_dataset)))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.660826Z","iopub.execute_input":"2022-07-11T01:34:05.66195Z","iopub.status.idle":"2022-07-11T01:34:05.669412Z","shell.execute_reply.started":"2022-07-11T01:34:05.661909Z","shell.execute_reply":"2022-07-11T01:34:05.668522Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Lets look at the new training dataset\n\ntrain_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.670699Z","iopub.execute_input":"2022-07-11T01:34:05.671438Z","iopub.status.idle":"2022-07-11T01:34:05.700161Z","shell.execute_reply.started":"2022-07-11T01:34:05.671404Z","shell.execute_reply":"2022-07-11T01:34:05.698965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 4.2 Drop and fill missing values","metadata":{}},{"cell_type":"markdown","source":"Here, we will drop the column - PassengerId from both the training and test sets.","metadata":{}},{"cell_type":"code","source":"### Dropping the columns - PassengerId from the training dataset\n\ntrain_dataset.drop(['PassengerId'], axis = 1, inplace = True)\ntrain_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.701468Z","iopub.execute_input":"2022-07-11T01:34:05.701883Z","iopub.status.idle":"2022-07-11T01:34:05.729515Z","shell.execute_reply.started":"2022-07-11T01:34:05.701855Z","shell.execute_reply":"2022-07-11T01:34:05.728539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Dropping the columns - PassengerId from the test dataset\n\ntest_dataset.drop(['PassengerId'], axis = 1, inplace = True)\ntest_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.730713Z","iopub.execute_input":"2022-07-11T01:34:05.731041Z","iopub.status.idle":"2022-07-11T01:34:05.75928Z","shell.execute_reply.started":"2022-07-11T01:34:05.731015Z","shell.execute_reply":"2022-07-11T01:34:05.758165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Looking at the missing values in the training set\n\ntrain_dataset.isnull().sum().sort_values(ascending = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.760796Z","iopub.execute_input":"2022-07-11T01:34:05.761115Z","iopub.status.idle":"2022-07-11T01:34:05.772826Z","shell.execute_reply.started":"2022-07-11T01:34:05.761085Z","shell.execute_reply":"2022-07-11T01:34:05.771657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the training dataset, we can see that there are missing values in the columns Cabin, Age, and Embarked. Here, we will focus on replacing the missing values in the columns Age and Embarked. We will take care of the column - Cabin during Feature Engineering.","metadata":{}},{"cell_type":"code","source":"### Looking at the missing values in the test set\n\ntest_dataset.isnull().sum().sort_values(ascending = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.775059Z","iopub.execute_input":"2022-07-11T01:34:05.775949Z","iopub.status.idle":"2022-07-11T01:34:05.786418Z","shell.execute_reply.started":"2022-07-11T01:34:05.775902Z","shell.execute_reply":"2022-07-11T01:34:05.785532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the test dataset, we can see that the missing values are in the columns - Cabin, Age and Fare. Here, we will focus on replacing the missing values in the columns Age and Fare. We will take care of the column - Cabin during Feature Engineering.","metadata":{}},{"cell_type":"markdown","source":"##### 4.2.1 Handling missing values - Age (training and test sets)","metadata":{}},{"cell_type":"markdown","source":"From the distribution of the column Age, we can see that the data is almost normally distributed. Hence, we will use the median value to replace the missing values of the column.","metadata":{}},{"cell_type":"code","source":"### Finding the median value of the column - Age in the training set\n\nage_index = list(~train_dataset['Age'].isnull())\nmedian_age = np.median(train_dataset['Age'].loc[age_index])\nmedian_age","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.787755Z","iopub.execute_input":"2022-07-11T01:34:05.789041Z","iopub.status.idle":"2022-07-11T01:34:05.799951Z","shell.execute_reply.started":"2022-07-11T01:34:05.788995Z","shell.execute_reply":"2022-07-11T01:34:05.798854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Replacing the missing values of the column - Age in the training dataset\n\ntrain_dataset['Age'].fillna(median_age, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.801279Z","iopub.execute_input":"2022-07-11T01:34:05.802134Z","iopub.status.idle":"2022-07-11T01:34:05.81295Z","shell.execute_reply.started":"2022-07-11T01:34:05.802102Z","shell.execute_reply":"2022-07-11T01:34:05.811953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Finding the median value of the column - Age in the test dataset\n\nage_index = list(~test_dataset['Age'].isnull())\nmedian_age = np.median(test_dataset['Age'].loc[age_index])\nmedian_age","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.814283Z","iopub.execute_input":"2022-07-11T01:34:05.814767Z","iopub.status.idle":"2022-07-11T01:34:05.825154Z","shell.execute_reply.started":"2022-07-11T01:34:05.814732Z","shell.execute_reply":"2022-07-11T01:34:05.824307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Replacing the missing values of the column - Age in the test dataset\n\ntest_dataset['Age'].fillna(median_age, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.826325Z","iopub.execute_input":"2022-07-11T01:34:05.826795Z","iopub.status.idle":"2022-07-11T01:34:05.835334Z","shell.execute_reply.started":"2022-07-11T01:34:05.826768Z","shell.execute_reply":"2022-07-11T01:34:05.834308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Checking if there are any missing values of Age in the training dataset\n\ntrain_dataset['Age'].isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.837174Z","iopub.execute_input":"2022-07-11T01:34:05.837632Z","iopub.status.idle":"2022-07-11T01:34:05.848269Z","shell.execute_reply.started":"2022-07-11T01:34:05.837602Z","shell.execute_reply":"2022-07-11T01:34:05.84742Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Checking if there are any missing values of Age in the test dataset\n\ntest_dataset['Age'].isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.849732Z","iopub.execute_input":"2022-07-11T01:34:05.850664Z","iopub.status.idle":"2022-07-11T01:34:05.862167Z","shell.execute_reply.started":"2022-07-11T01:34:05.850632Z","shell.execute_reply":"2022-07-11T01:34:05.860953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 4.2.2 Handling missing values - Embarked (training set)","metadata":{}},{"cell_type":"markdown","source":"From the analysis on the variable - Embarked, we can use the following decision logic to replace the missing values:\n\n> if (Pclass == 1 and Survived == 1) -> Embarked = C\n>\n> if (Pclass == 3 and Survived == 1) -> Embarked = Q\n>\n> else Embarked = S","metadata":{}},{"cell_type":"code","source":"### Finding the indices of the rows where Embarked is null\n\nindex_values = list(train_dataset['Embarked'].isnull())\nindex_values","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.863599Z","iopub.execute_input":"2022-07-11T01:34:05.863965Z","iopub.status.idle":"2022-07-11T01:34:05.889646Z","shell.execute_reply.started":"2022-07-11T01:34:05.863932Z","shell.execute_reply":"2022-07-11T01:34:05.888476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Replacing the missing values in the column Embarked using the decision logic\n\nfor index in range(len(train_dataset)):\n    if index_values[index]:\n        if train_dataset['Pclass'][index] == 1 and train_dataset['Survived'][index] == 1:\n            train_dataset['Embarked'][index] = 'C'\n        elif train_dataset['Pclass'][index] == 3 and train_dataset['Survived'][index] == 1:\n            train_dataset['Embarked'][index] = 'Q'\n        else:\n            train_dataset['Embarked'][index] = 'S'","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.891468Z","iopub.execute_input":"2022-07-11T01:34:05.892204Z","iopub.status.idle":"2022-07-11T01:34:05.90275Z","shell.execute_reply.started":"2022-07-11T01:34:05.892156Z","shell.execute_reply":"2022-07-11T01:34:05.901688Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Checking if there are any missing values of Embarked in the training dataset\n\ntrain_dataset['Embarked'].isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.903895Z","iopub.execute_input":"2022-07-11T01:34:05.90445Z","iopub.status.idle":"2022-07-11T01:34:05.914717Z","shell.execute_reply.started":"2022-07-11T01:34:05.904419Z","shell.execute_reply":"2022-07-11T01:34:05.913729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 4.2.3 Handling missing values - Fare (test set)","metadata":{}},{"cell_type":"markdown","source":"To replace, the missing values in the column - Fare, we use the medin fare of the data present in the test dataset.","metadata":{}},{"cell_type":"code","source":"### Finding the median value of the column - Fare in the test set\n\nfare_index = list(~test_dataset['Fare'].isnull())\nmedian_fare = np.median(test_dataset['Fare'].loc[fare_index])\nmedian_fare","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.916292Z","iopub.execute_input":"2022-07-11T01:34:05.916986Z","iopub.status.idle":"2022-07-11T01:34:05.928697Z","shell.execute_reply.started":"2022-07-11T01:34:05.916952Z","shell.execute_reply":"2022-07-11T01:34:05.927574Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Replacing the missing values of the column - Fare in the test dataset\n\ntest_dataset['Fare'].fillna(median_fare, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.930599Z","iopub.execute_input":"2022-07-11T01:34:05.931744Z","iopub.status.idle":"2022-07-11T01:34:05.938641Z","shell.execute_reply.started":"2022-07-11T01:34:05.931698Z","shell.execute_reply":"2022-07-11T01:34:05.937592Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Checking if there are any missing values of Fare in the test dataset\n\ntest_dataset['Fare'].isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.939842Z","iopub.execute_input":"2022-07-11T01:34:05.940905Z","iopub.status.idle":"2022-07-11T01:34:05.950555Z","shell.execute_reply.started":"2022-07-11T01:34:05.940872Z","shell.execute_reply":"2022-07-11T01:34:05.949547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Looking if the training dataset has any more missing values apart from Cabin\n\ntrain_dataset.isnull().sum().sort_values(ascending = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.95196Z","iopub.execute_input":"2022-07-11T01:34:05.952709Z","iopub.status.idle":"2022-07-11T01:34:05.969139Z","shell.execute_reply.started":"2022-07-11T01:34:05.952675Z","shell.execute_reply":"2022-07-11T01:34:05.968171Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Looking if the test dataset has any more missing values apart from Cabin\n\ntest_dataset.isnull().sum().sort_values(ascending = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.970575Z","iopub.execute_input":"2022-07-11T01:34:05.971151Z","iopub.status.idle":"2022-07-11T01:34:05.981671Z","shell.execute_reply.started":"2022-07-11T01:34:05.97112Z","shell.execute_reply":"2022-07-11T01:34:05.980585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since, there are no missing values in the data apart from the data in the column - Cabin (which we will deal in the Feature Engineering), we can proceed to perform Feature Engineering.","metadata":{}},{"cell_type":"markdown","source":"#### 4.3 Feature Engineering","metadata":{}},{"cell_type":"markdown","source":"Feature engineering is arguably the most important art in machine learning. It is the process of creating new features from existing features to better represent the underlying problem to the predictive models resulting in improved model accuracy on unseen data.","metadata":{}},{"cell_type":"markdown","source":"Here, we focus on creating new columns for:\n1. NewCabin - using the column Cabin\n2. Title - using the column Name\n3. TotalPassengers - using the columns SibSp, Parch\n4. AgeCategory - using the column Age","metadata":{}},{"cell_type":"markdown","source":"##### 4.3.1 NewCabin - using the column Cabin","metadata":{}},{"cell_type":"markdown","source":"From the training dataset, we can see that there are a lot of missing values in the column - Cabin. Hence, we convert the data in the column to strings i.e., nan -> \"nan\".","metadata":{}},{"cell_type":"code","source":"### Converting the data of the column Cabin to strings\n\ncabin_data = [str(cabin) for cabin in train_dataset['Cabin']]\n\n### Fetching the first character of the cabins other nan\n\nmodified_cabin_data = [cabin[0] if cabin != 'nan' else cabin for cabin in cabin_data]\nset(modified_cabin_data)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.983172Z","iopub.execute_input":"2022-07-11T01:34:05.98412Z","iopub.status.idle":"2022-07-11T01:34:05.994407Z","shell.execute_reply.started":"2022-07-11T01:34:05.984074Z","shell.execute_reply":"2022-07-11T01:34:05.993224Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Adding the modified cabin data to NewCabin column of the training dataset\n\ntrain_dataset['NewCabin'] = modified_cabin_data\ntrain_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:05.995955Z","iopub.execute_input":"2022-07-11T01:34:05.997177Z","iopub.status.idle":"2022-07-11T01:34:06.027902Z","shell.execute_reply.started":"2022-07-11T01:34:05.997131Z","shell.execute_reply":"2022-07-11T01:34:06.02657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Distribution of the column Pclass for each New Cabin type\n\nsns.factorplot('Pclass', col = 'NewCabin', data = train_dataset, kind = 'count', col_wrap = 3)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:06.029631Z","iopub.execute_input":"2022-07-11T01:34:06.030154Z","iopub.status.idle":"2022-07-11T01:34:07.210602Z","shell.execute_reply.started":"2022-07-11T01:34:06.030064Z","shell.execute_reply":"2022-07-11T01:34:07.209563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above graphs, we can see that most of the passengers travelling in the cabins starting with A, B, C, D, E, T belong to Pclass 1. Similarly, for the cabins starting with F belong to Pclass 2 and the cabins nan (unknown) have most of the passengers travelling in Pclass 3.","metadata":{}},{"cell_type":"code","source":"### Modifying the data using the above graphs\n\nnew_cabin_data = []\nfor cabin in modified_cabin_data:\n    if cabin in {'A', 'B', 'C', 'D', 'E', 'T'}:\n        new_cabin_data.append(1)\n    elif cabin == 'F':\n        new_cabin_data.append(2)\n    else:\n        new_cabin_data.append(3)\n        \nnew_cabin_data","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:07.212014Z","iopub.execute_input":"2022-07-11T01:34:07.212371Z","iopub.status.idle":"2022-07-11T01:34:07.238722Z","shell.execute_reply.started":"2022-07-11T01:34:07.21234Z","shell.execute_reply":"2022-07-11T01:34:07.237097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Adding the above data to the training data\n\ntrain_dataset['NewCabin'] = new_cabin_data\ntrain_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:07.240566Z","iopub.execute_input":"2022-07-11T01:34:07.241748Z","iopub.status.idle":"2022-07-11T01:34:07.27328Z","shell.execute_reply.started":"2022-07-11T01:34:07.241708Z","shell.execute_reply":"2022-07-11T01:34:07.271996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we add the column - NewCabin to the test dataset using the same steps followed in the above code blocks.","metadata":{}},{"cell_type":"code","source":"### Converting the data of the column Cabin to strings\n\ncabin_test_data = [str(cabin) for cabin in test_dataset['Cabin']]\n\n### Fetching the first character of the cabins other nan\n\nmodified_test_cabin_data = [cabin[0] if cabin != 'nan' else cabin for cabin in cabin_test_data]\nset(modified_test_cabin_data)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:07.274635Z","iopub.execute_input":"2022-07-11T01:34:07.275129Z","iopub.status.idle":"2022-07-11T01:34:07.284532Z","shell.execute_reply.started":"2022-07-11T01:34:07.275095Z","shell.execute_reply":"2022-07-11T01:34:07.283138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Modifying the data using the above graphs\n\nnew_cabin_test_data = []\nfor cabin in modified_test_cabin_data:\n    if cabin in {'A', 'B', 'C', 'D', 'E', 'T'}:\n        new_cabin_test_data.append(1)\n    elif cabin == 'F':\n        new_cabin_test_data.append(2)\n    else:\n        new_cabin_test_data.append(3)\n        \nnew_cabin_test_data","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:07.286812Z","iopub.execute_input":"2022-07-11T01:34:07.287646Z","iopub.status.idle":"2022-07-11T01:34:07.305243Z","shell.execute_reply.started":"2022-07-11T01:34:07.287598Z","shell.execute_reply":"2022-07-11T01:34:07.304004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Adding the above data to the test data\n\ntest_dataset['NewCabin'] = new_cabin_test_data\ntest_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:07.306609Z","iopub.execute_input":"2022-07-11T01:34:07.306925Z","iopub.status.idle":"2022-07-11T01:34:07.339139Z","shell.execute_reply.started":"2022-07-11T01:34:07.306896Z","shell.execute_reply":"2022-07-11T01:34:07.338016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since, we now added the column - NewCabin to both the datasets, we can remove the original column - Cabin from the datasets.","metadata":{}},{"cell_type":"code","source":"### Dropping the column - NewCabin from the training dataset\n\ntrain_dataset.drop(['Cabin'], axis = 1, inplace = True)\ntrain_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:07.34066Z","iopub.execute_input":"2022-07-11T01:34:07.340993Z","iopub.status.idle":"2022-07-11T01:34:07.370829Z","shell.execute_reply.started":"2022-07-11T01:34:07.340964Z","shell.execute_reply":"2022-07-11T01:34:07.369995Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Dropping the column - NewCabin from the test dataset\n\ntest_dataset.drop(['Cabin'], axis = 1, inplace = True)\ntest_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:07.371999Z","iopub.execute_input":"2022-07-11T01:34:07.372935Z","iopub.status.idle":"2022-07-11T01:34:07.402117Z","shell.execute_reply.started":"2022-07-11T01:34:07.3729Z","shell.execute_reply":"2022-07-11T01:34:07.401013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 4.3.2 Title - using the column Name","metadata":{}},{"cell_type":"code","source":"### Fetching the title from the name of a passenger\n\nname_data = train_dataset['Name']\ntitles = []\nfor name in name_data:\n    name_split = name.split(',')[1]\n    title = name_split[1 : name_split.index('.')]\n    titles.append(title)\n    \n### Looking at the unique titles\n\nset(titles)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:07.403463Z","iopub.execute_input":"2022-07-11T01:34:07.403898Z","iopub.status.idle":"2022-07-11T01:34:07.413516Z","shell.execute_reply.started":"2022-07-11T01:34:07.403869Z","shell.execute_reply":"2022-07-11T01:34:07.412123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we will categorize these titles to Special and Regular (Mr, Master, Miss, Mrs, Ms) and will be coded such that Special means 1 and Regular means 0, as if there is any other value to Regular it simply correlates with the column - Sex.","metadata":{}},{"cell_type":"code","source":"### Categorize titles to 2 categories - Special (1) and Regular (0)\n\ntitle_data = []\nfor title in titles:\n    if title in {'Mr', 'Master', 'Mrs', 'Miss', 'Ms'}:\n        title_data.append(0)\n    else:\n        title_data.append(1)\n        \n### Adding the title to the dataset\n\ntrain_dataset['Title'] = title_data\ntrain_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:07.414865Z","iopub.execute_input":"2022-07-11T01:34:07.415864Z","iopub.status.idle":"2022-07-11T01:34:07.447881Z","shell.execute_reply.started":"2022-07-11T01:34:07.415827Z","shell.execute_reply":"2022-07-11T01:34:07.446727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Mean of survival by Title\n\ntitle_survived = train_dataset[['Title', 'Survived']].groupby('Title', as_index = False).mean()\ntitle_survived","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:07.449672Z","iopub.execute_input":"2022-07-11T01:34:07.450089Z","iopub.status.idle":"2022-07-11T01:34:07.464652Z","shell.execute_reply.started":"2022-07-11T01:34:07.450029Z","shell.execute_reply":"2022-07-11T01:34:07.463503Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Survival Probability by Title\n\nsns.barplot(title_survived['Title'], title_survived['Survived'], alpha = 0.8)\nplt.title('Survival Probability by Title')\nplt.ylabel('Probability', fontsize = 12)\nplt.xlabel('Title', fontsize = 12)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:07.466179Z","iopub.execute_input":"2022-07-11T01:34:07.466912Z","iopub.status.idle":"2022-07-11T01:34:07.646247Z","shell.execute_reply.started":"2022-07-11T01:34:07.46688Z","shell.execute_reply":"2022-07-11T01:34:07.64514Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above graph, we can see that probability of survival in Female_Regular is higher than Male_Regular and Special. Now we add the column - Title to the test dataset using the same steps followed in the above code blocks.","metadata":{}},{"cell_type":"code","source":"### Fetching the title from the name of a passenger\n\nname_test_data = test_dataset['Name']\ntitles = []\nfor name in name_test_data:\n    name_split = name.split(',')[1]\n    title = name_split[1 : name_split.index('.')]\n    titles.append(title)\n    \n### Looking at the unique titles\n\nset(titles)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:07.647863Z","iopub.execute_input":"2022-07-11T01:34:07.648171Z","iopub.status.idle":"2022-07-11T01:34:07.657902Z","shell.execute_reply.started":"2022-07-11T01:34:07.648144Z","shell.execute_reply":"2022-07-11T01:34:07.656717Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Categorize titles to 2 categories - Special (1) and Regular (0)\n\ntitle_data = []\nfor title in titles:\n    if title in {'Mr', 'Master', 'Mrs', 'Miss', 'Ms'}:\n        title_data.append(0)\n    else:\n        title_data.append(1)\n        \n### Adding the title to the dataset\n\ntest_dataset['Title'] = title_data\ntest_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:07.66002Z","iopub.execute_input":"2022-07-11T01:34:07.660448Z","iopub.status.idle":"2022-07-11T01:34:07.689372Z","shell.execute_reply.started":"2022-07-11T01:34:07.660405Z","shell.execute_reply":"2022-07-11T01:34:07.688275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since, we now added the column - Title to both the datasets, we will now remove the parent column - Name from both the training and test datasets.","metadata":{}},{"cell_type":"code","source":"### Dropping the column - Name from the training dataset\n\ntrain_dataset.drop(['Name'], axis = 1, inplace = True)\ntrain_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:07.691107Z","iopub.execute_input":"2022-07-11T01:34:07.691442Z","iopub.status.idle":"2022-07-11T01:34:07.717163Z","shell.execute_reply.started":"2022-07-11T01:34:07.691413Z","shell.execute_reply":"2022-07-11T01:34:07.716069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Dropping the column - Name from the test dataset\n\ntest_dataset.drop(['Name'], axis = 1, inplace = True)\ntest_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:07.718583Z","iopub.execute_input":"2022-07-11T01:34:07.71958Z","iopub.status.idle":"2022-07-11T01:34:07.74477Z","shell.execute_reply.started":"2022-07-11T01:34:07.719543Z","shell.execute_reply":"2022-07-11T01:34:07.743227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 4.3.3 TotalPassengers - using the columns SibSp, Parch","metadata":{}},{"cell_type":"code","source":"### Adding the total passengers using SibSp + Parch + 1\n\ntrain_dataset['TotalPassengers'] = train_dataset['SibSp'] + train_dataset['Parch'] + 1\ntrain_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:07.74676Z","iopub.execute_input":"2022-07-11T01:34:07.747118Z","iopub.status.idle":"2022-07-11T01:34:07.776466Z","shell.execute_reply.started":"2022-07-11T01:34:07.747084Z","shell.execute_reply":"2022-07-11T01:34:07.7753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Mean of survival by TotalPassengers\n\ntotal_passengers_survived = train_dataset[['TotalPassengers', 'Survived']].groupby('TotalPassengers', as_index = False).mean()\ntotal_passengers_survived","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:07.778135Z","iopub.execute_input":"2022-07-11T01:34:07.77891Z","iopub.status.idle":"2022-07-11T01:34:07.793419Z","shell.execute_reply.started":"2022-07-11T01:34:07.778874Z","shell.execute_reply":"2022-07-11T01:34:07.792532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Survival Probability by TotalPassengers\n\nsns.barplot(total_passengers_survived['TotalPassengers'], total_passengers_survived['Survived'], alpha = 0.8)\nplt.title('Survival Probability by Total Passengers')\nplt.ylabel('Probability', fontsize = 12)\nplt.xlabel('Total Passengers', fontsize = 12)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:07.795465Z","iopub.execute_input":"2022-07-11T01:34:07.79666Z","iopub.status.idle":"2022-07-11T01:34:08.016218Z","shell.execute_reply.started":"2022-07-11T01:34:07.79651Z","shell.execute_reply":"2022-07-11T01:34:08.015443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above graph, we can say that probability of survival is high in the case where Total Passengers travelling are either 2, 3, or 4. Now we add the column - TotalPassengers to the test dataset using the same steps followed in the above code blocks.","metadata":{}},{"cell_type":"code","source":"### Adding the total passengers using SibSp + Parch + 1 to the test dataset\n\ntest_dataset['TotalPassengers'] = test_dataset['SibSp'] + test_dataset['Parch'] + 1\ntest_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:08.017242Z","iopub.execute_input":"2022-07-11T01:34:08.017754Z","iopub.status.idle":"2022-07-11T01:34:08.047196Z","shell.execute_reply.started":"2022-07-11T01:34:08.017719Z","shell.execute_reply":"2022-07-11T01:34:08.046191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since, we now added the column - TotalPassengers to both the datasets, we will now remove the parent columns - SibSp and Parch from both the training and test datasets.","metadata":{}},{"cell_type":"code","source":"### Dropping the columns - SibSp, Parch from the training dataset\n\ntrain_dataset.drop(['SibSp', 'Parch'], axis = 1, inplace = True)\ntrain_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:08.048634Z","iopub.execute_input":"2022-07-11T01:34:08.049301Z","iopub.status.idle":"2022-07-11T01:34:08.075838Z","shell.execute_reply.started":"2022-07-11T01:34:08.049255Z","shell.execute_reply":"2022-07-11T01:34:08.074673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Dropping the columns - SibSp, Parch from the test dataset\n\ntest_dataset.drop(['SibSp', 'Parch'], axis = 1, inplace = True)\ntest_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:08.0773Z","iopub.execute_input":"2022-07-11T01:34:08.077915Z","iopub.status.idle":"2022-07-11T01:34:08.105127Z","shell.execute_reply.started":"2022-07-11T01:34:08.077873Z","shell.execute_reply":"2022-07-11T01:34:08.104229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 4.3.4 AgeCategory - using the column Age","metadata":{}},{"cell_type":"markdown","source":"It was rumored that during the Titanic evacuation Females and Kids were given the first priority. We already know that half of the above statement is true as Females have a higher probability of survival. Now, we will verify the second part of the statement is a tautology or not. \n\nHere, we will keep an age threshold of 18 such that if the age of a passenger is less than 18 that passenger is young or else an adult. ","metadata":{}},{"cell_type":"code","source":"### Creating age slabs for the training data\n\nage_data = train_dataset['Age']\nslabs = []\nfor age in age_data:\n    if age < 18:\n        slabs.append('young')\n    else:\n        slabs.append('adult')\n\n### Adding the age slabs to the training dataset\n\ntrain_dataset['AgeSlab'] = slabs\ntrain_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:08.106704Z","iopub.execute_input":"2022-07-11T01:34:08.107016Z","iopub.status.idle":"2022-07-11T01:34:08.13408Z","shell.execute_reply.started":"2022-07-11T01:34:08.106989Z","shell.execute_reply":"2022-07-11T01:34:08.133188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Mean of survival by AgeSlab\n\nage_slab_survived = train_dataset[['AgeSlab', 'Survived']].groupby('AgeSlab', as_index = False).mean()\nage_slab_survived","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:08.135157Z","iopub.execute_input":"2022-07-11T01:34:08.136218Z","iopub.status.idle":"2022-07-11T01:34:08.150479Z","shell.execute_reply.started":"2022-07-11T01:34:08.136181Z","shell.execute_reply":"2022-07-11T01:34:08.149672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Survival Probability by AgeSlab\n\nsns.barplot(age_slab_survived['AgeSlab'], age_slab_survived['Survived'], alpha = 0.8)\nplt.title('Survival Probability by Age Slab')\nplt.ylabel('Probability', fontsize = 12)\nplt.xlabel('Age Slab', fontsize = 12)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:08.151614Z","iopub.execute_input":"2022-07-11T01:34:08.152519Z","iopub.status.idle":"2022-07-11T01:34:08.32364Z","shell.execute_reply.started":"2022-07-11T01:34:08.152465Z","shell.execute_reply":"2022-07-11T01:34:08.322674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above graph, we can say that probability of survival is higher in younger passengers than compared to adult passengers. Now we will add the AgeSlab column to the test dataset using the above code cells.","metadata":{}},{"cell_type":"code","source":"### Creating age slabs for the test data\n\nage_data = test_dataset['Age']\nslabs = []\nfor age in age_data:\n    if age < 18:\n        slabs.append('young')\n    else:\n        slabs.append('adult')\n\n### Adding the age slabs to the test dataset\n\ntest_dataset['AgeSlab'] = slabs\ntest_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:08.325749Z","iopub.execute_input":"2022-07-11T01:34:08.32638Z","iopub.status.idle":"2022-07-11T01:34:08.358375Z","shell.execute_reply.started":"2022-07-11T01:34:08.326333Z","shell.execute_reply":"2022-07-11T01:34:08.357308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since, we added the column - AgeSlab in both the datasets, we will now remove the parent column - Age from both the training and test datasets.","metadata":{}},{"cell_type":"code","source":"### Dropping the column - Age from the training dataset\n\ntrain_dataset.drop(['Age'], axis = 1, inplace = True)\ntrain_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:08.359806Z","iopub.execute_input":"2022-07-11T01:34:08.360225Z","iopub.status.idle":"2022-07-11T01:34:08.384967Z","shell.execute_reply.started":"2022-07-11T01:34:08.360193Z","shell.execute_reply":"2022-07-11T01:34:08.383919Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Dropping the column - Age from the test dataset\n\ntest_dataset.drop(['Age'], axis = 1, inplace = True)\ntest_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:08.386569Z","iopub.execute_input":"2022-07-11T01:34:08.386916Z","iopub.status.idle":"2022-07-11T01:34:08.413283Z","shell.execute_reply.started":"2022-07-11T01:34:08.386888Z","shell.execute_reply":"2022-07-11T01:34:08.411927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's look at the correlation between the numerical values of the training dataset. ","metadata":{}},{"cell_type":"code","source":"### Plotting the correlation between various columns of the training set\n\nplt.figure(figsize = (6, 6))\nheatmap = sns.heatmap(train_dataset.corr(), vmin = -1, vmax = 1, annot = True)\nheatmap.set_title('Correlation Heatmap', fontdict = {'fontsize' : 12}, pad = 12)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:08.416522Z","iopub.execute_input":"2022-07-11T01:34:08.417272Z","iopub.status.idle":"2022-07-11T01:34:08.910859Z","shell.execute_reply.started":"2022-07-11T01:34:08.417227Z","shell.execute_reply":"2022-07-11T01:34:08.91005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above matrix, we can see that the columns - NewCabin and Fare have a negative correlation between them.","metadata":{}},{"cell_type":"markdown","source":"#### 4.4 Data Transformation","metadata":{}},{"cell_type":"markdown","source":"In this section, we will remove the high positive degree of skewness present in the column - Fare by using a log transform on the data.","metadata":{}},{"cell_type":"code","source":"### Understanding the distribution of the column - Fare\n\nsns.distplot(train_dataset['Fare'], label = 'Skewness: %.2f'%(train_dataset['Fare'].skew()))\nplt.legend(loc = 'best')\nplt.title('Passenger Fare Distribution')","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:08.912283Z","iopub.execute_input":"2022-07-11T01:34:08.913348Z","iopub.status.idle":"2022-07-11T01:34:09.247758Z","shell.execute_reply.started":"2022-07-11T01:34:08.913308Z","shell.execute_reply":"2022-07-11T01:34:09.245635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Understanding the distribution of the data log(Fare)\n\nmodified_fare = [np.log(fare) if fare > 0 else fare for fare in train_dataset['Fare']]\ntrain_dataset['Fare'] = modified_fare\n\nsns.distplot(train_dataset['Fare'], label = 'Skewness: %.2f'%(train_dataset['Fare'].skew()))\nplt.legend(loc = 'best')\nplt.title('Passenger Fare Distribution')","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.248944Z","iopub.execute_input":"2022-07-11T01:34:09.249245Z","iopub.status.idle":"2022-07-11T01:34:09.534038Z","shell.execute_reply.started":"2022-07-11T01:34:09.249219Z","shell.execute_reply":"2022-07-11T01:34:09.532774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above graph, we can see that the degree of skewness is significantly reduced than compared to the skewness in the original data. Similarly, we transform the data present in the column Fare of the test data.","metadata":{}},{"cell_type":"code","source":"### Modifying the column - Fare in the test data\n\nmodified_fare = [np.log(fare) if fare > 0 else fare for fare in test_dataset['Fare']]\ntest_dataset['Fare'] = modified_fare","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.535729Z","iopub.execute_input":"2022-07-11T01:34:09.536062Z","iopub.status.idle":"2022-07-11T01:34:09.543259Z","shell.execute_reply.started":"2022-07-11T01:34:09.536032Z","shell.execute_reply":"2022-07-11T01:34:09.541957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 4.5 Feature Encoding","metadata":{}},{"cell_type":"markdown","source":"Feature encoding is the process of turning categorical data in a dataset into numerical data. It is essential that we perform feature encoding because most machine learning models can only interpret numerical data and not data in text form.","metadata":{}},{"cell_type":"markdown","source":"First we will encode the column Ticket such that if the Ticket value starts with a number it belongs to class 1, else class 2.","metadata":{}},{"cell_type":"code","source":"### Encoding the column - Ticket\n\nticket_data = train_dataset['Ticket']\nnew_ticket_data = []\n\nfor ticket in ticket_data:\n    if ticket[0].isdigit():\n        new_ticket_data.append(1)\n    else:\n        new_ticket_data.append(2)\n        \nnew_ticket_data","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.54492Z","iopub.execute_input":"2022-07-11T01:34:09.545254Z","iopub.status.idle":"2022-07-11T01:34:09.57184Z","shell.execute_reply.started":"2022-07-11T01:34:09.545223Z","shell.execute_reply":"2022-07-11T01:34:09.570463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Replacing the column - Ticket with the new data\n\ntrain_dataset['Ticket'] = new_ticket_data\ntrain_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.573535Z","iopub.execute_input":"2022-07-11T01:34:09.574626Z","iopub.status.idle":"2022-07-11T01:34:09.598572Z","shell.execute_reply.started":"2022-07-11T01:34:09.574588Z","shell.execute_reply":"2022-07-11T01:34:09.59767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we will modify the Ticket column in the test dataset using the same steps as done above.","metadata":{}},{"cell_type":"code","source":"### Encoding the column - Ticket\n\nticket_data = test_dataset['Ticket']\nnew_ticket_data = []\n\nfor ticket in ticket_data:\n    if ticket[0].isdigit():\n        new_ticket_data.append(1)\n    else:\n        new_ticket_data.append(2)\n        \nnew_ticket_data","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.599967Z","iopub.execute_input":"2022-07-11T01:34:09.600833Z","iopub.status.idle":"2022-07-11T01:34:09.621629Z","shell.execute_reply.started":"2022-07-11T01:34:09.600796Z","shell.execute_reply":"2022-07-11T01:34:09.620438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Replacing the column - Ticket with the new data\n\ntest_dataset['Ticket'] = new_ticket_data\ntest_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.623659Z","iopub.execute_input":"2022-07-11T01:34:09.624259Z","iopub.status.idle":"2022-07-11T01:34:09.646904Z","shell.execute_reply.started":"2022-07-11T01:34:09.624225Z","shell.execute_reply":"2022-07-11T01:34:09.646049Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we perform One Hot Encoding on the columns - Pclass, Sex, Ticket, Embarked, NewCabin, AgeSlab for both the Training and Test datasets.","metadata":{}},{"cell_type":"code","source":"### One Hot Encoding the columns - Pclass, Sex, Embarked, NewCabin, AgeSlab of the training set\n\nencoded_train_dataset = pd.get_dummies(data = train_dataset, \n                                       columns = ['Pclass', 'Sex', 'Embarked', 'NewCabin', 'AgeSlab'])\nencoded_train_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.648257Z","iopub.execute_input":"2022-07-11T01:34:09.649418Z","iopub.status.idle":"2022-07-11T01:34:09.681348Z","shell.execute_reply.started":"2022-07-11T01:34:09.649379Z","shell.execute_reply":"2022-07-11T01:34:09.680199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### One Hot Encoding the columns - Pclass, Sex, Embarked, NewCabin, AgeSlab of the test set\n\nencoded_test_dataset = pd.get_dummies(data = test_dataset, \n                                       columns = ['Pclass', 'Sex', 'Embarked', 'NewCabin', 'AgeSlab'])\nencoded_test_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.682843Z","iopub.execute_input":"2022-07-11T01:34:09.683178Z","iopub.status.idle":"2022-07-11T01:34:09.712601Z","shell.execute_reply.started":"2022-07-11T01:34:09.683149Z","shell.execute_reply":"2022-07-11T01:34:09.711424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now our datasets are ready for modelling.","metadata":{}},{"cell_type":"markdown","source":"### 5. Modelling","metadata":{}},{"cell_type":"markdown","source":"Scikit-learn is one of the most popular libraries for machine learning in Python and that is what we will use in the modelling part of this project.\n\nSince Titanic is a classfication problem, we will need to use classfication models, also known as classifiers, to train on our model to make predictions. I highly recommend checking out the scikit-learn documentation for more information on the different machine learning models available in their library. I have chosen the following classifiers for the job:\n\n1. Logistic regression\n2. Support vector classification\n3. K-nearest neighbours\n4. Naive Bayes\n5. Decision tree\n6. Random forest\n\nIn this section of the notebook, I will fit the models to the training set as outlined above and evaluate their accuracy at making predictions. Once the best model is determined, I will also do hyperparameter tuning to further boost the performance of the best model.","metadata":{}},{"cell_type":"markdown","source":"#### 5.1 Splitting the Training data ","metadata":{}},{"cell_type":"markdown","source":"Here, we will split the training data into X_train, X_test, Y_train, and Y_test so that they can be fed to the machine learning models that are used in the next section. Then the model with the best performance will be used to predict the result on the given test dataset.","metadata":{}},{"cell_type":"code","source":"### Splitting the data to the matrices X and Y using the training set.\n\nX = encoded_train_dataset.iloc[:, 1:].values\nY = encoded_train_dataset.iloc[:, 0].values","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.714098Z","iopub.execute_input":"2022-07-11T01:34:09.714657Z","iopub.status.idle":"2022-07-11T01:34:09.723185Z","shell.execute_reply.started":"2022-07-11T01:34:09.714624Z","shell.execute_reply":"2022-07-11T01:34:09.722057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Looking at the new training data - X\n\nX","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.724991Z","iopub.execute_input":"2022-07-11T01:34:09.725999Z","iopub.status.idle":"2022-07-11T01:34:09.733751Z","shell.execute_reply.started":"2022-07-11T01:34:09.725962Z","shell.execute_reply":"2022-07-11T01:34:09.732571Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Looking at the new test data - Y\n\nY","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.735021Z","iopub.execute_input":"2022-07-11T01:34:09.735447Z","iopub.status.idle":"2022-07-11T01:34:09.750898Z","shell.execute_reply.started":"2022-07-11T01:34:09.735418Z","shell.execute_reply":"2022-07-11T01:34:09.749651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Dividing the dataset into train and test in the ratio of 80 : 20\n\nX_train, X_test, Y_train, Y_test = train_test_split(X, Y, test_size = 0.2, random_state = 27, shuffle = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.752603Z","iopub.execute_input":"2022-07-11T01:34:09.753622Z","iopub.status.idle":"2022-07-11T01:34:09.763531Z","shell.execute_reply.started":"2022-07-11T01:34:09.75348Z","shell.execute_reply":"2022-07-11T01:34:09.762376Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.764863Z","iopub.execute_input":"2022-07-11T01:34:09.765721Z","iopub.status.idle":"2022-07-11T01:34:09.776252Z","shell.execute_reply.started":"2022-07-11T01:34:09.765681Z","shell.execute_reply":"2022-07-11T01:34:09.775088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.778074Z","iopub.execute_input":"2022-07-11T01:34:09.778742Z","iopub.status.idle":"2022-07-11T01:34:09.786522Z","shell.execute_reply.started":"2022-07-11T01:34:09.778707Z","shell.execute_reply":"2022-07-11T01:34:09.78546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Y_train","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.787687Z","iopub.execute_input":"2022-07-11T01:34:09.788478Z","iopub.status.idle":"2022-07-11T01:34:09.799886Z","shell.execute_reply.started":"2022-07-11T01:34:09.788441Z","shell.execute_reply":"2022-07-11T01:34:09.798892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Y_test","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.801504Z","iopub.execute_input":"2022-07-11T01:34:09.802879Z","iopub.status.idle":"2022-07-11T01:34:09.811797Z","shell.execute_reply.started":"2022-07-11T01:34:09.80284Z","shell.execute_reply":"2022-07-11T01:34:09.81058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we convert the original test dataset to a numpy array.","metadata":{}},{"cell_type":"code","source":"### Converting the original test dataset to a numpy array\n\nreal_test_data = encoded_test_dataset.iloc[:, :].values\nreal_test_data","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.813021Z","iopub.execute_input":"2022-07-11T01:34:09.813778Z","iopub.status.idle":"2022-07-11T01:34:09.822469Z","shell.execute_reply.started":"2022-07-11T01:34:09.813736Z","shell.execute_reply":"2022-07-11T01:34:09.821638Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 5.2 Fit Model","metadata":{}},{"cell_type":"markdown","source":"In this section, we use various machine learning models to predict the results for our sample test data (X_test). Based on the performance of each model, we select one best model to predict the results on the original test data (real_test_data). We will store the model and its accuracy so that we can tabulate them later for choosing the best model.","metadata":{}},{"cell_type":"code","source":"### Dictionary to store model and its accuracy\n\nmodel_performance = OrderedDict()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.823531Z","iopub.execute_input":"2022-07-11T01:34:09.824669Z","iopub.status.idle":"2022-07-11T01:34:09.829057Z","shell.execute_reply.started":"2022-07-11T01:34:09.824637Z","shell.execute_reply":"2022-07-11T01:34:09.828295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 5.2.1 Applying Logistic Regression","metadata":{}},{"cell_type":"code","source":"### Training the Logistic Regression model on the dataset\n\nlogistic_classifier = LogisticRegression(random_state = 27)\nlogistic_classifier.fit(X_train, Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.830342Z","iopub.execute_input":"2022-07-11T01:34:09.830928Z","iopub.status.idle":"2022-07-11T01:34:09.873116Z","shell.execute_reply.started":"2022-07-11T01:34:09.830891Z","shell.execute_reply":"2022-07-11T01:34:09.871794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Predicting the Test set results\n\nY_pred = logistic_classifier.predict(X_test)\nprint(np.concatenate((Y_pred.reshape(len(Y_pred), 1), Y_test.reshape(len(Y_test), 1)), 1))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.875277Z","iopub.execute_input":"2022-07-11T01:34:09.876029Z","iopub.status.idle":"2022-07-11T01:34:09.888652Z","shell.execute_reply.started":"2022-07-11T01:34:09.875977Z","shell.execute_reply":"2022-07-11T01:34:09.887268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Making the confusion matrix\n\ncm = confusion_matrix(Y_test, Y_pred)\nprint(cm)\n\n### Printing the accuracy of the model\n\nlogistic_accuracy = round(100 * accuracy_score(Y_test, Y_pred), 2)\nmodel_performance['Logistic Regression'] = logistic_accuracy\n\nprint('The accuracy of this model is {} %.'.format(logistic_accuracy))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.89045Z","iopub.execute_input":"2022-07-11T01:34:09.891233Z","iopub.status.idle":"2022-07-11T01:34:09.903784Z","shell.execute_reply.started":"2022-07-11T01:34:09.891179Z","shell.execute_reply":"2022-07-11T01:34:09.902546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 5.2.2 Applying Support Vector Classification - Linear","metadata":{}},{"cell_type":"code","source":"### Applying Linear SVM Classification model\n\nlinear_svm_classifier = SVC(kernel = 'linear', random_state = 27)\nlinear_svm_classifier.fit(X_train, Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.905432Z","iopub.execute_input":"2022-07-11T01:34:09.906156Z","iopub.status.idle":"2022-07-11T01:34:09.942734Z","shell.execute_reply.started":"2022-07-11T01:34:09.906114Z","shell.execute_reply":"2022-07-11T01:34:09.941381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Predicting the Test set results\n\nY_pred = linear_svm_classifier.predict(X_test)\nprint(np.concatenate((Y_pred.reshape(len(Y_pred), 1), Y_test.reshape(len(Y_test), 1)), 1))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.944525Z","iopub.execute_input":"2022-07-11T01:34:09.945309Z","iopub.status.idle":"2022-07-11T01:34:09.964144Z","shell.execute_reply.started":"2022-07-11T01:34:09.945256Z","shell.execute_reply":"2022-07-11T01:34:09.962763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Making the confusion matrix\n\ncm = confusion_matrix(Y_test, Y_pred)\nprint(cm)\n\n### Printing the accuracy of the model\n\nlinear_svc_accuracy = round(100 * accuracy_score(Y_test, Y_pred), 2)\nmodel_performance['Linear SVC'] = linear_svc_accuracy\nprint('The accuracy of this model is {} %.'.format(linear_svc_accuracy))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.966344Z","iopub.execute_input":"2022-07-11T01:34:09.967256Z","iopub.status.idle":"2022-07-11T01:34:09.979221Z","shell.execute_reply.started":"2022-07-11T01:34:09.967186Z","shell.execute_reply":"2022-07-11T01:34:09.97769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 5.2.3 Applying Support Vector Classification - Kernel","metadata":{}},{"cell_type":"code","source":"### Applying Kernel SVM Classification model\n\nkernel_svm_classifier = SVC(kernel = 'rbf', random_state = 27)\nkernel_svm_classifier.fit(X_train, Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:09.981675Z","iopub.execute_input":"2022-07-11T01:34:09.982705Z","iopub.status.idle":"2022-07-11T01:34:10.012336Z","shell.execute_reply.started":"2022-07-11T01:34:09.98265Z","shell.execute_reply":"2022-07-11T01:34:10.010842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Predicting the Test set results\n\nY_pred = kernel_svm_classifier.predict(X_test)\nprint(np.concatenate((Y_pred.reshape(len(Y_pred), 1), Y_test.reshape(len(Y_test), 1)), 1))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.014101Z","iopub.execute_input":"2022-07-11T01:34:10.015005Z","iopub.status.idle":"2022-07-11T01:34:10.027925Z","shell.execute_reply.started":"2022-07-11T01:34:10.014973Z","shell.execute_reply":"2022-07-11T01:34:10.026701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Making the confusion matrix\n\ncm = confusion_matrix(Y_test, Y_pred)\nprint(cm)\n\n### Printing the accuracy of the model\n\nkernel_svc_accuracy = round(100 * accuracy_score(Y_test, Y_pred), 2)\nmodel_performance['Kernel SVC'] = kernel_svc_accuracy\nprint('The accuracy of this model is {} %.'.format(kernel_svc_accuracy))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.029373Z","iopub.execute_input":"2022-07-11T01:34:10.030002Z","iopub.status.idle":"2022-07-11T01:34:10.03868Z","shell.execute_reply.started":"2022-07-11T01:34:10.029969Z","shell.execute_reply":"2022-07-11T01:34:10.03736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 5.2.4 Applying K-Nearest Neighbors (1-NN)","metadata":{}},{"cell_type":"code","source":"### Applying 1NN model\n\nclassifier_1nn = KNeighborsClassifier(n_neighbors = 1, algorithm = 'auto', p = 2, metric = 'minkowski')\nclassifier_1nn.fit(X_train, Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.040287Z","iopub.execute_input":"2022-07-11T01:34:10.041278Z","iopub.status.idle":"2022-07-11T01:34:10.049905Z","shell.execute_reply.started":"2022-07-11T01:34:10.041242Z","shell.execute_reply":"2022-07-11T01:34:10.048803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Predicting the Test set results\n\nY_pred = classifier_1nn.predict(X_test)\nprint(np.concatenate((Y_pred.reshape(len(Y_pred), 1), Y_test.reshape(len(Y_test), 1)), 1))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.051226Z","iopub.execute_input":"2022-07-11T01:34:10.052372Z","iopub.status.idle":"2022-07-11T01:34:10.078023Z","shell.execute_reply.started":"2022-07-11T01:34:10.052322Z","shell.execute_reply":"2022-07-11T01:34:10.076732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Making the confusion matrix\n\ncm = confusion_matrix(Y_test, Y_pred)\nprint(cm)\n\n### Printing the accuracy of the model\n\nnn1_accuracy = round(100 * accuracy_score(Y_test, Y_pred), 2)\nmodel_performance['1 Nearest Neighbors'] = nn1_accuracy\nprint('The accuracy of this model is {} %.'.format(nn1_accuracy))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.11031Z","iopub.execute_input":"2022-07-11T01:34:10.111336Z","iopub.status.idle":"2022-07-11T01:34:10.12316Z","shell.execute_reply.started":"2022-07-11T01:34:10.111279Z","shell.execute_reply":"2022-07-11T01:34:10.121677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 5.2.5 Applying K-Nearest Neighbors (3-NN)","metadata":{}},{"cell_type":"code","source":"### Applying 3NN model\n\nclassifier_3nn = KNeighborsClassifier(n_neighbors = 3, algorithm = 'auto', p = 2, metric = 'minkowski')\nclassifier_3nn.fit(X_train, Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.125636Z","iopub.execute_input":"2022-07-11T01:34:10.1267Z","iopub.status.idle":"2022-07-11T01:34:10.138405Z","shell.execute_reply.started":"2022-07-11T01:34:10.126649Z","shell.execute_reply":"2022-07-11T01:34:10.136914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Predicting the Test set results\n\nY_pred = classifier_3nn.predict(X_test)\nprint(np.concatenate((Y_pred.reshape(len(Y_pred), 1), Y_test.reshape(len(Y_test), 1)), 1))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.141187Z","iopub.execute_input":"2022-07-11T01:34:10.142157Z","iopub.status.idle":"2022-07-11T01:34:10.17175Z","shell.execute_reply.started":"2022-07-11T01:34:10.142105Z","shell.execute_reply":"2022-07-11T01:34:10.170464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Making the confusion matrix\n\ncm = confusion_matrix(Y_test, Y_pred)\nprint(cm)\n\n### Printing the accuracy of the model\n\nnn3_accuracy = round(100 * accuracy_score(Y_test, Y_pred), 2)\nmodel_performance['3 Nearest Neighbors'] = nn3_accuracy\nprint('The accuracy of this model is {} %.'.format(nn3_accuracy))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.173955Z","iopub.execute_input":"2022-07-11T01:34:10.174907Z","iopub.status.idle":"2022-07-11T01:34:10.186626Z","shell.execute_reply.started":"2022-07-11T01:34:10.174854Z","shell.execute_reply":"2022-07-11T01:34:10.185051Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 5.2.6 Applying K-Nearest Neighbors (5-NN)","metadata":{}},{"cell_type":"code","source":"### Applying 5NN model\n\nclassifier_5nn = KNeighborsClassifier(n_neighbors = 5, algorithm = 'auto', p = 2, metric = 'minkowski')\nclassifier_5nn.fit(X_train, Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.189295Z","iopub.execute_input":"2022-07-11T01:34:10.19029Z","iopub.status.idle":"2022-07-11T01:34:10.202924Z","shell.execute_reply.started":"2022-07-11T01:34:10.190239Z","shell.execute_reply":"2022-07-11T01:34:10.201657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Predicting the Test set results\n\nY_pred = classifier_5nn.predict(X_test)\nprint(np.concatenate((Y_pred.reshape(len(Y_pred), 1), Y_test.reshape(len(Y_test), 1)), 1))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.204844Z","iopub.execute_input":"2022-07-11T01:34:10.20552Z","iopub.status.idle":"2022-07-11T01:34:10.238514Z","shell.execute_reply.started":"2022-07-11T01:34:10.205461Z","shell.execute_reply":"2022-07-11T01:34:10.237091Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Making the confusion matrix\n\ncm = confusion_matrix(Y_test, Y_pred)\nprint(cm)\n\n### Printing the accuracy of the model\n\nnn5_accuracy = round(100 * accuracy_score(Y_test, Y_pred), 2)\nmodel_performance['5 Nearest Neighbors'] = nn5_accuracy\nprint('The accuracy of this model is {} %.'.format(nn5_accuracy))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.240912Z","iopub.execute_input":"2022-07-11T01:34:10.241929Z","iopub.status.idle":"2022-07-11T01:34:10.254423Z","shell.execute_reply.started":"2022-07-11T01:34:10.241873Z","shell.execute_reply":"2022-07-11T01:34:10.252903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 5.2.7 Applying K-Nearest Neighbors (7-NN)","metadata":{}},{"cell_type":"code","source":"### Applying 7NN model\n\nclassifier_7nn = KNeighborsClassifier(n_neighbors = 7, algorithm = 'auto', p = 2, metric = 'minkowski')\nclassifier_7nn.fit(X_train, Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.257297Z","iopub.execute_input":"2022-07-11T01:34:10.258317Z","iopub.status.idle":"2022-07-11T01:34:10.270562Z","shell.execute_reply.started":"2022-07-11T01:34:10.258261Z","shell.execute_reply":"2022-07-11T01:34:10.268983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Predicting the Test set results\n\nY_pred = classifier_7nn.predict(X_test)\nprint(np.concatenate((Y_pred.reshape(len(Y_pred), 1), Y_test.reshape(len(Y_test), 1)), 1))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.272954Z","iopub.execute_input":"2022-07-11T01:34:10.274059Z","iopub.status.idle":"2022-07-11T01:34:10.303007Z","shell.execute_reply.started":"2022-07-11T01:34:10.274007Z","shell.execute_reply":"2022-07-11T01:34:10.30171Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Making the confusion matrix\n\ncm = confusion_matrix(Y_test, Y_pred)\nprint(cm)\n\n### Printing the accuracy of the model\n\nnn7_accuracy = round(100 * accuracy_score(Y_test, Y_pred), 2)\nmodel_performance['7 Nearest Neighbors'] = nn7_accuracy\nprint('The accuracy of this model is {} %.'.format(nn7_accuracy))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.305271Z","iopub.execute_input":"2022-07-11T01:34:10.306228Z","iopub.status.idle":"2022-07-11T01:34:10.320419Z","shell.execute_reply.started":"2022-07-11T01:34:10.306175Z","shell.execute_reply":"2022-07-11T01:34:10.3191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 5.2.7 Applying K-Nearest Neighbors (9-NN)","metadata":{}},{"cell_type":"code","source":"### Applying 9NN model\n\nclassifier_9nn = KNeighborsClassifier(n_neighbors = 9, algorithm = 'auto', p = 2, metric = 'minkowski')\nclassifier_9nn.fit(X_train, Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.322773Z","iopub.execute_input":"2022-07-11T01:34:10.323737Z","iopub.status.idle":"2022-07-11T01:34:10.341138Z","shell.execute_reply.started":"2022-07-11T01:34:10.323681Z","shell.execute_reply":"2022-07-11T01:34:10.339547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Predicting the Test set results\n\nY_pred = classifier_9nn.predict(X_test)\nprint(np.concatenate((Y_pred.reshape(len(Y_pred), 1), Y_test.reshape(len(Y_test), 1)), 1))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.34358Z","iopub.execute_input":"2022-07-11T01:34:10.344628Z","iopub.status.idle":"2022-07-11T01:34:10.377122Z","shell.execute_reply.started":"2022-07-11T01:34:10.344571Z","shell.execute_reply":"2022-07-11T01:34:10.375868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Making the confusion matrix\n\ncm = confusion_matrix(Y_test, Y_pred)\nprint(cm)\n\n### Printing the accuracy of the model\n\nnn9_accuracy = round(100 * accuracy_score(Y_test, Y_pred), 2)\nmodel_performance['9 Nearest Neighbors'] = nn9_accuracy\nprint('The accuracy of this model is {} %.'.format(nn9_accuracy))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.379408Z","iopub.execute_input":"2022-07-11T01:34:10.380344Z","iopub.status.idle":"2022-07-11T01:34:10.396101Z","shell.execute_reply.started":"2022-07-11T01:34:10.380292Z","shell.execute_reply":"2022-07-11T01:34:10.39424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 5.2.8 Applying K-Nearest Neighbors (11-NN)","metadata":{}},{"cell_type":"code","source":"### Applying 11NN model\n\nclassifier_11nn = KNeighborsClassifier(n_neighbors = 11, algorithm = 'auto', p = 2, metric = 'minkowski')\nclassifier_11nn.fit(X_train, Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.398597Z","iopub.execute_input":"2022-07-11T01:34:10.399542Z","iopub.status.idle":"2022-07-11T01:34:10.411385Z","shell.execute_reply.started":"2022-07-11T01:34:10.399476Z","shell.execute_reply":"2022-07-11T01:34:10.409873Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Predicting the Test set results\n\nY_pred = classifier_11nn.predict(X_test)\nprint(np.concatenate((Y_pred.reshape(len(Y_pred), 1), Y_test.reshape(len(Y_test), 1)), 1))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.41364Z","iopub.execute_input":"2022-07-11T01:34:10.414553Z","iopub.status.idle":"2022-07-11T01:34:10.448009Z","shell.execute_reply.started":"2022-07-11T01:34:10.414503Z","shell.execute_reply":"2022-07-11T01:34:10.446668Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Making the confusion matrix\n\ncm = confusion_matrix(Y_test, Y_pred)\nprint(cm)\n\n### Printing the accuracy of the model\n\nnn11_accuracy = round(100 * accuracy_score(Y_test, Y_pred), 2)\nmodel_performance['11 Nearest Neighbors'] = nn11_accuracy\nprint('The accuracy of this model is {} %.'.format(nn11_accuracy))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.450926Z","iopub.execute_input":"2022-07-11T01:34:10.45293Z","iopub.status.idle":"2022-07-11T01:34:10.46918Z","shell.execute_reply.started":"2022-07-11T01:34:10.452846Z","shell.execute_reply":"2022-07-11T01:34:10.467223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 5.2.8 Applying K-Nearest Neighbors (13-NN)","metadata":{}},{"cell_type":"code","source":"### Applying 13NN model\n\nclassifier_13nn = KNeighborsClassifier(n_neighbors = 13, algorithm = 'auto', p = 2, metric = 'minkowski')\nclassifier_13nn.fit(X_train, Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.47131Z","iopub.execute_input":"2022-07-11T01:34:10.47218Z","iopub.status.idle":"2022-07-11T01:34:10.486907Z","shell.execute_reply.started":"2022-07-11T01:34:10.47213Z","shell.execute_reply":"2022-07-11T01:34:10.484799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Predicting the Test set results\n\nY_pred = classifier_13nn.predict(X_test)\nprint(np.concatenate((Y_pred.reshape(len(Y_pred), 1), Y_test.reshape(len(Y_test), 1)), 1))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.489193Z","iopub.execute_input":"2022-07-11T01:34:10.490159Z","iopub.status.idle":"2022-07-11T01:34:10.604829Z","shell.execute_reply.started":"2022-07-11T01:34:10.490106Z","shell.execute_reply":"2022-07-11T01:34:10.603551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Making the confusion matrix\n\ncm = confusion_matrix(Y_test, Y_pred)\nprint(cm)\n\n### Printing the accuracy of the model\n\nnn13_accuracy = round(100 * accuracy_score(Y_test, Y_pred), 2)\nmodel_performance['13 Nearest Neighbors'] = nn13_accuracy\nprint('The accuracy of this model is {} %.'.format(nn13_accuracy))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.60609Z","iopub.execute_input":"2022-07-11T01:34:10.606524Z","iopub.status.idle":"2022-07-11T01:34:10.625035Z","shell.execute_reply.started":"2022-07-11T01:34:10.60647Z","shell.execute_reply":"2022-07-11T01:34:10.623881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 5.2.9 Applying K-Nearest Neighbors (15-NN)","metadata":{}},{"cell_type":"code","source":"### Applying 15NN model\n\nclassifier_15nn = KNeighborsClassifier(n_neighbors = 15, algorithm = 'auto', p = 2, metric = 'minkowski')\nclassifier_15nn.fit(X_train, Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.626296Z","iopub.execute_input":"2022-07-11T01:34:10.627003Z","iopub.status.idle":"2022-07-11T01:34:10.649134Z","shell.execute_reply.started":"2022-07-11T01:34:10.626958Z","shell.execute_reply":"2022-07-11T01:34:10.647895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Predicting the Test set results\n\nY_pred = classifier_15nn.predict(X_test)\nprint(np.concatenate((Y_pred.reshape(len(Y_pred), 1), Y_test.reshape(len(Y_test), 1)), 1))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.651006Z","iopub.execute_input":"2022-07-11T01:34:10.651886Z","iopub.status.idle":"2022-07-11T01:34:10.684163Z","shell.execute_reply.started":"2022-07-11T01:34:10.651834Z","shell.execute_reply":"2022-07-11T01:34:10.682865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Making the confusion matrix\n\ncm = confusion_matrix(Y_test, Y_pred)\nprint(cm)\n\n### Printing the accuracy of the model\n\nnn15_accuracy = round(100 * accuracy_score(Y_test, Y_pred), 2)\nmodel_performance['15 Nearest Neighbors'] = nn15_accuracy\nprint('The accuracy of this model is {} %.'.format(nn15_accuracy))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.685904Z","iopub.execute_input":"2022-07-11T01:34:10.693081Z","iopub.status.idle":"2022-07-11T01:34:10.71688Z","shell.execute_reply.started":"2022-07-11T01:34:10.693012Z","shell.execute_reply":"2022-07-11T01:34:10.715302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 5.2.10 Applying K-Nearest Neighbors (17-NN)","metadata":{}},{"cell_type":"code","source":"### Applying 17NN model\n\nclassifier_17nn = KNeighborsClassifier(n_neighbors = 17, algorithm = 'auto', p = 2, metric = 'minkowski')\nclassifier_17nn.fit(X_train, Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.718927Z","iopub.execute_input":"2022-07-11T01:34:10.72463Z","iopub.status.idle":"2022-07-11T01:34:10.752683Z","shell.execute_reply.started":"2022-07-11T01:34:10.724569Z","shell.execute_reply":"2022-07-11T01:34:10.75101Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Predicting the Test set results\n\nY_pred = classifier_17nn.predict(X_test)\nprint(np.concatenate((Y_pred.reshape(len(Y_pred), 1), Y_test.reshape(len(Y_test), 1)), 1))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.754779Z","iopub.execute_input":"2022-07-11T01:34:10.761391Z","iopub.status.idle":"2022-07-11T01:34:10.825183Z","shell.execute_reply.started":"2022-07-11T01:34:10.761319Z","shell.execute_reply":"2022-07-11T01:34:10.824143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Making the confusion matrix\n\ncm = confusion_matrix(Y_test, Y_pred)\nprint(cm)\n\n### Printing the accuracy of the model\n\nnn17_accuracy = round(100 * accuracy_score(Y_test, Y_pred), 2)\nmodel_performance['17 Nearest Neighbors'] = nn17_accuracy\nprint('The accuracy of this model is {} %.'.format(nn17_accuracy))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.826818Z","iopub.execute_input":"2022-07-11T01:34:10.828916Z","iopub.status.idle":"2022-07-11T01:34:10.859725Z","shell.execute_reply.started":"2022-07-11T01:34:10.828867Z","shell.execute_reply":"2022-07-11T01:34:10.858224Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 5.2.11 Applying K-Nearest Neighbors (19-NN)","metadata":{}},{"cell_type":"code","source":"### Applying 19NN model\n\nclassifier_19nn = KNeighborsClassifier(n_neighbors = 19, algorithm = 'auto', p = 2, metric = 'minkowski')\nclassifier_19nn.fit(X_train, Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.864263Z","iopub.execute_input":"2022-07-11T01:34:10.869418Z","iopub.status.idle":"2022-07-11T01:34:10.887805Z","shell.execute_reply.started":"2022-07-11T01:34:10.869339Z","shell.execute_reply":"2022-07-11T01:34:10.886381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Predicting the Test set results\n\nY_pred = classifier_19nn.predict(X_test)\nprint(np.concatenate((Y_pred.reshape(len(Y_pred), 1), Y_test.reshape(len(Y_test), 1)), 1))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.892155Z","iopub.execute_input":"2022-07-11T01:34:10.89688Z","iopub.status.idle":"2022-07-11T01:34:10.924526Z","shell.execute_reply.started":"2022-07-11T01:34:10.896823Z","shell.execute_reply":"2022-07-11T01:34:10.923321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Making the confusion matrix\n\ncm = confusion_matrix(Y_test, Y_pred)\nprint(cm)\n\n### Printing the accuracy of the model\n\nnn19_accuracy = round(100 * accuracy_score(Y_test, Y_pred), 2)\nmodel_performance['19 Nearest Neighbors'] = nn19_accuracy\nprint('The accuracy of this model is {} %.'.format(nn19_accuracy))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.926641Z","iopub.execute_input":"2022-07-11T01:34:10.930476Z","iopub.status.idle":"2022-07-11T01:34:10.943278Z","shell.execute_reply.started":"2022-07-11T01:34:10.930421Z","shell.execute_reply":"2022-07-11T01:34:10.941413Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's plot the accuracies of all the nearest neighbor models. We can see that the accuracy first increases, reaches a peak and then decreases.","metadata":{}},{"cell_type":"code","source":"### Looking at the accuracy graph of all the nearest neighbors\n\nlabels = ['1NN', '3NN', '5NN', '7NN', '9NN', '11NN', '13NN', '15NN', '17NN', '19NN']\nvalues = [nn1_accuracy, nn3_accuracy, nn5_accuracy, nn7_accuracy, nn9_accuracy, \n          nn11_accuracy, nn13_accuracy, nn15_accuracy, nn17_accuracy, nn19_accuracy]\n\nplt.title('Accuracies of all the nearest neighbor models')\nplt.xlabel('Model')\nplt.ylabel('Accuracy')\nplt.plot(labels, values, '*-')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:10.945631Z","iopub.execute_input":"2022-07-11T01:34:10.94638Z","iopub.status.idle":"2022-07-11T01:34:11.21339Z","shell.execute_reply.started":"2022-07-11T01:34:10.946331Z","shell.execute_reply":"2022-07-11T01:34:11.212277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 5.2.12 Applying Gaussian Naive Bayes Classification ","metadata":{}},{"cell_type":"code","source":"### Applying Naive Bayes Classification model\n\nnaive_bayes_classifier = GaussianNB()\nnaive_bayes_classifier.fit(X_train, Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:11.214666Z","iopub.execute_input":"2022-07-11T01:34:11.215076Z","iopub.status.idle":"2022-07-11T01:34:11.22304Z","shell.execute_reply.started":"2022-07-11T01:34:11.215045Z","shell.execute_reply":"2022-07-11T01:34:11.222283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Predicting the Test set results\n\nY_pred = naive_bayes_classifier.predict(X_test)\nprint(np.concatenate((Y_pred.reshape(len(Y_pred), 1), Y_test.reshape(len(Y_test), 1)), 1))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:11.224196Z","iopub.execute_input":"2022-07-11T01:34:11.225168Z","iopub.status.idle":"2022-07-11T01:34:11.233561Z","shell.execute_reply.started":"2022-07-11T01:34:11.225134Z","shell.execute_reply":"2022-07-11T01:34:11.232752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Making the confusion matrix\n\ncm = confusion_matrix(Y_test, Y_pred)\nprint(cm)\n\n### Printing the accuracy of the model\n\nnaive_bayes_accuracy = round(100 * accuracy_score(Y_test, Y_pred), 2)\nmodel_performance['Gaussian Naive Bayes'] = naive_bayes_accuracy\nprint('The accuracy of this model is {} %.'.format(naive_bayes_accuracy))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:11.234972Z","iopub.execute_input":"2022-07-11T01:34:11.235972Z","iopub.status.idle":"2022-07-11T01:34:11.24624Z","shell.execute_reply.started":"2022-07-11T01:34:11.235938Z","shell.execute_reply":"2022-07-11T01:34:11.245398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 5.2.13 Applying Decision Tree Classification","metadata":{}},{"cell_type":"code","source":"### Applying Decision Tree Classification model\n\ndecision_tree_classifier = DecisionTreeClassifier(criterion = 'entropy', random_state = 27)\ndecision_tree_classifier.fit(X_train, Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:11.247797Z","iopub.execute_input":"2022-07-11T01:34:11.249035Z","iopub.status.idle":"2022-07-11T01:34:11.261392Z","shell.execute_reply.started":"2022-07-11T01:34:11.248987Z","shell.execute_reply":"2022-07-11T01:34:11.260266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Predicting the Test set results\n\nY_pred = decision_tree_classifier.predict(X_test)\nprint(np.concatenate((Y_pred.reshape(len(Y_pred), 1), Y_test.reshape(len(Y_test), 1)), 1))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:11.265021Z","iopub.execute_input":"2022-07-11T01:34:11.265719Z","iopub.status.idle":"2022-07-11T01:34:11.277302Z","shell.execute_reply.started":"2022-07-11T01:34:11.265669Z","shell.execute_reply":"2022-07-11T01:34:11.275969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Making the confusion matrix\n\ncm = confusion_matrix(Y_test, Y_pred)\nprint(cm)\n\n### Printing the accuracy of the model\n\ndecision_tree_accuracy = round(100 * accuracy_score(Y_test, Y_pred), 2)\nmodel_performance['Decision Tree'] = decision_tree_accuracy\nprint('The accuracy of this model is {} %.'.format(decision_tree_accuracy))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:11.279209Z","iopub.execute_input":"2022-07-11T01:34:11.279673Z","iopub.status.idle":"2022-07-11T01:34:11.290983Z","shell.execute_reply.started":"2022-07-11T01:34:11.279632Z","shell.execute_reply":"2022-07-11T01:34:11.289745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 5.2.14 Applying Random Forest Classification (10 trees)","metadata":{}},{"cell_type":"code","source":"### Applying Random Forest Classification model (10 trees)\n\nrandom_forest_10_classifier = RandomForestClassifier(n_estimators = 10, criterion = 'entropy', random_state = 27)\nrandom_forest_10_classifier.fit(X_train, Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:11.294014Z","iopub.execute_input":"2022-07-11T01:34:11.294875Z","iopub.status.idle":"2022-07-11T01:34:11.329803Z","shell.execute_reply.started":"2022-07-11T01:34:11.29484Z","shell.execute_reply":"2022-07-11T01:34:11.328543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Predicting the Test set results\n\nY_pred = random_forest_10_classifier.predict(X_test)\nprint(np.concatenate((Y_pred.reshape(len(Y_pred), 1), Y_test.reshape(len(Y_test), 1)), 1))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:11.331559Z","iopub.execute_input":"2022-07-11T01:34:11.332278Z","iopub.status.idle":"2022-07-11T01:34:11.346246Z","shell.execute_reply.started":"2022-07-11T01:34:11.332223Z","shell.execute_reply":"2022-07-11T01:34:11.344957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Making the confusion matrix\n\ncm = confusion_matrix(Y_test, Y_pred)\nprint(cm)\n\n### Printing the accuracy of the model\n\nrandom_forest_10_accuracy = round(100 * accuracy_score(Y_test, Y_pred), 2)\nmodel_performance['Random Forest (10 trees)'] = random_forest_10_accuracy\nprint('The accuracy of this model is {} %.'.format(random_forest_10_accuracy))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:11.347898Z","iopub.execute_input":"2022-07-11T01:34:11.348653Z","iopub.status.idle":"2022-07-11T01:34:11.359561Z","shell.execute_reply.started":"2022-07-11T01:34:11.348609Z","shell.execute_reply":"2022-07-11T01:34:11.358051Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 5.2.15 Applying Random Forest Classification (25 trees)","metadata":{}},{"cell_type":"code","source":"### Applying Random Forest Classification model (25 trees)\n\nrandom_forest_25_classifier = RandomForestClassifier(n_estimators = 25, criterion = 'entropy', random_state = 27)\nrandom_forest_25_classifier.fit(X_train, Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:11.3616Z","iopub.execute_input":"2022-07-11T01:34:11.362424Z","iopub.status.idle":"2022-07-11T01:34:11.431509Z","shell.execute_reply.started":"2022-07-11T01:34:11.362381Z","shell.execute_reply":"2022-07-11T01:34:11.430571Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Predicting the Test set results\n\nY_pred = random_forest_25_classifier.predict(X_test)\nprint(np.concatenate((Y_pred.reshape(len(Y_pred), 1), Y_test.reshape(len(Y_test), 1)), 1))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:11.435339Z","iopub.execute_input":"2022-07-11T01:34:11.435745Z","iopub.status.idle":"2022-07-11T01:34:11.455305Z","shell.execute_reply.started":"2022-07-11T01:34:11.435704Z","shell.execute_reply":"2022-07-11T01:34:11.454018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Making the confusion matrix\n\ncm = confusion_matrix(Y_test, Y_pred)\nprint(cm)\n\n### Printing the accuracy of the model\n\nrandom_forest_25_accuracy = round(100 * accuracy_score(Y_test, Y_pred), 2)\nmodel_performance['Random Forest (25 trees)'] = random_forest_25_accuracy\nprint('The accuracy of this model is {} %.'.format(random_forest_25_accuracy))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:11.457324Z","iopub.execute_input":"2022-07-11T01:34:11.458132Z","iopub.status.idle":"2022-07-11T01:34:11.467417Z","shell.execute_reply.started":"2022-07-11T01:34:11.458086Z","shell.execute_reply":"2022-07-11T01:34:11.466338Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 5.2.16 Applying Random Forest Classification (50 trees)","metadata":{}},{"cell_type":"code","source":"### Applying Random Forest Classification model (50 trees)\n\nrandom_forest_50_classifier = RandomForestClassifier(n_estimators = 50, criterion = 'entropy', random_state = 27)\nrandom_forest_50_classifier.fit(X_train, Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:11.469075Z","iopub.execute_input":"2022-07-11T01:34:11.469763Z","iopub.status.idle":"2022-07-11T01:34:11.598934Z","shell.execute_reply.started":"2022-07-11T01:34:11.469721Z","shell.execute_reply":"2022-07-11T01:34:11.597651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Predicting the Test set results\n\nY_pred = random_forest_50_classifier.predict(X_test)\nprint(np.concatenate((Y_pred.reshape(len(Y_pred), 1), Y_test.reshape(len(Y_test), 1)), 1))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:11.60252Z","iopub.execute_input":"2022-07-11T01:34:11.6032Z","iopub.status.idle":"2022-07-11T01:34:11.622192Z","shell.execute_reply.started":"2022-07-11T01:34:11.603154Z","shell.execute_reply":"2022-07-11T01:34:11.62088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Making the confusion matrix\n\ncm = confusion_matrix(Y_test, Y_pred)\nprint(cm)\n\n### Printing the accuracy of the model\n\nrandom_forest_50_accuracy = round(100 * accuracy_score(Y_test, Y_pred), 2)\nmodel_performance['Random Forest (50 trees)'] = random_forest_50_accuracy\nprint('The accuracy of this model is {} %.'.format(random_forest_50_accuracy))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:11.623737Z","iopub.execute_input":"2022-07-11T01:34:11.624844Z","iopub.status.idle":"2022-07-11T01:34:11.634238Z","shell.execute_reply.started":"2022-07-11T01:34:11.624797Z","shell.execute_reply":"2022-07-11T01:34:11.633434Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### 5.2.17 Applying Random Forest Classification (100 trees)","metadata":{}},{"cell_type":"code","source":"### Applying Random Forest Classification model (100 trees)\n\nrandom_forest_100_classifier = RandomForestClassifier(n_estimators = 100, criterion = 'entropy', random_state = 27)\nrandom_forest_100_classifier.fit(X_train, Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:11.635635Z","iopub.execute_input":"2022-07-11T01:34:11.635959Z","iopub.status.idle":"2022-07-11T01:34:11.864071Z","shell.execute_reply.started":"2022-07-11T01:34:11.635926Z","shell.execute_reply":"2022-07-11T01:34:11.862761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Predicting the Test set results\n\nY_pred = random_forest_100_classifier.predict(X_test)\nprint(np.concatenate((Y_pred.reshape(len(Y_pred), 1), Y_test.reshape(len(Y_test), 1)), 1))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:11.86591Z","iopub.execute_input":"2022-07-11T01:34:11.866624Z","iopub.status.idle":"2022-07-11T01:34:11.895596Z","shell.execute_reply.started":"2022-07-11T01:34:11.86657Z","shell.execute_reply":"2022-07-11T01:34:11.894245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Making the confusion matrix\n\ncm = confusion_matrix(Y_test, Y_pred)\nprint(cm)\n\n### Printing the accuracy of the model\n\nrandom_forest_100_accuracy = round(100 * accuracy_score(Y_test, Y_pred), 2)\nmodel_performance['Random Forest (100 trees)'] = random_forest_100_accuracy\nprint('The accuracy of this model is {} %.'.format(random_forest_100_accuracy))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:11.897345Z","iopub.execute_input":"2022-07-11T01:34:11.898098Z","iopub.status.idle":"2022-07-11T01:34:11.907258Z","shell.execute_reply.started":"2022-07-11T01:34:11.89805Z","shell.execute_reply":"2022-07-11T01:34:11.906446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 5.3 Model evaluation","metadata":{}},{"cell_type":"markdown","source":"Model evaluation is the process of using different evaluation metrics to understand a machine learning model's performance, as well as its strengths and weaknesses.","metadata":{}},{"cell_type":"markdown","source":"##### 5.3.1 Training accuracy of the models","metadata":{}},{"cell_type":"markdown","source":"Now we will tabulate all the models along with their accuracies. This data is stored in the model_performance dictionary. We will use the tabulate package for tabulating the results.","metadata":{}},{"cell_type":"code","source":"### Looking at the model performance dictionary\n\nmodel_performance","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:11.908602Z","iopub.execute_input":"2022-07-11T01:34:11.908927Z","iopub.status.idle":"2022-07-11T01:34:11.922844Z","shell.execute_reply.started":"2022-07-11T01:34:11.908897Z","shell.execute_reply":"2022-07-11T01:34:11.921608Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Tabulating the results\n\ntable = []\ntable.append(['S.No.', 'Classification Model', 'Model Accuracy'])\ncount = 1\n\nfor model in model_performance:\n    row = [count, model, model_performance[model]]\n    table.append(row)\n    count += 1\n    \nprint(tabulate(table, headers = 'firstrow', tablefmt = 'fancy_grid'))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:11.924702Z","iopub.execute_input":"2022-07-11T01:34:11.92564Z","iopub.status.idle":"2022-07-11T01:34:11.937506Z","shell.execute_reply.started":"2022-07-11T01:34:11.925593Z","shell.execute_reply":"2022-07-11T01:34:11.93636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above table, we can see that the models Kernel Support Vector Classification and 13 Nearest Neighbors have the same accuracy. ","metadata":{}},{"cell_type":"markdown","source":"##### 5.3.2 Applying K-fold Cross Validation","metadata":{}},{"cell_type":"markdown","source":"It is important to not get too carried away with models with impressive training accuracy as what we should focus on instead is the model's ability to predict out-of-samples data, in other words, data our model has not seen before. This is where k-fold cross validation comes in. K-fold cross validation is a technique whereby a subset of our training set is kept aside and will act as holdout set for testing purposes. ","metadata":{}},{"cell_type":"code","source":"### Create a list of classifiers\n\nclassifiers = []\nclassifiers.append(logistic_classifier)\nclassifiers.append(linear_svm_classifier)\nclassifiers.append(kernel_svm_classifier)\nclassifiers.append(classifier_1nn)\nclassifiers.append(classifier_3nn)\nclassifiers.append(classifier_5nn)\nclassifiers.append(classifier_7nn)\nclassifiers.append(classifier_9nn)\nclassifiers.append(classifier_11nn)\nclassifiers.append(classifier_13nn)\nclassifiers.append(classifier_15nn)\nclassifiers.append(classifier_17nn)\nclassifiers.append(classifier_19nn)\nclassifiers.append(naive_bayes_classifier)\nclassifiers.append(decision_tree_classifier)\nclassifiers.append(random_forest_10_classifier)\nclassifiers.append(random_forest_25_classifier)\nclassifiers.append(random_forest_50_classifier)\nclassifiers.append(random_forest_100_classifier)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:11.938794Z","iopub.execute_input":"2022-07-11T01:34:11.939708Z","iopub.status.idle":"2022-07-11T01:34:11.949059Z","shell.execute_reply.started":"2022-07-11T01:34:11.939674Z","shell.execute_reply":"2022-07-11T01:34:11.947632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will now use this list of classifiers to perform K-fold cross validation.","metadata":{}},{"cell_type":"code","source":"### Applying K-fold cross validation and tabulating the results\n\nvalidation_accuracies = []\nstandard_deviations = []\n\nfor each_classifier in classifiers:\n    accuracy = cross_val_score(estimator = each_classifier, X = X_train, y = Y_train, cv = 20)\n    validation_accuracies.append(np.mean(accuracy) * 100)\n    standard_deviations.append(accuracy.std() * 100)\n    \nvalidation_accuracies","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:11.950569Z","iopub.execute_input":"2022-07-11T01:34:11.950911Z","iopub.status.idle":"2022-07-11T01:34:22.822795Z","shell.execute_reply.started":"2022-07-11T01:34:11.950882Z","shell.execute_reply":"2022-07-11T01:34:22.82158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Tabulating the results\n\ntable = []\ntable.append(['S.No.', 'Classification Model', 'Model Accuracy', 'Validation Accuracy', 'Standard Deviation'])\ncount = 1\n\nfor model in model_performance:\n    row = [count, model, model_performance[model], validation_accuracies[count - 1], standard_deviations[count - 1]]\n    table.append(row)\n    count += 1\n    \nprint(tabulate(table, headers = 'firstrow', tablefmt = 'fancy_grid'))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:22.824278Z","iopub.execute_input":"2022-07-11T01:34:22.824991Z","iopub.status.idle":"2022-07-11T01:34:22.83663Z","shell.execute_reply.started":"2022-07-11T01:34:22.824953Z","shell.execute_reply":"2022-07-11T01:34:22.835284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above table, we can see that the validation accuracy is higher in Kernel Support Vector Classification. Hence, we will use the model - Kernel Support Vector Classification for our solution.","metadata":{}},{"cell_type":"markdown","source":"### 6. Preparing Data For Submission","metadata":{}},{"cell_type":"markdown","source":"In this section, we will predict the results using the Kernel Support Vector Classification model on the given test data. The predicted results will be stored in an excel sheet which will be later uploaded to the Kaggle's competition page to find out the accuracy.","metadata":{}},{"cell_type":"markdown","source":"#### 6.1 Predicting the results using Kernel SVM","metadata":{}},{"cell_type":"code","source":"### Predicting the Test set results\n\nfinal_predictions = kernel_svm_classifier.predict(real_test_data)\nprint(len(final_predictions))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:22.838333Z","iopub.execute_input":"2022-07-11T01:34:22.83872Z","iopub.status.idle":"2022-07-11T01:34:22.862127Z","shell.execute_reply.started":"2022-07-11T01:34:22.838684Z","shell.execute_reply":"2022-07-11T01:34:22.860735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 6.2 Creating the excel sheet for submission","metadata":{}},{"cell_type":"code","source":"### Creating a result dataframe using PassengerId and generated predictions\n\npassengerId = list(range(892, 1310))\nresult = pd.DataFrame(passengerId, columns = ['PassengerId'])\nresult['Survived'] = final_predictions\nresult","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:22.865594Z","iopub.execute_input":"2022-07-11T01:34:22.866069Z","iopub.status.idle":"2022-07-11T01:34:22.88143Z","shell.execute_reply.started":"2022-07-11T01:34:22.866034Z","shell.execute_reply":"2022-07-11T01:34:22.880314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Creating a CSV file of the predictions\n\ncompression_opts = dict(method = 'zip', archive_name = 'result_kernel_svm.csv')  \nresult.to_csv('result_kernel_svm.zip', index = False, compression = compression_opts)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:34:22.882847Z","iopub.execute_input":"2022-07-11T01:34:22.883183Z","iopub.status.idle":"2022-07-11T01:34:22.893854Z","shell.execute_reply.started":"2022-07-11T01:34:22.883154Z","shell.execute_reply":"2022-07-11T01:34:22.892759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 7. Conclusion","metadata":{}},{"cell_type":"markdown","source":"The above predictions resulted in an accuracy of 78.229 percent. ","metadata":{}}]}