{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <u> Survival Analysis On Titanic Dataset.<u>","metadata":{}},{"cell_type":"markdown","source":"*The sinking of the Titanic is one of the most infamous shipwrecks in history.*\n\n*This is a dataset where On April 15, 1912, during her maiden voyage, the widely considered “unsinkable” RMS Titanic sank after colliding with an iceberg. Unfortunately, there weren’t enough lifeboats for everyone onboard, resulting in the death of 1502 out of 2224 passengers and crew.*\n\n*While there was some element of luck involved in surviving, it seems some groups of people were more likely to survive than others.*\n\n***The Attributes of the dataset are as follows :***   \n- *Survived - Survival (0 = No, 1 = Yes)*\n- *pclass - A proxy for socio-economic status (1st = Upper, 2nd = Middle, 3rd = Lower)*\n- *Name - Names of the passangers on board*\n- *sex - Gender of the passangers on board (0 = Male, 1 = Female)*\n- *Age - Age of the passangers on board*\n- *sibsp - Number of sibblings/sprouse on board* \n    - *The dataset defines family relations in this way...*\n    - *Sibling = brother, sister, stepbrother, stepsister*\n    - *Spouse = husband, wife (mistresses and fiancés were ignored)*\n- *Parch - Number of parents/children on board*\n       -* The dataset defines family relations in this way...*\n       -* Parent = mother, father*\n       -* Child = daughter, son, stepdaughter, stepson*\n       -* Some children travelled only with a nanny, therefore parch=0 for them.*\n- *Ticket - Ticket Number* \n- *Fare - Passanger fare*\n- *Cabin - Cabin Number*\n- *Embarked - Port of Embarkation (C = Cherbourg, Q = Queenstown, S = Southampton)* \n ","metadata":{}},{"cell_type":"markdown","source":"## <u> Problem Statement: <u>\n#### ***To build a predictive model that answers the question: “what sorts of people were more likely to survive?” using passenger data (ie name, age, gender, socio-economic class, etc).***","metadata":{}},{"cell_type":"markdown","source":"## <u>Process involved in the process:<u>\n- ***i).Big Picture***\n- ***ii).Data Analysis and Exploration***\n- ***iii).Feature Engineering***\n- ***iv).Model Building***\n- ***v).Interpretation***","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport missingno as mp\nfrom matplotlib import colors\nimport statsmodels.api as sm\nimport pylab as py\nfrom sklearn.preprocessing import StandardScaler\nfrom imblearn.over_sampling import SMOTE\nfrom scipy.stats import probplot\nfrom sklearn.model_selection import train_test_split, cross_val_score\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.metrics import classification_report, confusion_matrix\nfrom collections import Counter\nfrom sklearn.neighbors import KNeighborsClassifier\n#import klib\nfrom sklearn.svm import SVC\nfrom sklearn.model_selection import GridSearchCV\n#from lazypredict.Supervised import LazyClassifier\nimport warnings\nfrom sklearn.model_selection import RandomizedSearchCV\nimport sklearn.metrics\nfrom sklearn.ensemble import AdaBoostClassifier\n%matplotlib inline\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:22:20.922578Z","iopub.execute_input":"2022-08-10T06:22:20.923196Z","iopub.status.idle":"2022-08-10T06:22:24.139137Z","shell.execute_reply.started":"2022-08-10T06:22:20.923060Z","shell.execute_reply":"2022-08-10T06:22:24.137595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Note: I have used 2 packages namely klib and Lazypredict. I suggest you use it in your local notebook as I am unable to install these packages here.***","metadata":{}},{"cell_type":"code","source":"path = r\"../input/titanic/train.csv\"\ntrain_data = pd.read_csv(path)\ndata2 = train_data\n\n# Reading the test data\npath2 = r\"../input/titanic/test.csv\"\ntest_data = pd.read_csv(path2)\ndata3 = test_data","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:22:49.333967Z","iopub.execute_input":"2022-08-10T06:22:49.334526Z","iopub.status.idle":"2022-08-10T06:22:49.371018Z","shell.execute_reply.started":"2022-08-10T06:22:49.334484Z","shell.execute_reply":"2022-08-10T06:22:49.369397Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:22:53.556796Z","iopub.execute_input":"2022-08-10T06:22:53.557439Z","iopub.status.idle":"2022-08-10T06:22:53.597909Z","shell.execute_reply.started":"2022-08-10T06:22:53.557385Z","shell.execute_reply":"2022-08-10T06:22:53.596754Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:22:56.076648Z","iopub.execute_input":"2022-08-10T06:22:56.077149Z","iopub.status.idle":"2022-08-10T06:22:56.105486Z","shell.execute_reply.started":"2022-08-10T06:22:56.077110Z","shell.execute_reply":"2022-08-10T06:22:56.104168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 1). Big Picture:","metadata":{}},{"cell_type":"markdown","source":"### *i).Determining the Structure and Summary of the dataset:*","metadata":{}},{"cell_type":"code","source":"train_data.describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:22:59.304904Z","iopub.execute_input":"2022-08-10T06:22:59.306573Z","iopub.status.idle":"2022-08-10T06:22:59.376342Z","shell.execute_reply.started":"2022-08-10T06:22:59.306500Z","shell.execute_reply":"2022-08-10T06:22:59.375064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:23:02.087696Z","iopub.execute_input":"2022-08-10T06:23:02.088164Z","iopub.status.idle":"2022-08-10T06:23:02.125756Z","shell.execute_reply.started":"2022-08-10T06:23:02.088127Z","shell.execute_reply":"2022-08-10T06:23:02.124116Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*From the above table, we obtain a very clear picture on the summary of the dataset, i.e The measures of central tendency and dispersion with respect to each variable.*","metadata":{}},{"cell_type":"markdown","source":"### ii). *Determining the unique values:*","metadata":{}},{"cell_type":"code","source":"train_data.nunique()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:23:05.198841Z","iopub.execute_input":"2022-08-10T06:23:05.199918Z","iopub.status.idle":"2022-08-10T06:23:05.215387Z","shell.execute_reply.started":"2022-08-10T06:23:05.199862Z","shell.execute_reply":"2022-08-10T06:23:05.213864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.nunique()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:23:12.923341Z","iopub.execute_input":"2022-08-10T06:23:12.924790Z","iopub.status.idle":"2022-08-10T06:23:12.937016Z","shell.execute_reply.started":"2022-08-10T06:23:12.924735Z","shell.execute_reply":"2022-08-10T06:23:12.935715Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*From the above line of code, we are able to understand that the variables 'PassengerId','Name','Ticket' have too many unique values which is not useful to the subject under study. Therefore we drop these variables from both train and test data.*","metadata":{}},{"cell_type":"code","source":"train_data = train_data.drop(['PassengerId','Ticket','Name'], axis = 1)\ntest_data = test_data.drop(['PassengerId','Ticket','Name'], axis = 1)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:23:17.089573Z","iopub.execute_input":"2022-08-10T06:23:17.090056Z","iopub.status.idle":"2022-08-10T06:23:17.099525Z","shell.execute_reply.started":"2022-08-10T06:23:17.090019Z","shell.execute_reply":"2022-08-10T06:23:17.097830Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### iii). *Determining the number of Categorical variables:*","metadata":{}},{"cell_type":"code","source":"# categorical varibales of train data\ncat_count = 0\nnum_count = 0\nfor i in train_data.dtypes:\n    if i == 'object':\n        cat_count = cat_count + 1\n    else:\n        num_count = num_count + 1\n\nprint(\"The Number of Numerical variables in Train Data :\",num_count)\nprint(\"The Number of qualitative variables in Train Data:\",cat_count)\n\n#categorical variables of test data\ncat_count = 0\nnum_count = 0\nfor i in test_data.dtypes:\n    if i == 'object':\n        cat_count = cat_count + 1\n    else:\n        num_count = num_count + 1\n\nprint(\"The Number of Numerical variables in Test Data:\",num_count)\nprint(\"The Number of qualitative variables in Test Data:\",cat_count)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:23:21.266300Z","iopub.execute_input":"2022-08-10T06:23:21.266771Z","iopub.status.idle":"2022-08-10T06:23:21.278688Z","shell.execute_reply.started":"2022-08-10T06:23:21.266733Z","shell.execute_reply":"2022-08-10T06:23:21.277621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### iv). *Determining the Structure of Dataset:*","metadata":{}},{"cell_type":"code","source":"train_data.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:23:28.401143Z","iopub.execute_input":"2022-08-10T06:23:28.401676Z","iopub.status.idle":"2022-08-10T06:23:28.428486Z","shell.execute_reply.started":"2022-08-10T06:23:28.401607Z","shell.execute_reply":"2022-08-10T06:23:28.425450Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:23:32.039514Z","iopub.execute_input":"2022-08-10T06:23:32.040128Z","iopub.status.idle":"2022-08-10T06:23:32.056279Z","shell.execute_reply.started":"2022-08-10T06:23:32.040075Z","shell.execute_reply":"2022-08-10T06:23:32.055057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2). Data Analysis and Exploration:","metadata":{}},{"cell_type":"markdown","source":"### *i). Determining the missing values :* ","metadata":{}},{"cell_type":"code","source":"count_missing_data = train_data.isnull().sum()\npercent_missing_data = round(train_data.isnull().sum()/len(train_data) * 100, 1)\nmissing_data = pd.concat([count_missing_data, percent_missing_data], axis = 1)\nmissing_data.columns = [\"Missing (count)\", \"Missing (%)\"]\nmissing_data","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:23:36.831569Z","iopub.execute_input":"2022-08-10T06:23:36.833163Z","iopub.status.idle":"2022-08-10T06:23:36.854751Z","shell.execute_reply.started":"2022-08-10T06:23:36.833104Z","shell.execute_reply":"2022-08-10T06:23:36.853436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- ***Train Data***: *There are missing values with respect to the varibles 'AGE','CABIN' and 'EMBARKED'.*","metadata":{}},{"cell_type":"code","source":"count_missing_data = test_data.isnull().sum()\npercent_missing_data = round(test_data.isnull().sum()/len(test_data) * 100, 1)\nmissing_data = pd.concat([count_missing_data, percent_missing_data], axis = 1)\nmissing_data.columns = [\"Missing (count)\", \"Missing (%)\"]\nmissing_data","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:23:40.381563Z","iopub.execute_input":"2022-08-10T06:23:40.382050Z","iopub.status.idle":"2022-08-10T06:23:40.402937Z","shell.execute_reply.started":"2022-08-10T06:23:40.381999Z","shell.execute_reply":"2022-08-10T06:23:40.401604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- ***Test Data***: *There are missing values with respect to teh variables 'AGE', 'CABIN' and 'FARE'.*","metadata":{}},{"cell_type":"markdown","source":"***Since the variable 'CABIN' has more than 25% missing values , this variable will be remove for the dataset.***","metadata":{}},{"cell_type":"code","source":"train_data = train_data.drop(['Cabin'], axis = 1)\ntest_data = test_data.drop(['Cabin'], axis = 1)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:23:43.162789Z","iopub.execute_input":"2022-08-10T06:23:43.163334Z","iopub.status.idle":"2022-08-10T06:23:43.173076Z","shell.execute_reply.started":"2022-08-10T06:23:43.163295Z","shell.execute_reply":"2022-08-10T06:23:43.171761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### *ii).Detecting the Presence of Outliers :* ","metadata":{}},{"cell_type":"markdown","source":"*An Outlier is an observation in a given dataset that lies far from the rest of the observations.Treating the outliers is very crucial as they can negatively affect the statistical analysis and the training process of a machine learning algorithm resulting in lower accuracy.*","metadata":{}},{"cell_type":"code","source":"warnings.filterwarnings(\"ignore\")\n\nplt.figure(figsize = (18,7))\n\n#Age analysis for both datasets:\n# Boxplot\nplt.subplot(2,4,1)\nage_train = sns.boxplot(x = train_data['Age'], color = 'blue', width = 0.4)\nplt.title(\"Age of Train Data\")\nplt.subplot(2,4,2)\nage_test = sns.boxplot(x = test_data['Age'], color = 'blue', width = 0.4)\nplt.title(\"Age of Test Data\")\n\n#QQ plot:\nplt.subplot(2,4,3)\nprobplot(train_data['Age'], dist = 'norm', plot=plt)\nplt.title(\"Age of Train Data\")\nplt.subplot(2,4,4)\nprobplot(test_data['Age'], dist = 'norm', plot=plt)\nplt.title(\"Age of Test Data\")","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:23:45.687925Z","iopub.execute_input":"2022-08-10T06:23:45.689518Z","iopub.status.idle":"2022-08-10T06:23:46.283758Z","shell.execute_reply.started":"2022-08-10T06:23:45.689439Z","shell.execute_reply":"2022-08-10T06:23:46.282173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"warnings.filterwarnings(\"ignore\")\n\nplt.figure(figsize = (18,7))\n\n#Fare analysis for both datasets:\n# Boxplot\nplt.subplot(2,4,1)\nage_train = sns.boxplot(x = train_data['Fare'], color = 'blue', width = 0.4)\nplt.title(\"Fare of Test Data\")\nplt.subplot(2,4,2)\nage_test = sns.boxplot(x = test_data['Fare'], color = 'blue', width = 0.4)\nplt.title(\"Fare of Test Data\")\n\n#QQ plot:\nplt.subplot(2,4,3)\nprobplot(train_data['Age'], dist = 'norm', plot=plt)\nplt.title(\"Fare of Train Data\")\nplt.subplot(2,4,4)\nprobplot(test_data['Age'], dist = 'norm', plot=plt)\nplt.title(\"Fare of Test Data\")","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:23:50.046172Z","iopub.execute_input":"2022-08-10T06:23:50.046646Z","iopub.status.idle":"2022-08-10T06:23:50.521387Z","shell.execute_reply.started":"2022-08-10T06:23:50.046594Z","shell.execute_reply":"2022-08-10T06:23:50.519903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***The above figures with respect to the variable 'FARE' and 'AGE' of both TRAIN and TEST data, helps us in detecting the presence of huge number of outliers which is intended to be treated later.***","metadata":{}},{"cell_type":"markdown","source":"### *iii). Determining the distribution of the variables :*","metadata":{}},{"cell_type":"markdown","source":"*Determining the distribution or skewness helps us to understand where the most information is lying and also helps in analyzing  the outliers in a given data. If the values of a specific independent variable (feature) are skewed, depending on the model, skewness may violate model assumptions or may reduce the interpretation of feature importance.*","metadata":{}},{"cell_type":"code","source":"train_variables = train_data[['Age','Fare']]\n#klib.dist_plot(train_variables)\n#plt.title(\"The distributions with respect to Train Data\")    # write big ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_variables = test_data[['Age','Fare']]\n#klib.dist_plot(test_variables)\n#plt.title(\"The distributions with respect to Test Data\")    # write big ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Clearly, the variable 'FARE' are heavily skewed to the right in both TEST and TRAIN data.***","metadata":{}},{"cell_type":"markdown","source":"### *iv). Imputation of the Missing Values:*","metadata":{}},{"cell_type":"markdown","source":"### *a). EMBARKED VARIABLE:* ","metadata":{}},{"cell_type":"markdown","source":"***As mentioned earlier, there are missing values with respect to 'Embarked' only in the train data which is to be imputed by the most frequent value(mode) of this variable.***","metadata":{}},{"cell_type":"code","source":"train_data['Embarked'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:24:09.879053Z","iopub.execute_input":"2022-08-10T06:24:09.879564Z","iopub.status.idle":"2022-08-10T06:24:09.890499Z","shell.execute_reply.started":"2022-08-10T06:24:09.879521Z","shell.execute_reply":"2022-08-10T06:24:09.889473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data['Embarked'] = np.where((train_data['Embarked'].isnull() == True),'S', train_data['Embarked'])\ntrain_data['Embarked'].isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:24:11.501111Z","iopub.execute_input":"2022-08-10T06:24:11.502273Z","iopub.status.idle":"2022-08-10T06:24:11.512500Z","shell.execute_reply.started":"2022-08-10T06:24:11.502218Z","shell.execute_reply":"2022-08-10T06:24:11.511199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### *b). AGE VARIABLE:*","metadata":{}},{"cell_type":"code","source":"bar_pclass = sns.countplot(x = 'Survived', hue = 'Pclass', data = train_data, palette = 'PuRd')","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:24:13.446187Z","iopub.execute_input":"2022-08-10T06:24:13.447030Z","iopub.status.idle":"2022-08-10T06:24:13.693021Z","shell.execute_reply.started":"2022-08-10T06:24:13.446965Z","shell.execute_reply":"2022-08-10T06:24:13.691187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***From the above figure, its is very evident that passengers belonging to the 1st class(richer class) have more chances of being survived.So therefore we can make an assumption that the first preferneces/priority was given to the passengers belonging to the 1st class.***","metadata":{}},{"cell_type":"code","source":"bar_gender = sns.countplot(x = 'Survived',hue = 'Sex', data = train_data, palette = 'RdPu')","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:24:16.596928Z","iopub.execute_input":"2022-08-10T06:24:16.597410Z","iopub.status.idle":"2022-08-10T06:24:17.010301Z","shell.execute_reply.started":"2022-08-10T06:24:16.597368Z","shell.execute_reply":"2022-08-10T06:24:17.009431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***The graph tells us that majority who died were male passengers than female passengers.***","metadata":{}},{"cell_type":"code","source":"bar_gender_pclass = sns.countplot(x = 'Pclass', hue = 'Sex', data = train_data, palette = 'PuRd')","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:24:21.496430Z","iopub.execute_input":"2022-08-10T06:24:21.496916Z","iopub.status.idle":"2022-08-10T06:24:21.727290Z","shell.execute_reply.started":"2022-08-10T06:24:21.496872Z","shell.execute_reply":"2022-08-10T06:24:21.725745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***From the above figure, it is very evident that strength of the female passengers are more in each of the Pclass.***","metadata":{}},{"cell_type":"code","source":"data2 = train_data.dropna()\n\ngroup_age = round(data2.groupby(['Pclass','Sex']).aggregate({'Age':'median'}),0)\nprint(group_age)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:24:23.813008Z","iopub.execute_input":"2022-08-10T06:24:23.815062Z","iopub.status.idle":"2022-08-10T06:24:23.833075Z","shell.execute_reply.started":"2022-08-10T06:24:23.814991Z","shell.execute_reply":"2022-08-10T06:24:23.831733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***We will impute the missing values in the variable 'Age' by mean of age with respect to 'Pclass' and 'Gender' for the following two reasons:***    \n- ***a). The count of the female passengers in each of the Pclass is more than that of the male.***  \n- ***b). The surival rate for the passengers travelling in 1st pclass is more than any other Pclasses.***","metadata":{}},{"cell_type":"code","source":"#Treating AGE of traindata:\ndata = train_data\n\ndata['Age'] = np.where((data['Age'].isnull() == True) & (data['Sex'] == 'female') & (data['Pclass'] == 1),35.0, data['Age'])\ndata['Age'] = np.where((data['Age'].isnull() == True) & (data['Sex'] == 'male') & (data['Pclass'] == 1),40.0, data['Age'])\ndata['Age'] = np.where((data['Age'].isnull() == True) & (data['Sex'] == 'female') & (data['Pclass'] == 2),29.0, data['Age'])\ndata['Age'] = np.where((data['Age'].isnull() == True) & (data['Sex'] == 'male') & (data['Pclass'] == 2),18.0, data['Age'])\ndata['Age'] = np.where((data['Age'].isnull() == True) & (data['Sex'] == 'female') & (data['Pclass'] == 3),24.0, data['Age'])\ndata['Age'] = np.where((data['Age'].isnull() == True) & (data['Sex'] == 'male') & (data['Pclass'] == 3),25.0, data['Age'])\n\ntrain_data = data","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:24:26.212040Z","iopub.execute_input":"2022-08-10T06:24:26.212655Z","iopub.status.idle":"2022-08-10T06:24:26.236763Z","shell.execute_reply.started":"2022-08-10T06:24:26.212581Z","shell.execute_reply":"2022-08-10T06:24:26.235738Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Treating AGE of test data :\ndata = test_data\ndata['Age'] = np.where((data['Age'].isnull() == True) & (data['Sex'] == 'female') & (data['Pclass'] == 1),35.0, data['Age'])\ndata['Age'] = np.where((data['Age'].isnull() == True) & (data['Sex'] == 'male') & (data['Pclass'] == 1),40.0, data['Age'])\ndata['Age'] = np.where((data['Age'].isnull() == True) & (data['Sex'] == 'female') & (data['Pclass'] == 2),29.0, data['Age'])\ndata['Age'] = np.where((data['Age'].isnull() == True) & (data['Sex'] == 'male') & (data['Pclass'] == 2),18.0, data['Age'])\ndata['Age'] = np.where((data['Age'].isnull() == True) & (data['Sex'] == 'female') & (data['Pclass'] == 3),24.0, data['Age'])\ndata['Age'] = np.where((data['Age'].isnull() == True) & (data['Sex'] == 'male') & (data['Pclass'] == 3),25.0, data['Age'])\n\ntest_data = data","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:24:27.936293Z","iopub.execute_input":"2022-08-10T06:24:27.936781Z","iopub.status.idle":"2022-08-10T06:24:27.964235Z","shell.execute_reply.started":"2022-08-10T06:24:27.936737Z","shell.execute_reply":"2022-08-10T06:24:27.961567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data['Age'].isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:24:32.799091Z","iopub.execute_input":"2022-08-10T06:24:32.800249Z","iopub.status.idle":"2022-08-10T06:24:32.812148Z","shell.execute_reply.started":"2022-08-10T06:24:32.800187Z","shell.execute_reply":"2022-08-10T06:24:32.810562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data['Age'].isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:24:30.580824Z","iopub.execute_input":"2022-08-10T06:24:30.581257Z","iopub.status.idle":"2022-08-10T06:24:30.591752Z","shell.execute_reply.started":"2022-08-10T06:24:30.581220Z","shell.execute_reply":"2022-08-10T06:24:30.590112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### *c). FARE VARIABLE:*","metadata":{}},{"cell_type":"markdown","source":"***As observed earlier, there is 1 missing values with respect to the 'FARE' variable of test data which will be imputed by the median of the same due to the presence of the outliers.***","metadata":{}},{"cell_type":"code","source":"median_fare = round(test_data['Fare'].median(),0)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:24:35.342258Z","iopub.execute_input":"2022-08-10T06:24:35.342952Z","iopub.status.idle":"2022-08-10T06:24:35.350999Z","shell.execute_reply.started":"2022-08-10T06:24:35.342887Z","shell.execute_reply":"2022-08-10T06:24:35.349582Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data['Fare'] = np.where((test_data['Fare'].isnull() == True), median_fare, test_data['Fare'])","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:24:37.236690Z","iopub.execute_input":"2022-08-10T06:24:37.237177Z","iopub.status.idle":"2022-08-10T06:24:37.245093Z","shell.execute_reply.started":"2022-08-10T06:24:37.237136Z","shell.execute_reply":"2022-08-10T06:24:37.243693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data['Fare'].isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:24:38.867725Z","iopub.execute_input":"2022-08-10T06:24:38.868174Z","iopub.status.idle":"2022-08-10T06:24:38.878868Z","shell.execute_reply.started":"2022-08-10T06:24:38.868138Z","shell.execute_reply":"2022-08-10T06:24:38.877176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:24:40.259173Z","iopub.execute_input":"2022-08-10T06:24:40.259877Z","iopub.status.idle":"2022-08-10T06:24:40.275469Z","shell.execute_reply.started":"2022-08-10T06:24:40.259813Z","shell.execute_reply":"2022-08-10T06:24:40.274281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:24:42.268369Z","iopub.execute_input":"2022-08-10T06:24:42.268874Z","iopub.status.idle":"2022-08-10T06:24:42.282039Z","shell.execute_reply.started":"2022-08-10T06:24:42.268828Z","shell.execute_reply":"2022-08-10T06:24:42.280736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Clearly, all the missing values have been created.***","metadata":{}},{"cell_type":"markdown","source":"## *v). Outlier Treatment :*","metadata":{}},{"cell_type":"markdown","source":"- ***From the above histograms and boxplots, it is very evident that 'FARE' variable is positively skewed due to which we will be treating the ouliers with IQR.***\n- ***On the other hand, the variable 'AGE' is approximately normally distributed because of which we will be using the z-score to detect the outliers of the same.***","metadata":{}},{"cell_type":"markdown","source":"### *a). FARE VARAIBLE:*","metadata":{}},{"cell_type":"code","source":"outliers = []\ndef detect_outliers_iqr(data):\n    \"\"\" Function to detect and handle outliers using IQR.\"\"\"\n    \n    data = sorted(data)\n    q1 = np.percentile(data, 25)\n    q3 = np.percentile(data, 75)\n    \n    IQR = q3-q1\n    \n    lwr_bound = q1-(1.5*IQR)\n    upr_bound = q3+(1.5*IQR)\n  \n    data = np.where(data < lwr_bound, lwr_bound, data)\n    data = np.where(data > upr_bound, upr_bound, data)\n    \n    return data ","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:24:45.874075Z","iopub.execute_input":"2022-08-10T06:24:45.874532Z","iopub.status.idle":"2022-08-10T06:24:45.884959Z","shell.execute_reply.started":"2022-08-10T06:24:45.874494Z","shell.execute_reply":"2022-08-10T06:24:45.883754Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#calling the function for train_data:\ntrain_data['Fare'] = detect_outliers_iqr(train_data['Fare'])\n\n#calling the function for testdata:\ntest_data['Fare'] = detect_outliers_iqr(test_data['Fare'])","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:24:48.087400Z","iopub.execute_input":"2022-08-10T06:24:48.087902Z","iopub.status.idle":"2022-08-10T06:24:48.101409Z","shell.execute_reply.started":"2022-08-10T06:24:48.087857Z","shell.execute_reply":"2022-08-10T06:24:48.099228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize = (10,5))\n\nplt.subplot(1,2,1)\nsns.boxplot(x = data3['Fare'], color = 'yellow')\nplt.title(\"TrainData(Fare) - Before treatment\")\n\nplt.subplot(1,2,2)\nsns.boxplot(x = train_data['Fare'], color = 'greenyellow')\nplt.title('TrainData(Fare) - After treatment')","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:24:49.723115Z","iopub.execute_input":"2022-08-10T06:24:49.723957Z","iopub.status.idle":"2022-08-10T06:24:50.021512Z","shell.execute_reply.started":"2022-08-10T06:24:49.723909Z","shell.execute_reply":"2022-08-10T06:24:50.020037Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize = (10,5))\nplt.subplot(1,2,1)\nsns.boxplot(x = data2['Fare'], color = 'yellow')\nplt.title(\"TestData(Fare) - Before treatment\")\n\n\nplt.subplot(1,2,2)\nsns.boxplot(x = test_data['Fare'], color = 'yellow')\nplt.title(\"TestData(Fare) - Before treatment\")","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:24:52.467888Z","iopub.execute_input":"2022-08-10T06:24:52.469115Z","iopub.status.idle":"2022-08-10T06:24:52.760185Z","shell.execute_reply.started":"2022-08-10T06:24:52.469061Z","shell.execute_reply":"2022-08-10T06:24:52.758753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### *b). AGE VARIABLE:*","metadata":{}},{"cell_type":"code","source":"p = []\ndef zscore_treatment(data):\n    '''Function for oulier Treatment using z-score.'''\n    \n    mean1 = np.mean(data)\n    std1 = np.std(data)\n    \n    fifth = np.percentile(sorted(data), 5)\n    ninty_fifth = np.percentile(sorted(data), 95)\n    \n    for i in data:\n        z_score = (i - mean1)/std1\n    \n        if (z_score < -3):\n            data = data.replace([i], fifth)\n        elif (z_score > 3):\n            data = data.replace([i],ninty_fifth )\n        else:\n            pass\n        \n    return data          ","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:24:55.203243Z","iopub.execute_input":"2022-08-10T06:24:55.204199Z","iopub.status.idle":"2022-08-10T06:24:55.213400Z","shell.execute_reply.started":"2022-08-10T06:24:55.204148Z","shell.execute_reply":"2022-08-10T06:24:55.211780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# calling the function for train data\ntrain_data['Age'] = zscore_treatment(train_data['Age'])\n\n#calling the function for test data:\ntest_data['Age'] = zscore_treatment(test_data['Age'])","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:25:02.236414Z","iopub.execute_input":"2022-08-10T06:25:02.237510Z","iopub.status.idle":"2022-08-10T06:25:02.252095Z","shell.execute_reply.started":"2022-08-10T06:25:02.237432Z","shell.execute_reply":"2022-08-10T06:25:02.250676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize = (10,5))\n\nplt.subplot(1,2,1)\nsns.boxplot(x = data3['Age'], color = 'yellow')\nplt.title(\"TrainData(Age) - Before treatment\")\n\nplt.subplot(1,2,2)\nsns.boxplot(x = train_data['Age'], color = 'greenyellow')\nplt.title('TrainData(Age) - After treatment')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize = (10,5))\n\nplt.subplot(1,2,1)\nsns.boxplot(x = data2['Age'], color = 'yellow')\nplt.title(\"TestData(Age) - Before treatment\")\n\nplt.subplot(1,2,2)\nsns.boxplot(x = test_data['Age'], color = 'greenyellow')\nplt.title('TestData(Age) - After treatment')","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:25:03.869067Z","iopub.execute_input":"2022-08-10T06:25:03.869535Z","iopub.status.idle":"2022-08-10T06:25:04.170260Z","shell.execute_reply.started":"2022-08-10T06:25:03.869497Z","shell.execute_reply":"2022-08-10T06:25:04.168736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3). Feature Engineering:","metadata":{}},{"cell_type":"markdown","source":"### *i). Creation of New Features:*","metadata":{}},{"cell_type":"code","source":"train_data['Familymembers_onboard'] = train_data['SibSp'] + train_data['Parch'] \ntest_data['Familymembers_onboard'] = test_data['SibSp'] + test_data['Parch']","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:25:06.735085Z","iopub.execute_input":"2022-08-10T06:25:06.735539Z","iopub.status.idle":"2022-08-10T06:25:06.744922Z","shell.execute_reply.started":"2022-08-10T06:25:06.735502Z","shell.execute_reply":"2022-08-10T06:25:06.743460Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***A new attribute Familymembers_onboard has been generated from Sibsp and Parch. Therefore we need to drop Sibsp and Parch from the data as it will lead to multicollinearity.***","metadata":{}},{"cell_type":"code","source":"train_data = train_data.drop(['Parch','SibSp'], axis = 1)\ntest_data = test_data.drop(['Parch','SibSp'], axis = 1)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:25:08.365396Z","iopub.execute_input":"2022-08-10T06:25:08.365859Z","iopub.status.idle":"2022-08-10T06:25:08.377539Z","shell.execute_reply.started":"2022-08-10T06:25:08.365822Z","shell.execute_reply":"2022-08-10T06:25:08.376154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### *ii). Dummy Enconding:*","metadata":{}},{"cell_type":"code","source":"train_data = pd.get_dummies(train_data)\ntest_data = pd.get_dummies(test_data)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:25:10.153193Z","iopub.execute_input":"2022-08-10T06:25:10.153644Z","iopub.status.idle":"2022-08-10T06:25:10.173219Z","shell.execute_reply.started":"2022-08-10T06:25:10.153595Z","shell.execute_reply":"2022-08-10T06:25:10.171963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:25:11.874560Z","iopub.execute_input":"2022-08-10T06:25:11.875473Z","iopub.status.idle":"2022-08-10T06:25:11.894604Z","shell.execute_reply.started":"2022-08-10T06:25:11.875313Z","shell.execute_reply":"2022-08-10T06:25:11.893122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:25:14.197552Z","iopub.execute_input":"2022-08-10T06:25:14.198313Z","iopub.status.idle":"2022-08-10T06:25:14.213260Z","shell.execute_reply.started":"2022-08-10T06:25:14.198267Z","shell.execute_reply":"2022-08-10T06:25:14.211915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### *ii). Removing MultiCollinearity:*","metadata":{}},{"cell_type":"code","source":"#klib.corr_plot(train_data)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:25:20.252260Z","iopub.execute_input":"2022-08-10T06:25:20.252774Z","iopub.status.idle":"2022-08-10T06:25:20.259965Z","shell.execute_reply.started":"2022-08-10T06:25:20.252733Z","shell.execute_reply":"2022-08-10T06:25:20.258114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = train_data.drop(['Sex_male','Embarked_S'],axis = 1)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:25:22.111310Z","iopub.execute_input":"2022-08-10T06:25:22.111993Z","iopub.status.idle":"2022-08-10T06:25:22.121916Z","shell.execute_reply.started":"2022-08-10T06:25:22.111938Z","shell.execute_reply":"2022-08-10T06:25:22.120373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#klib.corr_plot(test_data)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data = test_data.drop(['Sex_male','Embarked_S'],axis = 1)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:25:30.222240Z","iopub.execute_input":"2022-08-10T06:25:30.222748Z","iopub.status.idle":"2022-08-10T06:25:30.232520Z","shell.execute_reply.started":"2022-08-10T06:25:30.222704Z","shell.execute_reply":"2022-08-10T06:25:30.230758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## *-> Splitting the Dataset:*","metadata":{}},{"cell_type":"code","source":"x = train_data.drop(['Survived'], axis = 1)\ny = train_data['Survived']","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:25:32.459735Z","iopub.execute_input":"2022-08-10T06:25:32.460294Z","iopub.status.idle":"2022-08-10T06:25:32.470387Z","shell.execute_reply.started":"2022-08-10T06:25:32.460245Z","shell.execute_reply":"2022-08-10T06:25:32.468904Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x_train, x_val, y_train, y_val = train_test_split(x, y, random_state = 42, test_size = 0.20)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:25:35.190978Z","iopub.execute_input":"2022-08-10T06:25:35.191875Z","iopub.status.idle":"2022-08-10T06:25:35.202946Z","shell.execute_reply.started":"2022-08-10T06:25:35.191819Z","shell.execute_reply":"2022-08-10T06:25:35.201783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### *-> Feature Scaling:*","metadata":{}},{"cell_type":"code","source":"scaler = StandardScaler()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:25:37.359176Z","iopub.execute_input":"2022-08-10T06:25:37.359807Z","iopub.status.idle":"2022-08-10T06:25:37.371995Z","shell.execute_reply.started":"2022-08-10T06:25:37.359753Z","shell.execute_reply":"2022-08-10T06:25:37.369622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:25:39.103859Z","iopub.execute_input":"2022-08-10T06:25:39.104340Z","iopub.status.idle":"2022-08-10T06:25:39.119978Z","shell.execute_reply.started":"2022-08-10T06:25:39.104301Z","shell.execute_reply":"2022-08-10T06:25:39.118958Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Fitting the scaler to the train data:\nx_train1 = scaler.fit_transform(x_train[['Age','Fare','Familymembers_onboard','Pclass']])\nx_train1 = pd.DataFrame(x_train1, columns = ['Age','Fare','Familymembers_onboard','Pclass'])\n\n#creating a new dataframe:\nx_train_new = x_train[['Sex_female','Embarked_Q','Embarked_C']].reset_index()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:25:40.944570Z","iopub.execute_input":"2022-08-10T06:25:40.945591Z","iopub.status.idle":"2022-08-10T06:25:40.960046Z","shell.execute_reply.started":"2022-08-10T06:25:40.945520Z","shell.execute_reply":"2022-08-10T06:25:40.958738Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x_train_new[['Age','Fare','Familymembers_onboard','Pclass']] = x_train1[['Age','Fare','Familymembers_onboard','Pclass']]\nx_train_new = pd.DataFrame(x_train_new)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:25:46.381798Z","iopub.execute_input":"2022-08-10T06:25:46.382267Z","iopub.status.idle":"2022-08-10T06:25:46.395884Z","shell.execute_reply.started":"2022-08-10T06:25:46.382223Z","shell.execute_reply":"2022-08-10T06:25:46.393753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Scaling w.r.t test data :\nx_val1 = scaler.transform(x_val[['Age','Fare','Familymembers_onboard','Pclass']])\nx_val1 = pd.DataFrame(x_val1, columns = ['Age','Fare','Familymembers_onboard','Pclass'])\n\n#creating a new dataframe:\nx_val_new = x_val[['Sex_female','Embarked_Q','Embarked_C']].reset_index()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:25:48.048426Z","iopub.execute_input":"2022-08-10T06:25:48.049008Z","iopub.status.idle":"2022-08-10T06:25:48.064112Z","shell.execute_reply.started":"2022-08-10T06:25:48.048956Z","shell.execute_reply":"2022-08-10T06:25:48.062657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x_val_new[['Age','Fare','Familymembers_onboard','Pclass']] = x_val1[['Age','Fare','Familymembers_onboard','Pclass']]\nx_val_new = pd.DataFrame(x_val_new)\nx_val = x_val_new","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:25:49.944379Z","iopub.execute_input":"2022-08-10T06:25:49.944932Z","iopub.status.idle":"2022-08-10T06:25:49.958573Z","shell.execute_reply.started":"2022-08-10T06:25:49.944888Z","shell.execute_reply":"2022-08-10T06:25:49.957452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## -> *Balancing the Dataset:*","metadata":{}},{"cell_type":"markdown","source":"*SMOTE stands for Synthetic Minority Oversampling Technique, is an oversampling technique that creates synthetic minority class data points to balance the dataset.It uses k-nearest neighbor algorithm to create synthetic data points.*","metadata":{}},{"cell_type":"code","source":"y_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:25:52.136396Z","iopub.execute_input":"2022-08-10T06:25:52.137218Z","iopub.status.idle":"2022-08-10T06:25:52.149919Z","shell.execute_reply.started":"2022-08-10T06:25:52.137159Z","shell.execute_reply":"2022-08-10T06:25:52.147754Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x_train_new.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:25:55.155308Z","iopub.execute_input":"2022-08-10T06:25:55.155777Z","iopub.status.idle":"2022-08-10T06:25:55.172360Z","shell.execute_reply.started":"2022-08-10T06:25:55.155738Z","shell.execute_reply":"2022-08-10T06:25:55.171379Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"The value Counts of the target class is :\")\nv = Counter(y_train)\nprint(v)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:25:57.064010Z","iopub.execute_input":"2022-08-10T06:25:57.064461Z","iopub.status.idle":"2022-08-10T06:25:57.072420Z","shell.execute_reply.started":"2022-08-10T06:25:57.064422Z","shell.execute_reply":"2022-08-10T06:25:57.071267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"smt = SMOTE()\nx_train ,y_train = smt.fit_resample(x_train_new, y_train)\nprint(\"After Balancing the data:\", Counter(y_train))","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:25:58.909758Z","iopub.execute_input":"2022-08-10T06:25:58.910199Z","iopub.status.idle":"2022-08-10T06:25:58.929527Z","shell.execute_reply.started":"2022-08-10T06:25:58.910163Z","shell.execute_reply":"2022-08-10T06:25:58.927868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Assigning it to a new variable:\nx_train = x_train[['Pclass','Sex_female','Embarked_Q','Embarked_C','Age','Fare','Familymembers_onboard']]\nx_val = x_val[['Pclass','Sex_female','Embarked_Q','Embarked_C','Age','Fare','Familymembers_onboard']]","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:26:00.629950Z","iopub.execute_input":"2022-08-10T06:26:00.630447Z","iopub.status.idle":"2022-08-10T06:26:00.639912Z","shell.execute_reply.started":"2022-08-10T06:26:00.630394Z","shell.execute_reply":"2022-08-10T06:26:00.638434Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4). Model Building :","metadata":{}},{"cell_type":"markdown","source":"***We will now use LazyPredict package which help us in camparing different algorithms for the dataset under consideration, after which the best performing algorithm is to be chosen.***","metadata":{}},{"cell_type":"code","source":"#warnings.filterwarnings(\"ignore\")\n#from lazypredict.Supervised import LazyClassifier\n#clf = LazyClassifier(predictions=True, random_state = 42)\n#models, predictions = clf.fit(x_train, x_val, y_train, y_val)\n#models","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:26:13.075072Z","iopub.execute_input":"2022-08-10T06:26:13.080374Z","iopub.status.idle":"2022-08-10T06:26:13.094164Z","shell.execute_reply.started":"2022-08-10T06:26:13.080189Z","shell.execute_reply":"2022-08-10T06:26:13.092848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Clearly, from the above table, we observe that AdaBooster is the best algorithm to be considered. Therefore we will now tune the parameters for AdaBooster.***","metadata":{}},{"cell_type":"markdown","source":"## -> Hyper Parameter Tuning of AdaBoost:","metadata":{}},{"cell_type":"code","source":"model = AdaBoostClassifier()\nparam = {'n_estimators':[10,50,100,500],\n         'learning_rate': [0.0001, 0.001, 0.01, 0.1, 1.0]}","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:26:15.994167Z","iopub.execute_input":"2022-08-10T06:26:15.994985Z","iopub.status.idle":"2022-08-10T06:26:16.003710Z","shell.execute_reply.started":"2022-08-10T06:26:15.994926Z","shell.execute_reply":"2022-08-10T06:26:16.002448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search = GridSearchCV(estimator = model, param_grid = param, cv = 10, n_jobs = -1)\ngrid_result = grid_search.fit(x_train,y_train)\nprint(\"The Best parameters will be :\",\n      grid_result.best_params_)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:26:17.739248Z","iopub.execute_input":"2022-08-10T06:26:17.739987Z","iopub.status.idle":"2022-08-10T06:26:50.824337Z","shell.execute_reply.started":"2022-08-10T06:26:17.739942Z","shell.execute_reply":"2022-08-10T06:26:50.822090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## -> Applying the updated Algorithm:","metadata":{}},{"cell_type":"code","source":"# Applying the tuned algorithm to train and checking for the performance.\nmodel1 = grid_result.predict(x_train)\naccu_train = sklearn.metrics.accuracy_score(y_train, model1)\nf1_train = sklearn.metrics.f1_score(y_train, model1)\n\n# Applying the tuned algorithm to test data and checking for the performance.\nmodel2 = grid_result.predict(x_val)\naccu_test = sklearn.metrics.accuracy_score(y_val, model2)\nf1_test = sklearn.metrics.f1_score(y_val, model2)\n\nprint(\"The Accuracy of AdaBoost Classifier with respect to TRAINING data :\",accu_train)\nprint(\"The F1 Score of AdaBoost Classifier with respect to TRAINING data :\",f1_train)\nprint(\"The Accuracy of AdaBoost Classifier with respect to TESTING data :\",accu_test)\nprint(\"The F1 Score of AdaBoost Classifier with respect to TESTING data :\",f1_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:26:54.057710Z","iopub.execute_input":"2022-08-10T06:26:54.058173Z","iopub.status.idle":"2022-08-10T06:26:54.325815Z","shell.execute_reply.started":"2022-08-10T06:26:54.058133Z","shell.execute_reply":"2022-08-10T06:26:54.324280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Here we can clearly observe that our performace of our model is comparitively decreased after tuning the parameters.Therefore we will now use the default parameters as it will help us build a generalized model.***","metadata":{}},{"cell_type":"markdown","source":"## -> Applying the Default Agorithm:","metadata":{}},{"cell_type":"code","source":"ada_clf = AdaBoostClassifier()\nada_clf.fit(x_train, y_train)\nprint(\"The mean Accuracy Score after Cross-validation is:\",\n     cross_val_score(ada_clf, x_train, y_train, cv = 20, scoring = 'accuracy').mean())","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:26:57.476618Z","iopub.execute_input":"2022-08-10T06:26:57.477661Z","iopub.status.idle":"2022-08-10T06:26:59.872222Z","shell.execute_reply.started":"2022-08-10T06:26:57.477585Z","shell.execute_reply":"2022-08-10T06:26:59.871229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Appling the model for test and train data :\nada_clf_train = ada_clf.predict(x_train)\nada_clf_test = ada_clf.predict(x_val)\n\nacc_score_train = sklearn.metrics.accuracy_score(y_train,ada_clf_train)\nf1_score_train = sklearn.metrics.f1_score(y_train, ada_clf_train)\nacc_score_test = sklearn.metrics.accuracy_score(y_val,ada_clf_test)\nf1_score_test = sklearn.metrics.f1_score(y_val, ada_clf_test)\n\nprint(\"The Accuracy of AdaBoost Classifier with respect to TRAINING data :\",acc_score_train)\nprint(\"The F1 Score of AdaBoost Classifier with respect to TRAINING data :\",f1_score_train)\nprint(\"The Accuracy of AdaBoost Classifier with respect to TESTING data :\",acc_score_test)\nprint(\"The F1 Score of AdaBoost Classifier with respect to TESTING data :\",f1_score_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:26:59.874478Z","iopub.execute_input":"2022-08-10T06:26:59.876167Z","iopub.status.idle":"2022-08-10T06:26:59.922195Z","shell.execute_reply.started":"2022-08-10T06:26:59.876103Z","shell.execute_reply":"2022-08-10T06:26:59.920746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5). Interpretation:","metadata":{}},{"cell_type":"markdown","source":"***After employing many different algorithms on this dataset, we finally choose AdaBoostClassifier as our predictive model thereby predicting which passengers survived the shipwreck.***","metadata":{}},{"cell_type":"markdown","source":"### -> Appling the model to the test dataset:","metadata":{}},{"cell_type":"code","source":"xtest = test_data\na = ada_clf.predict(xtest)\n#print(a)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:27:15.108067Z","iopub.execute_input":"2022-08-10T06:27:15.108525Z","iopub.status.idle":"2022-08-10T06:27:15.137640Z","shell.execute_reply.started":"2022-08-10T06:27:15.108489Z","shell.execute_reply":"2022-08-10T06:27:15.136097Z"},"trusted":true},"execution_count":null,"outputs":[]}]}