{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Titanic EDA","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:14:38.906671Z","iopub.execute_input":"2021-11-17T13:14:38.907253Z","iopub.status.idle":"2021-11-17T13:14:38.918995Z","shell.execute_reply.started":"2021-11-17T13:14:38.907215Z","shell.execute_reply":"2021-11-17T13:14:38.917981Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"First, we will import all the necessary libraries.","metadata":{"editable":false}},{"cell_type":"code","source":"import pandas as pd\nfrom pandas_profiling import ProfileReport\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport missingno as mno\nimport numpy as np\nimport math\n%matplotlib inline\nsns.set_style('whitegrid')","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:14:38.920901Z","iopub.execute_input":"2021-11-17T13:14:38.921607Z","iopub.status.idle":"2021-11-17T13:14:40.66206Z","shell.execute_reply.started":"2021-11-17T13:14:38.921556Z","shell.execute_reply":"2021-11-17T13:14:40.661013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's import the data","metadata":{"editable":false}},{"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/titanic/train.csv')\ntest_df = pd.read_csv('/kaggle/input/titanic/test.csv')","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:14:40.664592Z","iopub.execute_input":"2021-11-17T13:14:40.664912Z","iopub.status.idle":"2021-11-17T13:14:40.697841Z","shell.execute_reply.started":"2021-11-17T13:14:40.664883Z","shell.execute_reply":"2021-11-17T13:14:40.696902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As you can see, we have two datasets, train.csv and test.csv. Now lets just see the top 5 values of the train dataset and the shape for both of these datasets.","metadata":{"editable":false}},{"cell_type":"code","source":"train_df.head()","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:14:40.700211Z","iopub.execute_input":"2021-11-17T13:14:40.70055Z","iopub.status.idle":"2021-11-17T13:14:40.727793Z","shell.execute_reply.started":"2021-11-17T13:14:40.700518Z","shell.execute_reply":"2021-11-17T13:14:40.726948Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that there are some missing values in this dataset.","metadata":{"editable":false}},{"cell_type":"code","source":"train_df.shape","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:14:40.729032Z","iopub.execute_input":"2021-11-17T13:14:40.729542Z","iopub.status.idle":"2021-11-17T13:14:40.734658Z","shell.execute_reply.started":"2021-11-17T13:14:40.729508Z","shell.execute_reply":"2021-11-17T13:14:40.733892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.shape","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:14:40.735845Z","iopub.execute_input":"2021-11-17T13:14:40.736348Z","iopub.status.idle":"2021-11-17T13:14:40.746049Z","shell.execute_reply.started":"2021-11-17T13:14:40.736317Z","shell.execute_reply":"2021-11-17T13:14:40.745101Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Missing Numbers\n\nNow lets see which columns having missing numbers and form up a strategy to either drop those rows with missing numbers, or fill them up. To see which columns have missing numbers, we will use the missingno library which will help us recogonize the columns with missing values using data viz.","metadata":{"editable":false}},{"cell_type":"code","source":"mno.matrix(train_df)\nplt.show()","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:14:40.747648Z","iopub.execute_input":"2021-11-17T13:14:40.748158Z","iopub.status.idle":"2021-11-17T13:14:41.200114Z","shell.execute_reply.started":"2021-11-17T13:14:40.748124Z","shell.execute_reply":"2021-11-17T13:14:41.199389Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mno.matrix(test_df)\nplt.show()","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:14:41.201114Z","iopub.execute_input":"2021-11-17T13:14:41.201533Z","iopub.status.idle":"2021-11-17T13:14:41.576382Z","shell.execute_reply.started":"2021-11-17T13:14:41.201492Z","shell.execute_reply":"2021-11-17T13:14:41.57539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As you can see, the columns age and cabin have a lot of missing values. There also are 2 missing values in the embarked column and 1 in the fare column which can be dropped.\nTo handle the missing values in the train and the test dataset, we will first join both of these datasets since both of them are to be dealt with missing values.","metadata":{"editable":false}},{"cell_type":"code","source":"df = pd.concat([train_df, test_df])\ndf.shape","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:14:41.578648Z","iopub.execute_input":"2021-11-17T13:14:41.579266Z","iopub.status.idle":"2021-11-17T13:14:41.596399Z","shell.execute_reply.started":"2021-11-17T13:14:41.579218Z","shell.execute_reply":"2021-11-17T13:14:41.595278Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Filling up the missing values","metadata":{"editable":false}},{"cell_type":"code","source":"df = df.dropna(subset=['Embarked'])\ndf.shape","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:14:41.598868Z","iopub.execute_input":"2021-11-17T13:14:41.599912Z","iopub.status.idle":"2021-11-17T13:14:41.631045Z","shell.execute_reply.started":"2021-11-17T13:14:41.599851Z","shell.execute_reply":"2021-11-17T13:14:41.629909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have a missing fare value in the test dataset which needs to be filled. Lets just fill it with the mode of the fare column.","metadata":{"editable":false}},{"cell_type":"code","source":"df['Fare'] = df['Fare'].fillna(float(8.05)) #8.05 is the mode for the fare column. Can be calculated using df['Fare'].mode()\ndf.isna().sum()","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:14:41.634014Z","iopub.execute_input":"2021-11-17T13:14:41.634528Z","iopub.status.idle":"2021-11-17T13:14:41.648743Z","shell.execute_reply.started":"2021-11-17T13:14:41.634479Z","shell.execute_reply":"2021-11-17T13:14:41.646372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We dropped the 2 rows with nan values in the embarked columns. Now let's do some EDA to see how to deal with the missing values in the age column. For age, we cannot just straight away select the mean and fill in the missing values since a lot of missing values are there and some or the other feature does affect the age parameter in some way.","metadata":{"editable":false}},{"cell_type":"code","source":"sns.countplot(x='Survived', data=df.iloc[0:889, :], palette='GnBu')\nplt.show()","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:14:41.985614Z","iopub.execute_input":"2021-11-17T13:14:41.986218Z","iopub.status.idle":"2021-11-17T13:14:42.123337Z","shell.execute_reply.started":"2021-11-17T13:14:41.986156Z","shell.execute_reply":"2021-11-17T13:14:42.122293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot(x='Survived', data=df.iloc[0:889, :], palette='GnBu', hue='Sex')\nplt.show()","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:14:42.139431Z","iopub.execute_input":"2021-11-17T13:14:42.139813Z","iopub.status.idle":"2021-11-17T13:14:42.300195Z","shell.execute_reply.started":"2021-11-17T13:14:42.139776Z","shell.execute_reply":"2021-11-17T13:14:42.298938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot(x='Survived', data=df.iloc[0:889, :], palette='GnBu', hue='Pclass')\nplt.show()","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:14:42.828338Z","iopub.execute_input":"2021-11-17T13:14:42.828749Z","iopub.status.idle":"2021-11-17T13:14:43.032505Z","shell.execute_reply.started":"2021-11-17T13:14:42.828716Z","shell.execute_reply":"2021-11-17T13:14:43.031186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Pclass had a little or moderate effect on the survival rate. All the people in Pclass = 1 would have been nearest to the emergency boats and were saved first. Now let's see the mean age of people in all the Pclass.","metadata":{"editable":false}},{"cell_type":"code","source":"sns.boxplot(x='Pclass', y='Age', data=df, palette='GnBu')\nplt.show()","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:14:43.034635Z","iopub.execute_input":"2021-11-17T13:14:43.035071Z","iopub.status.idle":"2021-11-17T13:14:43.231387Z","shell.execute_reply.started":"2021-11-17T13:14:43.035027Z","shell.execute_reply":"2021-11-17T13:14:43.23051Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From here, we can see that for Pclass = 1, mean age is 39, 29 for Pclass = 2 and 25 for Pclass = 3. Pclass might be the feature that must have affected age the most and therefore we will replace the missing values in age with these values for the specific Pclass.\n\nI saw this in a video by krish naik, the link will be mentioned below.\n\nLink: https://www.youtube.com/watch?v=P_iMSYQnqac","metadata":{"editable":false}},{"cell_type":"code","source":"Age = list(df.Age)\nPclass = list(df.Pclass)\nfor i in range(len(Age)):\n    if math.isnan(Age[i]) == True:\n        if Pclass[i] == 1:\n            Age[i] = 39\n        elif Pclass[i] == 2:\n            Age[i] = 29\n        else:\n            Age[i] = 25\ndf['Age'] = Age\nprint(df.isna().sum())","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:14:43.233878Z","iopub.execute_input":"2021-11-17T13:14:43.234309Z","iopub.status.idle":"2021-11-17T13:14:43.251259Z","shell.execute_reply.started":"2021-11-17T13:14:43.234267Z","shell.execute_reply":"2021-11-17T13:14:43.249932Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As you can see, we have filled up all the nan values in the age column. Now we have 1014 missing values left in the cabin column. The total size of the dataframe is (1307, 12). The number of missing values is very large in the cabin column and thus we will drop the cabin column. We will also drop the ticket column and the passengerid column as they are just unique features.","metadata":{"editable":false}},{"cell_type":"code","source":"PassengerId = list(test_df['PassengerId'])","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:14:43.494878Z","iopub.execute_input":"2021-11-17T13:14:43.495374Z","iopub.status.idle":"2021-11-17T13:14:43.50034Z","shell.execute_reply.started":"2021-11-17T13:14:43.495342Z","shell.execute_reply":"2021-11-17T13:14:43.499307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = df.drop(['Cabin','Ticket','PassengerId'], axis=1)","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:14:43.680496Z","iopub.execute_input":"2021-11-17T13:14:43.681077Z","iopub.status.idle":"2021-11-17T13:14:43.687068Z","shell.execute_reply.started":"2021-11-17T13:14:43.681026Z","shell.execute_reply":"2021-11-17T13:14:43.686093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.columns","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:14:43.907153Z","iopub.execute_input":"2021-11-17T13:14:43.907805Z","iopub.status.idle":"2021-11-17T13:14:43.914885Z","shell.execute_reply.started":"2021-11-17T13:14:43.907748Z","shell.execute_reply":"2021-11-17T13:14:43.913867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Pandas Report\n\nLets use the pandas-profiling library to generate a report of the dataset that will help us get more insights.","metadata":{"editable":false}},{"cell_type":"code","source":"profile = ProfileReport(df, title=\"Titanic Dataset Report\")\nprofile.to_file(output_file=\"TitanicDatasetReport.html\")","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:14:44.545033Z","iopub.execute_input":"2021-11-17T13:14:44.545614Z","iopub.status.idle":"2021-11-17T13:15:01.854283Z","shell.execute_reply.started":"2021-11-17T13:14:44.545565Z","shell.execute_reply":"2021-11-17T13:15:01.853275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# EDA\n\nNow, let's do some data viz and understand the rest of the features in this dataset.","metadata":{"editable":false}},{"cell_type":"code","source":"sns.swarmplot(x=\"Pclass\", data=df, y=\"Fare\", hue=\"Survived\", palette=\"GnBu\")\nplt.show()","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:15:01.857057Z","iopub.execute_input":"2021-11-17T13:15:01.857564Z","iopub.status.idle":"2021-11-17T13:15:08.159221Z","shell.execute_reply.started":"2021-11-17T13:15:01.857516Z","shell.execute_reply":"2021-11-17T13:15:08.15832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.swarmplot(x='Survived', y=\"Fare\", data=df.iloc[0:889,:], palette=\"GnBu\")\nplt.show()","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:15:08.160768Z","iopub.execute_input":"2021-11-17T13:15:08.161068Z","iopub.status.idle":"2021-11-17T13:15:12.023975Z","shell.execute_reply.started":"2021-11-17T13:15:08.161037Z","shell.execute_reply":"2021-11-17T13:15:12.022922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From this graph, we can see that many people with same type of fare like the survived people have not survived. Another insight that we get from the two graphs above are that only people from Pclass 1 travelled with a fare > 100 and the survival rate of people with fare > 100 is good.","metadata":{"editable":false}},{"cell_type":"code","source":"sns.swarmplot(x='Sex', y=\"Fare\", hue=\"Survived\", data=df.iloc[0:889,:], palette=\"GnBu\")\nplt.show()","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:15:12.025399Z","iopub.execute_input":"2021-11-17T13:15:12.025777Z","iopub.status.idle":"2021-11-17T13:15:17.684363Z","shell.execute_reply.started":"2021-11-17T13:15:12.025745Z","shell.execute_reply":"2021-11-17T13:15:17.68359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"All females who paid more than 200 dollars survived. The survival rate for females is pretty good.","metadata":{"editable":false}},{"cell_type":"code","source":"sns.distplot(df.iloc[0:889,4:5], kde=False, bins=10, color=\"darkred\")\nplt.show()","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:15:17.686911Z","iopub.execute_input":"2021-11-17T13:15:17.687316Z","iopub.status.idle":"2021-11-17T13:15:17.990319Z","shell.execute_reply.started":"2021-11-17T13:15:17.687285Z","shell.execute_reply":"2021-11-17T13:15:17.989482Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The graph above shows the Age distribution.","metadata":{"editable":false}},{"cell_type":"code","source":"sns.countplot(x=\"SibSp\", data=df.iloc[0:889,:], palette=\"GnBu\", hue=\"Survived\")\nplt.show()","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:15:17.992627Z","iopub.execute_input":"2021-11-17T13:15:17.99312Z","iopub.status.idle":"2021-11-17T13:15:18.238994Z","shell.execute_reply.started":"2021-11-17T13:15:17.993088Z","shell.execute_reply":"2021-11-17T13:15:18.238136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot(x=\"Parch\", data=df.iloc[0:889,:], palette=\"GnBu\", hue=\"Survived\")\nplt.show()","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:15:18.240078Z","iopub.execute_input":"2021-11-17T13:15:18.240515Z","iopub.status.idle":"2021-11-17T13:15:18.469801Z","shell.execute_reply.started":"2021-11-17T13:15:18.240448Z","shell.execute_reply":"2021-11-17T13:15:18.46895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.scatterplot(x=\"Age\", y=\"Fare\", hue=\"Survived\", style=\"Sex\", data=df.iloc[0:889,:], palette=\"GnBu\")\nplt.show()","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:15:18.471088Z","iopub.execute_input":"2021-11-17T13:15:18.47135Z","iopub.status.idle":"2021-11-17T13:15:18.849542Z","shell.execute_reply.started":"2021-11-17T13:15:18.471324Z","shell.execute_reply":"2021-11-17T13:15:18.848546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"According to this plot, all the female survivors were the ones who paid a fare > 100 with 2 or 3 exceptions.","metadata":{"editable":false}},{"cell_type":"code","source":"sns.countplot(x=\"Embarked\", hue=\"Survived\", data=df.iloc[0:889,:], palette=\"GnBu\")\nplt.show()","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:15:18.850758Z","iopub.execute_input":"2021-11-17T13:15:18.851037Z","iopub.status.idle":"2021-11-17T13:15:19.031523Z","shell.execute_reply.started":"2021-11-17T13:15:18.851Z","shell.execute_reply":"2021-11-17T13:15:19.030501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot(x=\"Embarked\", hue=\"Pclass\", data=df.iloc[0:889,:], palette=\"GnBu\")\nplt.show()\ndf.head()","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:15:19.033144Z","iopub.execute_input":"2021-11-17T13:15:19.033446Z","iopub.status.idle":"2021-11-17T13:15:19.235154Z","shell.execute_reply.started":"2021-11-17T13:15:19.033417Z","shell.execute_reply":"2021-11-17T13:15:19.234108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Many form S = Southampton did not survive even though a lot of them were from Pclass = 1","metadata":{"editable":false}},{"cell_type":"markdown","source":"Finally, we will draw a heatmap to see the correlation between different features.","metadata":{"editable":false}},{"cell_type":"code","source":"Sex = list(df['Sex'])\nfor i in range(len(Sex)):\n          if Sex[i] == 'male':\n               Sex[i] = 1\n          elif Sex[i] == 'female':\n               Sex[i] = 0\ndf['Sex'] = Sex\nEmbarked = list(df['Embarked'])\nfor i in range(len(Embarked)):\n    if Embarked[i] == 'S':\n        Embarked[i] = 1\n    elif Embarked[i] == 'C':\n        Embarked[i] = 2\n    elif Embarked[i] == 'Q':\n        Embarked[i] = 3\ndf['Embarked'] = Embarked","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:15:19.236496Z","iopub.execute_input":"2021-11-17T13:15:19.236785Z","iopub.status.idle":"2021-11-17T13:15:19.251146Z","shell.execute_reply.started":"2021-11-17T13:15:19.236757Z","shell.execute_reply":"2021-11-17T13:15:19.249867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We replaced the embarked column with numerical value. Replacing the S,C and Q values with 1,2 and 3 doesn't change the feature type. Embarked still is a categorical feature. Same goes for the Sex Column. We did this so that we can plot a correlation heatmap.","metadata":{"editable":false}},{"cell_type":"markdown","source":"****With this, we end the EDA and as you can see, we gained a lot of information from all the visualizations.****","metadata":{"editable":false}},{"cell_type":"markdown","source":"# Feature Engineering\n\nTill now in the analysis we did, the name column wasn't useful at all. That is because everyone had a unique name, but even during that time, people were referred using Mr./Master or Mrs./Miss. So let's create a feature that will extract this information from the name column and let's see if this feature is important. The new feature that we will be creating would be a categorical feature.","metadata":{"editable":false}},{"cell_type":"code","source":"Names = list(df['Name'])\nTitles = []\nfor i in range(len(Names)):\n    title = Names[i].split(',')[1].split()[0]\n    Titles.append(title)\ndf['Title'] = Titles\nprint(df['Title'].value_counts())\ndf['Title'].loc[df['Title']=='the'] = 'Miss.'\ndf['Title'].loc[df['Title']=='Ms.'] = 'Miss.'\nprint(df['Title'].unique())","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:15:19.253021Z","iopub.execute_input":"2021-11-17T13:15:19.253483Z","iopub.status.idle":"2021-11-17T13:15:19.275924Z","shell.execute_reply.started":"2021-11-17T13:15:19.253415Z","shell.execute_reply":"2021-11-17T13:15:19.274937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"columns = ['Survived', 'Pclass','Title', 'Name', 'Sex', 'Age', 'SibSp', 'Parch', 'Fare',\n       'Embarked']\ndf = df[columns]\ndf = df.drop('Name', axis=1)\ndf.head()","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:15:19.277419Z","iopub.execute_input":"2021-11-17T13:15:19.278065Z","iopub.status.idle":"2021-11-17T13:15:19.297905Z","shell.execute_reply.started":"2021-11-17T13:15:19.278018Z","shell.execute_reply":"2021-11-17T13:15:19.297154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As you can see, we made a title column which tells us the title of the person. You might not understand some titles but they are just the french or spanish form of Miss or Mr. What we will do is keep them as it is till we train our model and validate it. After that is done, we will replace these french and spanish form of words to english and again train our model, see if the accuracy improves or not. Now lets do some EDA for this feature.","metadata":{"editable":false}},{"cell_type":"markdown","source":"# EDA(Continued)","metadata":{"editable":false}},{"cell_type":"code","source":"sns.countplot(x='Title', data=df.iloc[0:889,:], hue='Survived', palette='GnBu')\nplt.xticks(rotation=90)\nplt.show()","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:15:19.298919Z","iopub.execute_input":"2021-11-17T13:15:19.299301Z","iopub.status.idle":"2021-11-17T13:15:19.641917Z","shell.execute_reply.started":"2021-11-17T13:15:19.299272Z","shell.execute_reply":"2021-11-17T13:15:19.641046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.preprocessing import LabelEncoder\nlabelencoder = LabelEncoder()\ndf['Title'] = labelencoder.fit_transform(df['Title'])\ndf.head()","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:15:19.643243Z","iopub.execute_input":"2021-11-17T13:15:19.643547Z","iopub.status.idle":"2021-11-17T13:15:19.718814Z","shell.execute_reply.started":"2021-11-17T13:15:19.643518Z","shell.execute_reply":"2021-11-17T13:15:19.7177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"What we did above is known as label encoding. It is basically what we did with the sex column(replaced female and male with 0 and 1 respectively). Label encoding is done before training the model during the feature preprocessing. Here, we are doing it now for plotting a correlation heatmap as seaborn heatmaps only take numerical values\nNote: The values of the Title column are now numerical but it is still a categorical feature.","metadata":{"editable":false}},{"cell_type":"code","source":"columns = list(df.columns)\ncolumns.remove('Sex')\ncolumns\nheatmap = sns.heatmap(df[columns].iloc[0:889,:].corr(), cmap='GnBu')\nheatmap.set_title('Correlation Matrix')\nplt.xticks(rotation=90)\nplt.show()","metadata":{"editable":false,"execution":{"iopub.status.busy":"2021-11-17T13:15:19.720351Z","iopub.execute_input":"2021-11-17T13:15:19.720774Z","iopub.status.idle":"2021-11-17T13:15:20.009431Z","shell.execute_reply.started":"2021-11-17T13:15:19.720731Z","shell.execute_reply":"2021-11-17T13:15:20.00855Z"},"trusted":true},"execution_count":null,"outputs":[]}]}