{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"This is a solution for classic beginner's project 'Titanic - Machine Learning from Disaster'. From the history we know, that british passenger liner 'Titanic' sank on 15 april 1912 after striking an iceberg. From 2224 people aboard more than 1500 died. This famous catastrophe draw world's attention and was described as precise as possible. Today we have a data set of the passengers, who were aboard, with confirmation of which one survived or not. Based on this, we will work with  the test data to see if we can predict survivance of the passengers.\n","metadata":{}},{"cell_type":"markdown","source":"# 1. Importing libraries","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n%matplotlib inline","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:23:39.784509Z","iopub.execute_input":"2022-07-13T11:23:39.785177Z","iopub.status.idle":"2022-07-13T11:23:40.936975Z","shell.execute_reply.started":"2022-07-13T11:23:39.785083Z","shell.execute_reply":"2022-07-13T11:23:40.935618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For now we imported libraries needed for data wrangling. We'll download models later","metadata":{}},{"cell_type":"markdown","source":"# 2. Data wrangling\n## 2.1. Loading data","metadata":{}},{"cell_type":"code","source":"train_data = pd.read_csv('../input/titanic/train.csv')\ntest_data = pd.read_csv('../input/titanic/test.csv')","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:23:48.758481Z","iopub.execute_input":"2022-07-13T11:23:48.759622Z","iopub.status.idle":"2022-07-13T11:23:48.787581Z","shell.execute_reply.started":"2022-07-13T11:23:48.759576Z","shell.execute_reply":"2022-07-13T11:23:48.786562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's see what our data are consisted of","metadata":{}},{"cell_type":"markdown","source":"## 2.2. Data exploration","metadata":{}},{"cell_type":"code","source":"display(train_data.head())\ndisplay(test_data.head())","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:23:51.965389Z","iopub.execute_input":"2022-07-13T11:23:51.965826Z","iopub.status.idle":"2022-07-13T11:23:52.005832Z","shell.execute_reply.started":"2022-07-13T11:23:51.965790Z","shell.execute_reply":"2022-07-13T11:23:52.004923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(train_data.info())\ndisplay(test_data.info())","metadata":{"execution":{"iopub.status.busy":"2022-07-12T15:22:19.305885Z","iopub.execute_input":"2022-07-12T15:22:19.306734Z","iopub.status.idle":"2022-07-12T15:22:19.345703Z","shell.execute_reply.started":"2022-07-12T15:22:19.306688Z","shell.execute_reply":"2022-07-12T15:22:19.344777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 12 columns (11 for test data) with different type of entities (int, float, object).\nAlso we can see that some data are missing. For example, Age column  has 177 missing values for the train data  and 86 for the test data. We will fill it later. Embarked column for the train data and Fare column for the test data are missing 2 and 1 value. \n<p>The big problem is the Cabin column. It is almost empty for both data sets. We will drop this column</p>","metadata":{}},{"cell_type":"code","source":"train_data = train_data.drop('Cabin',axis=1)\ntest_data = test_data.drop('Cabin',axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:24:01.830323Z","iopub.execute_input":"2022-07-13T11:24:01.830708Z","iopub.status.idle":"2022-07-13T11:24:01.845625Z","shell.execute_reply.started":"2022-07-13T11:24:01.830655Z","shell.execute_reply":"2022-07-13T11:24:01.844621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's fill missing values","metadata":{}},{"cell_type":"markdown","source":"## 2.2. Filling missing values","metadata":{}},{"cell_type":"markdown","source":"We start with  Embarked column for the train data and Fare column for the test data","metadata":{}},{"cell_type":"code","source":"train_data['Embarked'].fillna(method = 'ffill',inplace=True)\ntest_data['Fare'].fillna(method = 'ffill',inplace=True)\ndisplay(train_data.info())\ndisplay(test_data.info())","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:24:05.828742Z","iopub.execute_input":"2022-07-13T11:24:05.829588Z","iopub.status.idle":"2022-07-13T11:24:05.866531Z","shell.execute_reply.started":"2022-07-13T11:24:05.829553Z","shell.execute_reply":"2022-07-13T11:24:05.865749Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, let's handle Age column for both datasets. The simplest way is to fill by mean value of all rows. But we will go further and explore dependency of Age on different parameters.\n<p>Before we continue with Age column it is better to transform object data into int/float one. Obviously, we can't transform names so let's just drop the column</p>","metadata":{}},{"cell_type":"code","source":"train_data = train_data.drop('Name',axis=1)\ntest_data = test_data.drop('Name',axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:24:07.928311Z","iopub.execute_input":"2022-07-13T11:24:07.929081Z","iopub.status.idle":"2022-07-13T11:24:07.935863Z","shell.execute_reply.started":"2022-07-13T11:24:07.929040Z","shell.execute_reply":"2022-07-13T11:24:07.934965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next target is the Sex column. We will transform it by command *get_dummies*","metadata":{}},{"cell_type":"code","source":"train_data = pd.get_dummies(train_data,columns=['Sex'],drop_first = True)\ntrain_data.rename(columns = {'Sex_male':'Sex'}, inplace = True)\ntest_data = pd.get_dummies(test_data,columns=['Sex'],drop_first = True)\ntest_data.rename(columns = {'Sex_male':'Sex'}, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:24:19.903947Z","iopub.execute_input":"2022-07-13T11:24:19.904338Z","iopub.status.idle":"2022-07-13T11:24:19.923339Z","shell.execute_reply.started":"2022-07-13T11:24:19.904297Z","shell.execute_reply":"2022-07-13T11:24:19.922324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, we will explore the Ticket column.","metadata":{}},{"cell_type":"code","source":"display(train_data['Ticket'].nunique())\ndisplay(test_data['Ticket'].nunique())","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:24:23.013696Z","iopub.execute_input":"2022-07-13T11:24:23.014310Z","iopub.status.idle":"2022-07-13T11:24:23.024808Z","shell.execute_reply.started":"2022-07-13T11:24:23.014261Z","shell.execute_reply":"2022-07-13T11:24:23.023717Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, most of the Tickets are unique, so we can't replace them with integers. Let's drop this column.","metadata":{}},{"cell_type":"code","source":"train_data = train_data.drop('Ticket',axis=1)\ntest_data = test_data.drop('Ticket',axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:24:25.573113Z","iopub.execute_input":"2022-07-13T11:24:25.573721Z","iopub.status.idle":"2022-07-13T11:24:25.581452Z","shell.execute_reply.started":"2022-07-13T11:24:25.573685Z","shell.execute_reply":"2022-07-13T11:24:25.580159Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lastly, we will transform Embarked column.","metadata":{}},{"cell_type":"code","source":"train_data.loc[train_data['Embarked'] == 'S', 'Embarked'] = 0\ntrain_data.loc[train_data['Embarked'] == 'C', 'Embarked'] = 1\ntrain_data.loc[train_data['Embarked'] == 'Q', 'Embarked'] = 2\ntest_data.loc[test_data['Embarked'] == 'S', 'Embarked'] = 0\ntest_data.loc[test_data['Embarked'] == 'C', 'Embarked'] = 1\ntest_data.loc[test_data['Embarked'] == 'Q', 'Embarked'] = 2","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:24:28.170737Z","iopub.execute_input":"2022-07-13T11:24:28.171139Z","iopub.status.idle":"2022-07-13T11:24:28.185186Z","shell.execute_reply.started":"2022-07-13T11:24:28.171108Z","shell.execute_reply":"2022-07-13T11:24:28.183949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.to_numeric(train_data['Embarked'])\ntrain_data['Embarked'] = train_data['Embarked'].astype(int)\npd.to_numeric(test_data['Embarked'])\ntest_data['Embarked'] = test_data['Embarked'].astype(int)","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:24:30.872474Z","iopub.execute_input":"2022-07-13T11:24:30.873637Z","iopub.status.idle":"2022-07-13T11:24:30.882164Z","shell.execute_reply.started":"2022-07-13T11:24:30.873580Z","shell.execute_reply":"2022-07-13T11:24:30.881043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(train_data.info())\ndisplay(test_data.info())","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:24:34.716884Z","iopub.execute_input":"2022-07-13T11:24:34.717279Z","iopub.status.idle":"2022-07-13T11:24:34.745100Z","shell.execute_reply.started":"2022-07-13T11:24:34.717249Z","shell.execute_reply":"2022-07-13T11:24:34.743935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we have transformed all data into numbers. Let's explore the correlation between Age and other parameters.","metadata":{}},{"cell_type":"code","source":"train_data.corr(method ='pearson')['Age']","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:24:37.849423Z","iopub.execute_input":"2022-07-13T11:24:37.850105Z","iopub.status.idle":"2022-07-13T11:24:37.862608Z","shell.execute_reply.started":"2022-07-13T11:24:37.850064Z","shell.execute_reply":"2022-07-13T11:24:37.861307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, there is a quite good correlation with Pclass, SibSp, Parch, Fare and Sex. Let's explore number of unique values for SibSp,Parch and Fare.","metadata":{}},{"cell_type":"code","source":"display(train_data['SibSp'].nunique())\ndisplay(train_data['Parch'].nunique())\ndisplay(train_data['Fare'].nunique())","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:24:40.269800Z","iopub.execute_input":"2022-07-13T11:24:40.270198Z","iopub.status.idle":"2022-07-13T11:24:40.285351Z","shell.execute_reply.started":"2022-07-13T11:24:40.270167Z","shell.execute_reply":"2022-07-13T11:24:40.283833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There is too much unique elements for Fare column to make our Age predictions based on it. SibSp and Parch columns also have quite big numbers. We will fill Age column based on Pclass and Sex columns.","metadata":{}},{"cell_type":"code","source":"train_data.groupby(['Pclass','Sex'], as_index=False)['Age'].mean().round(0)","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:24:43.029850Z","iopub.execute_input":"2022-07-13T11:24:43.030254Z","iopub.status.idle":"2022-07-13T11:24:43.048692Z","shell.execute_reply.started":"2022-07-13T11:24:43.030220Z","shell.execute_reply":"2022-07-13T11:24:43.047897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def impute_age(cols):\n    Age = cols[0]\n    Pclass = cols[1]\n    Sex = cols[2]\n    \n    if pd.isnull(Age):\n\n        if Sex == 0:\n            if Pclass == 1:\n                return 35\n\n            elif Pclass == 2:\n                return 29\n\n            else:\n                return 23\n        else:\n            if Pclass == 1:\n                return 41\n\n            elif Pclass == 2:\n                return 31\n\n            else:\n                return 26\n\n    else:\n        return Age","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:24:46.749446Z","iopub.execute_input":"2022-07-13T11:24:46.750225Z","iopub.status.idle":"2022-07-13T11:24:46.758943Z","shell.execute_reply.started":"2022-07-13T11:24:46.750173Z","shell.execute_reply":"2022-07-13T11:24:46.757731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data['Age'] = train_data[['Age','Pclass','Sex']].apply(impute_age,axis=1)\ntest_data['Age'] = test_data[['Age','Pclass','Sex']].apply(impute_age,axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:24:49.069920Z","iopub.execute_input":"2022-07-13T11:24:49.070717Z","iopub.status.idle":"2022-07-13T11:24:49.107687Z","shell.execute_reply.started":"2022-07-13T11:24:49.070649Z","shell.execute_reply":"2022-07-13T11:24:49.106637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(train_data.info())\ndisplay(test_data.info())","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:24:51.588266Z","iopub.execute_input":"2022-07-13T11:24:51.588699Z","iopub.status.idle":"2022-07-13T11:24:51.620028Z","shell.execute_reply.started":"2022-07-13T11:24:51.588650Z","shell.execute_reply":"2022-07-13T11:24:51.618762Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, as we filled and transformed all columns, let's create some graphs!","metadata":{}},{"cell_type":"markdown","source":"# 3. Data visualisation","metadata":{}},{"cell_type":"markdown","source":"Firstly, we delete Id column from our train data","metadata":{}},{"cell_type":"code","source":"train_data = train_data.drop('PassengerId',axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:24:55.155104Z","iopub.execute_input":"2022-07-13T11:24:55.155914Z","iopub.status.idle":"2022-07-13T11:24:55.161842Z","shell.execute_reply.started":"2022-07-13T11:24:55.155871Z","shell.execute_reply":"2022-07-13T11:24:55.160466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's start with heatmap of correlations. It'll help us to understand, which parameter is more correlated with the  survival data.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(18,12))\nheat = sns.heatmap(data=train_data.corr(), annot=True, cmap='coolwarm',alpha=0.5)\nheat.tick_params(axis='x', labelsize=14)\nheat.tick_params(axis='y', labelsize=14)\nheat.set_title('Training Set Correlations', size=20)","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:24:58.970076Z","iopub.execute_input":"2022-07-13T11:24:58.970482Z","iopub.status.idle":"2022-07-13T11:24:59.632019Z","shell.execute_reply.started":"2022-07-13T11:24:58.970449Z","shell.execute_reply":"2022-07-13T11:24:59.630629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, survival rate has the best correlation with Sex, then Pclass and then Fare. Other parameters show lesser influence on the survival rate.\n<p>Now let's see how many passengers survived.</p>","metadata":{}},{"cell_type":"code","source":"survived = train_data['Survived'].value_counts()[1]\nnot_survived = train_data['Survived'].value_counts()[0]\nsurvived_per = survived / train_data.shape[0] * 100\nnot_survived_per = not_survived / train_data.shape[0] * 100\n\n\n\nplt.figure(figsize=(10, 8))\nsns.set_style('darkgrid')\nsns.countplot(x = train_data['Survived'], palette = 'coolwarm')\n\nplt.xlabel('Survival', size=15, labelpad=15)\nplt.ylabel('Passenger Count', size=15, labelpad=15)\nplt.xticks((0, 1), ['Not Survived ({0:.2f}%)'.format(not_survived_per), 'Survived ({0:.2f}%)'.format(survived_per)])\nplt.tick_params(axis='x', labelsize=13)\nplt.tick_params(axis='y', labelsize=13)\n\nplt.title('Training Set Survival Distribution', size=15, y=1.05)\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:25:05.801280Z","iopub.execute_input":"2022-07-13T11:25:05.801701Z","iopub.status.idle":"2022-07-13T11:25:05.992063Z","shell.execute_reply.started":"2022-07-13T11:25:05.801647Z","shell.execute_reply":"2022-07-13T11:25:05.990847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Only 38% of  passengers, which is a huge disaster!","metadata":{}},{"cell_type":"markdown","source":"Now, let's explore different parameters. We will start with the Pclass. It's interesting to see the distribution of passengers between classes and survival rate between classes.","metadata":{}},{"cell_type":"code","source":"f,ax=plt.subplots(1,2,figsize=(18,8))\nsns.set_style('darkgrid')\n\ncc = sns.countplot(x='Pclass',data = train_data, palette='coolwarm', ax = ax[0])\nfor p in cc.patches:\n        cc.annotate('{:.1f}'.format(p.get_height()), (p.get_x()+0.1, p.get_height()+5))\nax[0].set_title('Number of passengers for each class')\nax[0].set_ylabel('Passenger count')\n\ncs = sns.countplot(x='Pclass',data = train_data,hue = 'Survived', palette='coolwarm',ax=ax[1])\nfor p in cs.patches:\n        cs.annotate('{:.1f}'.format(p.get_height()), (p.get_x()+0.1, p.get_height()+5))\nax[1].set_title('Survival distribution of passengers for each class')\nax[1].set_ylabel('Passenger count')\nax[1].legend(['Not Survived', 'Survived'], loc='upper left', prop={'size': 12})\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:25:11.328684Z","iopub.execute_input":"2022-07-13T11:25:11.329475Z","iopub.status.idle":"2022-07-13T11:25:11.726829Z","shell.execute_reply.started":"2022-07-13T11:25:11.329428Z","shell.execute_reply":"2022-07-13T11:25:11.725724Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Most of the passengers were in the 3rd class, where survival rate is horrible (only 24% of passenges from 3rd class survived). The 2nd class has almost 50/50 survival rate and the 1st class has 63% survival rate.","metadata":{}},{"cell_type":"markdown","source":"Let's explore age distribution of passengers.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(20,16))\nsns.set_style('darkgrid')\ntotal = len(train_data['Age'])*1.\nca = sns.histplot(x='Age',hue = train_data['Survived'],data=train_data, palette = 'coolwarm')\nfor p in ca.patches:\n        ca.annotate('{:.1f}%'.format(100*p.get_height()/total), (p.get_x()+0.5, p.get_height()+2))\nca.set_ylabel('Passenger count')\nca.legend(['Survived', 'Not survived'], loc='upper right', prop={'size': 15})","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:25:16.335357Z","iopub.execute_input":"2022-07-13T11:25:16.336068Z","iopub.status.idle":"2022-07-13T11:25:17.091808Z","shell.execute_reply.started":"2022-07-13T11:25:16.336029Z","shell.execute_reply":"2022-07-13T11:25:17.090547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Most of the passengers were in the age between 20 to 30 years. This age category also has the smallest survival rate.","metadata":{}},{"cell_type":"markdown","source":"Now we will explore SibSp and Parch columns. Let's combine them in two subplots.","metadata":{}},{"cell_type":"code","source":"f,ax=plt.subplots(1,2,figsize=(18,8))\nsns.set_style('darkgrid')\ntotal = len(train_data['SibSp'])*1.\nca = sns.countplot(x='SibSp',hue = train_data['Survived'],data=train_data, palette='coolwarm', ax=ax[0])\nfor p in ca.patches:\n        ca.annotate('{:.1f}%'.format(100*p.get_height()/total), (p.get_x()+0.1, p.get_height()+2))\nax[0].set_title('Survival distribution of passengers for SibSp')\nax[0].set_ylabel('Passenger count')\nax[0].legend(['Not Survived', 'Survived'], loc='upper right', prop={'size': 12})\n\nca = sns.countplot(x='Parch',hue = train_data['Survived'],data=train_data, palette='coolwarm', ax=ax[1])\nfor p in ca.patches:\n        ca.annotate('{:.1f}%'.format(100*p.get_height()/total), (p.get_x()+0, p.get_height()+2))\nax[1].set_title('Survival distribution of passengers for Parch')\nax[1].set_ylabel('Passenger count')\nax[1].legend(['Not Survived', 'Survived'], loc='upper right', prop={'size': 12})","metadata":{"execution":{"iopub.status.busy":"2022-07-13T12:26:53.445259Z","iopub.execute_input":"2022-07-13T12:26:53.446023Z","iopub.status.idle":"2022-07-13T12:26:54.047495Z","shell.execute_reply.started":"2022-07-13T12:26:53.445983Z","shell.execute_reply":"2022-07-13T12:26:54.045943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Most of the passengers were alone. Interesting, that survival rate of the passengers with 1 sibling/spouse and 1 parent/children is the highest.","metadata":{}},{"cell_type":"markdown","source":"Next column is the Fare. Let's see passenger distribution along with the survival rate.","metadata":{}},{"cell_type":"code","source":"fig, axs = plt.subplots(figsize=(22, 9))\nsns.countplot(x=pd.qcut(train_data['Fare'], 13), hue='Survived', data=train_data, palette='coolwarm')\n\nplt.xlabel('Fare', size=15, labelpad=20)\nplt.ylabel('Passenger Count', size=15, labelpad=20)\nplt.tick_params(axis='x', labelsize=10)\nplt.tick_params(axis='y', labelsize=15)\n\nplt.legend(['Not Survived', 'Survived'], loc='upper right', prop={'size': 15})\nplt.title('Count of Survival in {} Feature'.format('Fare'), size=15, y=1.05)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:25:29.026346Z","iopub.execute_input":"2022-07-13T11:25:29.027249Z","iopub.status.idle":"2022-07-13T11:25:29.403306Z","shell.execute_reply.started":"2022-07-13T11:25:29.027200Z","shell.execute_reply":"2022-07-13T11:25:29.402137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The higher the price of the ticket - the higher survival rate. The heatmap we created earlier showed high correlation between the Fare and the Pclass. Let's explore it then.","metadata":{}},{"cell_type":"code","source":"fig, axs = plt.subplots(figsize=(22, 9))\ncs = sns.countplot(x=pd.qcut(train_data['Fare'], 13), hue='Pclass', data=train_data, palette='coolwarm')\n\nplt.xlabel('Fare', size=15, labelpad=20)\nplt.ylabel('Passenger Count', size=15, labelpad=20)\nplt.tick_params(axis='x', labelsize=10)\nplt.tick_params(axis='y', labelsize=15)\nfor p in cs.patches:\n        cs.annotate('{:.1f}'.format(p.get_height()), (p.get_x()+0, p.get_height()+2))\n\nplt.legend(['1', '2', '3'], loc='upper right', prop={'size': 15})\nplt.title('Fare distribution between classes', size=15, y=1.05)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:25:33.839761Z","iopub.execute_input":"2022-07-13T11:25:33.841034Z","iopub.status.idle":"2022-07-13T11:25:34.422130Z","shell.execute_reply.started":"2022-07-13T11:25:33.840979Z","shell.execute_reply":"2022-07-13T11:25:34.420746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Except 12 odd values for the 1st and the 2nd classes (within range 0-7.229), the plot confirms correlation between the cost of the ticket and the class.","metadata":{}},{"cell_type":"markdown","source":"Now, we will explore the Embarked column.","metadata":{}},{"cell_type":"code","source":"f,ax=plt.subplots(1,2,figsize=(18,8))\nsns.set_style('darkgrid')\n\ncc = sns.countplot(x='Embarked',data = train_data, palette='coolwarm', ax = ax[0])\nfor p in cc.patches:\n        cc.annotate('{:.1f}'.format(p.get_height()), (p.get_x()+0.1, p.get_height()+5))\nax[0].set_title('Number of passengers for each port')\nax[0].set_xticklabels(['S', 'C', 'Q'])\nax[0].set_ylabel('Passenger count')\n\ncs = sns.countplot(x='Embarked',data = train_data,hue = 'Survived', palette='coolwarm',ax=ax[1])\nfor p in cs.patches:\n        cs.annotate('{:.1f}'.format(p.get_height()), (p.get_x()+0.1, p.get_height()+5))\nax[1].set_title('Survival distribution of passengers for each port')\nax[1].set_xticklabels(['S', 'C', 'Q'])\nax[1].set_ylabel('Passenger count')\nax[1].legend(['Not Survived', 'Survived'], loc='upper right', prop={'size': 12})\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:25:39.283607Z","iopub.execute_input":"2022-07-13T11:25:39.284027Z","iopub.status.idle":"2022-07-13T11:25:39.684686Z","shell.execute_reply.started":"2022-07-13T11:25:39.283994Z","shell.execute_reply":"2022-07-13T11:25:39.683360Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Most of the passengers were from the \"S\" port. Suprisingly, the \"C\" port shows higher ratio of survived. As we know for now, highest correlation with the \"Survival\"  have \"Sex\" and \"Pclass\" parameters. Let's explore them.","metadata":{}},{"cell_type":"code","source":"f,ax=plt.subplots(1,2,figsize=(18,8))\nsns.set_style('darkgrid')\ncs = sns.countplot(x='Embarked',data = train_data,hue = 'Pclass', palette='coolwarm', ax=ax[0])\nfor p in cs.patches:\n        cs.annotate('{:.1f}'.format(p.get_height()), (p.get_x()+0.1, p.get_height()+5))\nax[0].set_title('Class distribution  for each port')\nax[0].set_xticklabels(['S', 'C', 'Q'])\nax[0].set_ylabel('Passenger count')\nax[0].legend(['1', '2', '3'], loc='upper right', prop={'size': 12})\n\ncs = sns.countplot(x='Embarked',data = train_data,hue = 'Sex', palette='coolwarm', ax=ax[1])\nfor p in cs.patches:\n        cs.annotate('{:.1f}'.format(p.get_height()), (p.get_x()+0.1, p.get_height()+5))\nax[1].set_title('Sex distribution  for each port')\nax[1].set_xticklabels(['S', 'C', 'Q'])\nax[1].set_ylabel('Passenger count')\nax[1].legend(['Female', 'Male'], loc='upper right', prop={'size': 12})\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:25:43.593201Z","iopub.execute_input":"2022-07-13T11:25:43.593619Z","iopub.status.idle":"2022-07-13T11:25:44.185100Z","shell.execute_reply.started":"2022-07-13T11:25:43.593587Z","shell.execute_reply":"2022-07-13T11:25:44.183940Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see, that the class distribution for the \"C\" port is different then for other two. Sex distribution shows, that both ports \"C\" and \"Q\" have almost equal numbers of male and female passengers while number of male passengers from \"S\" port is more than twice the number of female passengers.","metadata":{}},{"cell_type":"markdown","source":"Now we continue with \"Sex\" column.","metadata":{}},{"cell_type":"code","source":"\nfig, axes = plt.subplots(1, 3, figsize=(20,8))\n\naxes[0].pie(train_data.loc[train_data['Sex'],'Sex'].value_counts(),labels = ['Male','Female'],autopct='%1.0f%%', colors = ['#ff9950','#66b3ff'],wedgeprops={'alpha':0.5})\naxes[0].set_title('Sex distribution', fontsize = 15)\n\naxes[1].pie(train_data.loc[train_data[\"Sex\"] == 1, \"Survived\"].value_counts(),labels = ['No survived','Survived'],autopct='%1.0f%%',colors = ['#ff9950','#66b3ff'],wedgeprops={'alpha':0.5})\naxes[1].set_title('% of survived male', fontsize = 15)\n\naxes[2].pie(train_data.loc[train_data[\"Sex\"] == 0, \"Survived\"].value_counts(),labels = ['Survived','No survived'], autopct='%1.0f%%',colors = ['#ff9950','#66b3ff'],wedgeprops={'alpha':0.5})\naxes[2].set_title('% of survived female', fontsize = 15)\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:25:47.569630Z","iopub.execute_input":"2022-07-13T11:25:47.570069Z","iopub.status.idle":"2022-07-13T11:25:47.873209Z","shell.execute_reply.started":"2022-07-13T11:25:47.570035Z","shell.execute_reply":"2022-07-13T11:25:47.871819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As it was non-directly confirmed earlier, the number of male passengers is bigger than the number of female passengers. However, the survival rate is higher for female passengers.","metadata":{}},{"cell_type":"markdown","source":"For the end, let's create some violin plots. For example, Pclass-Age and Sex-Age distribution with Survived as parameter.","metadata":{}},{"cell_type":"code","source":"f,ax=plt.subplots(1,2,figsize=(18,8))\nsns.set_style('darkgrid')\nsns.violinplot(x = \"Pclass\",y = \"Age\", hue=\"Survived\", data=train_data,palette='coolwarm',split=True,ax=ax[0])\nax[0].set_title('Pclass and Age vs Survived')\nax[0].set_yticks(range(0,110,10))\nax[0].legend(handles=ax[0].legend_.legendHandles, labels=['Not survived', 'Survived'])\nsns.violinplot(x = \"Sex\",y = \"Age\", hue=\"Survived\", data=train_data,palette = 'coolwarm',split=True,ax=ax[1])\nax[1].set_title('Sex and Age vs Survived')\nax[1].set_yticks(range(0,110,10))\nax[1].legend(handles=ax[1].legend_.legendHandles, labels=['Not survived', 'Survived'])\nax[1].set_xticklabels(['Female', 'Male'])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:25:52.801459Z","iopub.execute_input":"2022-07-13T11:25:52.801888Z","iopub.status.idle":"2022-07-13T11:25:53.321864Z","shell.execute_reply.started":"2022-07-13T11:25:52.801852Z","shell.execute_reply":"2022-07-13T11:25:53.320575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, older passenger have higher class, but younger passenger have better survival rate for all classes.","metadata":{}},{"cell_type":"markdown","source":"Now, let's train our model and make a predictions!","metadata":{}},{"cell_type":"markdown","source":"# 4. Predictons\n## 4.1. Models import ","metadata":{}},{"cell_type":"markdown","source":"We will import some base models for making predictions.","metadata":{}},{"cell_type":"code","source":"from sklearn.linear_model import LogisticRegression \nfrom sklearn.ensemble import RandomForestClassifier \nfrom sklearn.neighbors import KNeighborsClassifier \nfrom sklearn.tree import DecisionTreeClassifier \nfrom sklearn.model_selection import train_test_split \nfrom sklearn import metrics \nfrom sklearn.metrics import classification_report\nfrom sklearn.metrics import confusion_matrix ","metadata":{"execution":{"iopub.status.busy":"2022-07-13T12:39:39.238640Z","iopub.execute_input":"2022-07-13T12:39:39.239084Z","iopub.status.idle":"2022-07-13T12:39:39.244746Z","shell.execute_reply.started":"2022-07-13T12:39:39.239047Z","shell.execute_reply":"2022-07-13T12:39:39.243803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we will split our train data so we can observe which model has better score.","metadata":{}},{"cell_type":"code","source":"X = train_data.drop('Survived', axis =1)\ny = train_data['Survived']\nX_train, X_test,y_train, y_test = train_test_split(X,y,test_size = 0.3, random_state = 101)","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:26:04.690176Z","iopub.execute_input":"2022-07-13T11:26:04.690585Z","iopub.status.idle":"2022-07-13T11:26:04.700587Z","shell.execute_reply.started":"2022-07-13T11:26:04.690551Z","shell.execute_reply":"2022-07-13T11:26:04.699640Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"First model will be logistic regression.","metadata":{}},{"cell_type":"code","source":"logreg = LogisticRegression()\n\nlogreg.fit(X_train, y_train)\n\ny_pred = logreg.predict(X_test)\n\nprint(classification_report(y_test,y_pred))\nprint(confusion_matrix(y_test,y_pred))\n","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:31:38.323994Z","iopub.execute_input":"2022-07-13T11:31:38.324383Z","iopub.status.idle":"2022-07-13T11:31:38.369364Z","shell.execute_reply.started":"2022-07-13T11:31:38.324352Z","shell.execute_reply":"2022-07-13T11:31:38.367999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Logistic regression shows pretty good score. We will continue with K neighbor classifier. Additionally, we will observe for which number of neighbors the model shows the best results.","metadata":{}},{"cell_type":"markdown","source":"Start the model with for loop.","metadata":{}},{"cell_type":"code","source":"error = []\nfor i in range (1,40):\n    \n    knn = KNeighborsClassifier(n_neighbors=i)\n    knn.fit(X_train, y_train)\n    y_predi = knn.predict(X_test)\n    error.append(np.mean(y_predi != y_test))\n","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:26:14.731182Z","iopub.execute_input":"2022-07-13T11:26:14.731792Z","iopub.status.idle":"2022-07-13T11:26:15.320849Z","shell.execute_reply.started":"2022-07-13T11:26:14.731757Z","shell.execute_reply":"2022-07-13T11:26:15.319638Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, let's print error distribution for the number of neighbors","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10,6))\ner = plt.plot(range(1,40),error,color='blue',marker='o',markerfacecolor='orange',markersize=10)\nplt.title('Error distribution  for the different number of neighbors')\nplt.xlabel('Number of neighbors')\nplt.ylabel('Error')","metadata":{"execution":{"iopub.status.busy":"2022-07-13T12:46:25.569220Z","iopub.execute_input":"2022-07-13T12:46:25.569602Z","iopub.status.idle":"2022-07-13T12:46:25.848385Z","shell.execute_reply.started":"2022-07-13T12:46:25.569571Z","shell.execute_reply":"2022-07-13T12:46:25.847296Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The smallest error is for 7 neighbors. We will use this number for our model.","metadata":{}},{"cell_type":"code","source":"knn = KNeighborsClassifier(n_neighbors=7)\nknn.fit(X_train, y_train)\ny_pred = knn.predict(X_test)\n\nprint(classification_report(y_test,y_pred))\nprint(confusion_matrix(y_test,y_pred))","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:32:47.222833Z","iopub.execute_input":"2022-07-13T11:32:47.223219Z","iopub.status.idle":"2022-07-13T11:32:47.261600Z","shell.execute_reply.started":"2022-07-13T11:32:47.223188Z","shell.execute_reply":"2022-07-13T11:32:47.260374Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Even with the best model parameters K neighbor classifier model is worse in predictions then logistic regression model.","metadata":{}},{"cell_type":"markdown","source":"Now, let's use the decision tree classifier model.","metadata":{}},{"cell_type":"code","source":"dtree = DecisionTreeClassifier()\ndtree.fit(X_train, y_train)\ny_pred = dtree.predict(X_test)\nprint(classification_report(y_test,y_pred))\nprint(confusion_matrix(y_test,y_pred))","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:32:55.611015Z","iopub.execute_input":"2022-07-13T11:32:55.611457Z","iopub.status.idle":"2022-07-13T11:32:55.630842Z","shell.execute_reply.started":"2022-07-13T11:32:55.611421Z","shell.execute_reply":"2022-07-13T11:32:55.629628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This model shows the same accuracy, precision and f1-score as logistic regression.","metadata":{}},{"cell_type":"markdown","source":"Now, we will use our last model: the random forestt classifier.","metadata":{}},{"cell_type":"code","source":"random_forest = RandomForestClassifier(n_estimators=200)\n\nrandom_forest.fit(X_train, y_train)\n\ny_pred = random_forest.predict(X_test)\n\nprint(classification_report(y_test,y_pred))\nprint(confusion_matrix(y_test,y_pred))","metadata":{"execution":{"iopub.status.busy":"2022-07-13T11:56:23.639797Z","iopub.execute_input":"2022-07-13T11:56:23.640747Z","iopub.status.idle":"2022-07-13T11:56:24.098867Z","shell.execute_reply.started":"2022-07-13T11:56:23.640707Z","shell.execute_reply":"2022-07-13T11:56:24.097959Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The Random Forest Classifier shows the best scores. We will use it for our predictions.","metadata":{}},{"cell_type":"code","source":"X_data = test_data.drop('PassengerId', axis=1).copy()\nX_train = train_data.drop('Survived', axis =1)\ny_train = train_data['Survived']","metadata":{"execution":{"iopub.status.busy":"2022-07-13T12:54:01.216648Z","iopub.execute_input":"2022-07-13T12:54:01.217078Z","iopub.status.idle":"2022-07-13T12:54:01.225154Z","shell.execute_reply.started":"2022-07-13T12:54:01.217043Z","shell.execute_reply":"2022-07-13T12:54:01.224074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"random_forest = RandomForestClassifier(n_estimators=200)\nrandom_forest.fit(X_train, y_train)\nY_pred = random_forest.predict(X_data)","metadata":{"execution":{"iopub.status.busy":"2022-07-13T12:56:22.751172Z","iopub.execute_input":"2022-07-13T12:56:22.751879Z","iopub.status.idle":"2022-07-13T12:56:23.244800Z","shell.execute_reply.started":"2022-07-13T12:56:22.751836Z","shell.execute_reply":"2022-07-13T12:56:23.243567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"PassengerId = test_data['PassengerId']","metadata":{"execution":{"iopub.status.busy":"2022-07-13T12:54:47.552891Z","iopub.execute_input":"2022-07-13T12:54:47.553384Z","iopub.status.idle":"2022-07-13T12:54:47.560598Z","shell.execute_reply.started":"2022-07-13T12:54:47.553343Z","shell.execute_reply":"2022-07-13T12:54:47.559783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame({'PassengerId': PassengerId, 'Survived':Y_pred })\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-13T12:56:27.039277Z","iopub.execute_input":"2022-07-13T12:56:27.040374Z","iopub.status.idle":"2022-07-13T12:56:27.051344Z","shell.execute_reply.started":"2022-07-13T12:56:27.040329Z","shell.execute_reply":"2022-07-13T12:56:27.050439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv(\"Titanic_predictions_Kolomijec.csv\", index = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-13T12:57:37.328108Z","iopub.execute_input":"2022-07-13T12:57:37.329412Z","iopub.status.idle":"2022-07-13T12:57:37.340080Z","shell.execute_reply.started":"2022-07-13T12:57:37.329361Z","shell.execute_reply":"2022-07-13T12:57:37.339011Z"},"trusted":true},"execution_count":null,"outputs":[]}]}