{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-31T17:53:43.062395Z","iopub.execute_input":"2022-07-31T17:53:43.063066Z","iopub.status.idle":"2022-07-31T17:53:43.094378Z","shell.execute_reply.started":"2022-07-31T17:53:43.062947Z","shell.execute_reply":"2022-07-31T17:53:43.093337Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Titanic survival prediction","metadata":{}},{"cell_type":"markdown","source":"## Problem Statement\n<br><br>\nThe sinking of the Titanic is one of the most infamous shipwrecks in history.\n\nOn April 15, 1912, during her maiden voyage, the widely considered “unsinkable” RMS Titanic sank after colliding with an iceberg. Unfortunately, there weren’t enough lifeboats for everyone onboard, resulting in the death of 1502 out of 2224 passengers and crew.\n\nWhile there was some element of luck involved in surviving, it seems some groups of people were more likely to survive than others.\n\n**what sorts of people were more likely to survive ?**","metadata":{"execution":{"iopub.status.busy":"2022-07-31T17:54:39.518427Z","iopub.execute_input":"2022-07-31T17:54:39.518836Z","iopub.status.idle":"2022-07-31T17:54:39.527486Z","shell.execute_reply.started":"2022-07-31T17:54:39.518803Z","shell.execute_reply":"2022-07-31T17:54:39.526389Z"}}},{"cell_type":"markdown","source":"### Model Objectve<br>\nPredict the Titanic Passenger survival chances by developing machine learning model ","metadata":{}},{"cell_type":"markdown","source":"## About Data\n<br><br>\n\n|Id|variable|Definition|Key|\n|:--|:--|:--|:--|\n|01|survival|Survival|0=No, 1=Yes| \n|02| pclass | Ticket class||\n|03| sex | Sex||\n|04| Age| Age in years||   \n|05| sibsp| # of siblings / spouses aboard the Titanic||\n|06| parch| # of parents / children aboard the Titanic||\n|07| ticket| Ticket number||\n|08| fare| Passenger fare||\n|09| cabin | Cabin number||\n|10| embarked| Port of Embarkation|C = Cherbourg, Q = Queenstown, S = Southampton|","metadata":{}},{"cell_type":"markdown","source":"# Contents\n<br><br>\n\n- [Import Libraries](#step_one)\n- [Collect and understand the data](#step_two)\n- [Data Preprocessing](#step_three)\n- [Exploratory Data Analysis](#step_four)\n    - [Fill missing value (Age)](#mid_step_four_one)\n    - [Fill missing values from Embarked](#mid_step_four_two)\n- [Feature Engineering](#step_five)\n- [Feature Correlation](#step_six)\n- [Post processing and Data Preparation for prediction](#step_seven)\n- [Prediction](#step_eight)\n    - [1. Logistic Regression](#step_eight_one)\n    - [2 Decision Tree](#step_eight_two)\n    - [3.Random Forest](#step_eight_three)\n- [Classification Report](#step_nine)\n- [grid search CV](#step_ten)\n- [Save model](#step_eleven)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T17:57:05.002301Z","iopub.execute_input":"2022-07-31T17:57:05.002752Z","iopub.status.idle":"2022-07-31T17:57:05.011837Z","shell.execute_reply.started":"2022-07-31T17:57:05.002717Z","shell.execute_reply":"2022-07-31T17:57:05.010688Z"}}},{"cell_type":"markdown","source":"<a id='step_one'></a>\n##  1.   Import Libraries","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.graph_objs as go\nimport plotly.express as px\nimport plotly.figure_factory as ff\nfrom plotly import tools\nfrom plotly.subplots import make_subplots","metadata":{"execution":{"iopub.status.busy":"2022-07-31T17:57:56.801405Z","iopub.execute_input":"2022-07-31T17:57:56.801822Z","iopub.status.idle":"2022-07-31T17:57:59.497446Z","shell.execute_reply.started":"2022-07-31T17:57:56.801788Z","shell.execute_reply":"2022-07-31T17:57:59.496566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.set_option('display.max_columns', None)                          # Unfolding hidden features if the cardinality is high\npd.set_option('display.max_colwidth', None)                         # Unfolding the max feature width for better clearity\npd.set_option('display.max_rows', None)                             # Unfolding hidden data points if the cardinality is high\npd.set_option('mode.chained_assignment', None)                      # Removing restriction over chained assignments operations\nimport warnings                                                     # Importing warning to disable runtime warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"execution":{"iopub.status.busy":"2022-07-31T17:58:22.873856Z","iopub.execute_input":"2022-07-31T17:58:22.874273Z","iopub.status.idle":"2022-07-31T17:58:22.881256Z","shell.execute_reply.started":"2022-07-31T17:58:22.874240Z","shell.execute_reply":"2022-07-31T17:58:22.880023Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn import preprocessing\nfrom sklearn.preprocessing import StandardScaler","metadata":{"execution":{"iopub.status.busy":"2022-07-31T17:58:51.623912Z","iopub.execute_input":"2022-07-31T17:58:51.624328Z","iopub.status.idle":"2022-07-31T17:58:51.689151Z","shell.execute_reply.started":"2022-07-31T17:58:51.624293Z","shell.execute_reply":"2022-07-31T17:58:51.688276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='step_two'></a>\n##  2.   Collect and understand the data","metadata":{}},{"cell_type":"code","source":"data = pd.read_csv('../input/titanic/train.csv')","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:00:10.007942Z","iopub.execute_input":"2022-07-31T18:00:10.008317Z","iopub.status.idle":"2022-07-31T18:00:10.027104Z","shell.execute_reply.started":"2022-07-31T18:00:10.008287Z","shell.execute_reply":"2022-07-31T18:00:10.025949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.head()   #to view the data","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:00:30.679808Z","iopub.execute_input":"2022-07-31T18:00:30.680196Z","iopub.status.idle":"2022-07-31T18:00:30.704462Z","shell.execute_reply.started":"2022-07-31T18:00:30.680158Z","shell.execute_reply":"2022-07-31T18:00:30.703685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"Total rows/entries in data : {data.shape[0]}\")\nprint(\"Total columns/features in data : {}\".format(data.shape[1]))\ndata.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:00:51.185878Z","iopub.execute_input":"2022-07-31T18:00:51.186738Z","iopub.status.idle":"2022-07-31T18:00:51.195837Z","shell.execute_reply.started":"2022-07-31T18:00:51.186687Z","shell.execute_reply":"2022-07-31T18:00:51.194936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Total **12 columns**","metadata":{}},{"cell_type":"code","source":"data.columns","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:01:18.291679Z","iopub.execute_input":"2022-07-31T18:01:18.292372Z","iopub.status.idle":"2022-07-31T18:01:18.301520Z","shell.execute_reply.started":"2022-07-31T18:01:18.292337Z","shell.execute_reply":"2022-07-31T18:01:18.299456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='step_three'></a>\n##  3.   Data Preprocessing","metadata":{}},{"cell_type":"markdown","source":"#### Null entries in dataset","metadata":{}},{"cell_type":"code","source":"data.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:02:57.144402Z","iopub.execute_input":"2022-07-31T18:02:57.144815Z","iopub.status.idle":"2022-07-31T18:02:57.158562Z","shell.execute_reply.started":"2022-07-31T18:02:57.144783Z","shell.execute_reply":"2022-07-31T18:02:57.157373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"In Percentage\\n\\n\")\nprint((data.isnull().sum()/data.shape[0])*100)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:03:04.615850Z","iopub.execute_input":"2022-07-31T18:03:04.616225Z","iopub.status.idle":"2022-07-31T18:03:04.627366Z","shell.execute_reply.started":"2022-07-31T18:03:04.616195Z","shell.execute_reply":"2022-07-31T18:03:04.626170Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Age, Cabin and Embarked** has missing values","metadata":{}},{"cell_type":"code","source":"data.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:03:19.023852Z","iopub.execute_input":"2022-07-31T18:03:19.024997Z","iopub.status.idle":"2022-07-31T18:03:19.050476Z","shell.execute_reply.started":"2022-07-31T18:03:19.024948Z","shell.execute_reply":"2022-07-31T18:03:19.049370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:03:27.521248Z","iopub.execute_input":"2022-07-31T18:03:27.522223Z","iopub.status.idle":"2022-07-31T18:03:27.556246Z","shell.execute_reply.started":"2022-07-31T18:03:27.522183Z","shell.execute_reply":"2022-07-31T18:03:27.555181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p>Primary Observations<br>\n    Age :  maximum = 80   Minimum = 0.42  Average = 29.699<br>\n    Fare : Maximum = 512.329  Minimum = 0.0 Average = 32.2042\n    <p>\n","metadata":{}},{"cell_type":"code","source":"def unique_values(data):\n    '''This function will print the nunique values for all columns present in the dataset\n    eg. dataset_name is the target dataset. Pass the data to the function\n    \n    unique_values(dataset_name)\n    \n    It will print the result\n    '''\n    for column in data.columns:\n        print('Unique Values in feature',column)\n        print(data[column].nunique())\n        if data[column].nunique()<10:\n            print('\\nValues are -> ',data[column].unique())\n        print('-'*50)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:03:46.281879Z","iopub.execute_input":"2022-07-31T18:03:46.282295Z","iopub.status.idle":"2022-07-31T18:03:46.289433Z","shell.execute_reply.started":"2022-07-31T18:03:46.282262Z","shell.execute_reply":"2022-07-31T18:03:46.288262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"unique_values(data)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:03:52.250321Z","iopub.execute_input":"2022-07-31T18:03:52.250916Z","iopub.status.idle":"2022-07-31T18:03:52.273408Z","shell.execute_reply.started":"2022-07-31T18:03:52.250867Z","shell.execute_reply":"2022-07-31T18:03:52.272127Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Types of Features<br><br>\n**Categorical Features** : sex,Embarked<br>\n**Ordinal Features** : PClass,SibSp,Parch<br>\n**Continuous Features** : Age,Fare","metadata":{}},{"cell_type":"markdown","source":"<a id='step_four'></a>\n##  4.   EDA (Exploratory Data Analysis)","metadata":{}},{"cell_type":"markdown","source":"**Q How many Survived**","metadata":{}},{"cell_type":"code","source":"colors = sns.color_palette('pastel')\ncolors","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:05:03.875090Z","iopub.execute_input":"2022-07-31T18:05:03.875528Z","iopub.status.idle":"2022-07-31T18:05:03.884786Z","shell.execute_reply.started":"2022-07-31T18:05:03.875492Z","shell.execute_reply":"2022-07-31T18:05:03.883465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels= ['Not Survived','Survived']\ncolors = sns.color_palette('pastel')[3:1:-1]   # Some coding to get red and green colobrs\nfig,ax = plt.subplots(1,2,figsize=(15,5))\nb = ax[0].pie(data['Survived'].value_counts(),labels=labels,colors=colors,autopct='%.1f%%')\nax[0].set_title('Pie Chart Survived Vs Not Survived')\n\na = ax[1].bar(labels,data['Survived'].value_counts(),color=colors)\nax[1].spines.right.set_visible(False)\nax[1].spines.top.set_visible(False)\nax[1].bar_label(a,fmt='%0.2f',fontsize=10)\nax[1].set_title('Bar Chart Survived Vs Not Survived')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:05:13.375950Z","iopub.execute_input":"2022-07-31T18:05:13.376368Z","iopub.status.idle":"2022-07-31T18:05:13.826744Z","shell.execute_reply.started":"2022-07-31T18:05:13.376333Z","shell.execute_reply":"2022-07-31T18:05:13.825790Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*plotly code*","metadata":{}},{"cell_type":"code","source":"# fig = make_subplots(rows=1,cols=2,specs=[[{\"type\": \"pie\"}, {\"type\": \"bar\"}]],column_titles=['Donut Chart','Bar Chart'])\n# trace1 = go.Pie(labels = ['Not Survived','Survived'],\n#                 values=data.Survived.value_counts(),\n#                 textinfo='label+percent',\n#                 hole=0.4,\n#                 marker=dict(colors=['red','green'])\n#                )\n# trace2 = go.Bar(x=['Not Survived','Survived'],\n#                 y=data.Survived.value_counts().values,\n#                 text=data.Survived.value_counts().values,\n#                 marker_color=['red','green'],\n#                 showlegend=False)\n# fig.append_trace(trace1,1,1)\n# fig.append_trace(trace2,1,2)\n# fig.update_layout(title={'text':'Survived (Passengers) vs Not (survived Passengers)'})\n# fig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:05:35.196988Z","iopub.execute_input":"2022-07-31T18:05:35.197375Z","iopub.status.idle":"2022-07-31T18:05:35.202822Z","shell.execute_reply.started":"2022-07-31T18:05:35.197344Z","shell.execute_reply":"2022-07-31T18:05:35.201702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Survived Passengers : 342 (38.4%)<br>\nNot Survived passengers : 549 (61.6%)<br>\nTotal Passengers : 891<br>\nIt is evident that not many passengers survived the accident.","metadata":{}},{"cell_type":"markdown","source":"### Feature Analysis : Sex","metadata":{}},{"cell_type":"code","source":"data.groupby(by=['Sex','Survived'])['Survived'].count()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:06:03.407325Z","iopub.execute_input":"2022-07-31T18:06:03.407760Z","iopub.status.idle":"2022-07-31T18:06:03.420155Z","shell.execute_reply.started":"2022-07-31T18:06:03.407725Z","shell.execute_reply":"2022-07-31T18:06:03.419262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data[data['Sex']=='female']['Survived'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:06:14.897563Z","iopub.execute_input":"2022-07-31T18:06:14.897947Z","iopub.status.idle":"2022-07-31T18:06:14.911325Z","shell.execute_reply.started":"2022-07-31T18:06:14.897918Z","shell.execute_reply":"2022-07-31T18:06:14.910335Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels= ['Survived','Not Survived']\ncolors = sns.color_palette('pastel')[2:4]\nfig,ax = plt.subplots(1,2,figsize=(15,5))\n\na = ax[0].pie(data[data['Sex']=='female']['Survived'].value_counts()/data.shape[0]*100,\n              labels=labels,\n              colors=colors,\n              autopct='%.1f%%')\nax[0].set_title('Pie Chart Female Passengers Survived Vs Not Survived percentage')\n\nb = ax[1].bar(labels,data[data['Sex']=='female']['Survived'].value_counts(),color=colors)\nax[1].set_title('Female Passengers survived')\nax[1].spines.right.set_visible(False)\nax[1].spines.top.set_visible(False)\nax[1].bar_label(b,fmt='%0.2f',fontsize=10)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:06:22.099670Z","iopub.execute_input":"2022-07-31T18:06:22.100017Z","iopub.status.idle":"2022-07-31T18:06:22.339802Z","shell.execute_reply.started":"2022-07-31T18:06:22.099989Z","shell.execute_reply":"2022-07-31T18:06:22.338600Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*same code for male just change* data[data['Sex']=='male']['Survived']","metadata":{}},{"cell_type":"code","source":"labels= ['Survived','Not Survived']\ncolors = sns.color_palette('pastel')[2:4]\nfig,ax = plt.subplots(1,2,figsize=(15,5))\n\na = ax[0].pie(data[data['Sex']=='male']['Survived'].value_counts()[::-1]/data.shape[0]*100, \n              labels=labels,\n              colors=colors,\n              autopct='%.1f%%')\nax[0].set_title('Pie Chart Female Passengers Survived Vs Not Survived percentage')\n\nb = ax[1].bar(labels,data[data['Sex']=='male']['Survived'].value_counts()[::-1],color=colors)\nax[1].set_title('Female Passengers survived')\nax[1].spines.right.set_visible(False)\nax[1].spines.top.set_visible(False)\nax[1].bar_label(b,fmt='%0.2f',fontsize=10)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:06:40.888271Z","iopub.execute_input":"2022-07-31T18:06:40.889388Z","iopub.status.idle":"2022-07-31T18:06:41.136342Z","shell.execute_reply.started":"2022-07-31T18:06:40.889346Z","shell.execute_reply":"2022-07-31T18:06:41.135200Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*plotly code*","metadata":{}},{"cell_type":"code","source":"# fig = make_subplots(rows=1,cols=2,column_titles=['Percentage Survived vs Sex','Sex : Survived vs Dead'])\n# fig.add_trace(\n#     go.Bar(x=['Female','Male'],\n#            y = data[['Sex','Survived']].groupby(['Sex']).mean().values.flatten()*100,\n#            text=['Female','Male'],\n#            marker=dict(color=['rgb(232,123,123)','rgb(157, 245, 173)']),\n#            showlegend=False),\n#     row=1,\n#     col=1\n#     )\n# fig.add_trace(\n#     go.Bar(x=data[data.Survived==0]['Sex'].value_counts().index,\n#            y=data[data.Survived==0]['Sex'].value_counts(),name='Not Survived',\n#            text=data[data.Survived==0]['Sex'].value_counts()),\n#            row=1,\n#            col=2)\n# fig.add_trace(go.Bar(x=data[data.Survived==1]['Sex'].value_counts().index,\n#                      y=data[data.Survived==1]['Sex'].value_counts(),name='Survived',\n#                     text=data[data.Survived==1]['Sex'].value_counts()),\n#               row=1,\n#               col=2)\n# fig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:06:58.856311Z","iopub.execute_input":"2022-07-31T18:06:58.857002Z","iopub.status.idle":"2022-07-31T18:06:58.862684Z","shell.execute_reply.started":"2022-07-31T18:06:58.856959Z","shell.execute_reply":"2022-07-31T18:06:58.861590Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Even though number of male passenger is high, the number of female passenger survived is higher than twice the number of male passengers survived<br>\nFemale Survival Rate -> **74.2%**<br>\nMale Survival Rate -> **18.9%**<br><br>","metadata":{}},{"cell_type":"markdown","source":"### Feature Analysis : Pclass","metadata":{}},{"cell_type":"code","source":"color = sns.color_palette('colorblind')\ncolor","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:07:31.805396Z","iopub.execute_input":"2022-07-31T18:07:31.806000Z","iopub.status.idle":"2022-07-31T18:07:31.815635Z","shell.execute_reply.started":"2022-07-31T18:07:31.805938Z","shell.execute_reply":"2022-07-31T18:07:31.814499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.crosstab(data.Pclass,data.Survived)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:07:38.508219Z","iopub.execute_input":"2022-07-31T18:07:38.508649Z","iopub.status.idle":"2022-07-31T18:07:38.535896Z","shell.execute_reply.started":"2022-07-31T18:07:38.508612Z","shell.execute_reply":"2022-07-31T18:07:38.534767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data['Pclass'].value_counts().sort_index().index.to_list()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:07:45.666288Z","iopub.execute_input":"2022-07-31T18:07:45.666704Z","iopub.status.idle":"2022-07-31T18:07:45.676180Z","shell.execute_reply.started":"2022-07-31T18:07:45.666670Z","shell.execute_reply":"2022-07-31T18:07:45.675125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig,ax = plt.subplots(1,2,figsize=(15,5))\ncolor = sns.color_palette('bright')\nlabels = [\"1\",\"2\",\"3\"]\na = ax[0].bar(labels,\n             data['Pclass'].value_counts().sort_index())\nax[0].set_title('number of passengers per class')\nax[0].spines.top.set_visible(False)\nax[0].spines.right.set_visible(False)\nax[0].bar_label(a,fmt='%.0f',fontsize=10)\n\nwidth = 0.4\nx = np.arange(3)\nb = ax[1].bar(x-0.2,\n             pd.crosstab(data.Pclass,data.Survived)[0],width,color=color[3],label='Not Survived')\nc = ax[1].bar(x+0.2,\n             pd.crosstab(data.Pclass,data.Survived)[1],width,color=color[2],label='Survived')\nax[1].set_title('number of passengers per class')\nax[1].spines.top.set_visible(False)\nax[1].spines.right.set_visible(False)\nax[1].bar_label(b,fmt='%.0f',fontsize=10)\nax[1].bar_label(c,fmt='%.0f',fontsize=10)\nax[1].legend()\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:07:53.014456Z","iopub.execute_input":"2022-07-31T18:07:53.014957Z","iopub.status.idle":"2022-07-31T18:07:53.386771Z","shell.execute_reply.started":"2022-07-31T18:07:53.014922Z","shell.execute_reply":"2022-07-31T18:07:53.385698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*plotly code*","metadata":{}},{"cell_type":"code","source":"# fig = make_subplots(rows=1,cols=2,column_titles=['Number of Passengers per class','Pclass : Survived vs Dead'])\n# fig.add_trace(\n#     go.Bar(x=['3','1','2'],\n#            y = data[['Pclass','Survived']].groupby(['Pclass']).count().sort_values(by='Survived',ascending=False).values.flatten(),\n#            text=data[['Pclass','Survived']].groupby(['Pclass']).count().sort_values(by='Survived',ascending=False).values.flatten(),\n#            showlegend=False),\n#     row=1,\n#     col=1\n#     )\n# fig.add_trace(\n#     go.Bar(x=data[data.Survived==0]['Pclass'].value_counts().index,\n#            y=data[data.Survived==0]['Pclass'].value_counts(),name='Not Survived',\n#            text=data[data.Survived==0]['Pclass'].value_counts()\n#           ),\n#            row=1,\n#            col=2)\n# fig.add_trace(go.Bar(x=data[data.Survived==1]['Pclass'].value_counts().index,\n#                      y=data[data.Survived==1]['Pclass'].value_counts(),name='Survived',\n#                     text=data[data.Survived==1]['Pclass'].value_counts()),\n#               row=1,\n#               col=2)\n# fig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:08:10.637767Z","iopub.execute_input":"2022-07-31T18:08:10.638151Z","iopub.status.idle":"2022-07-31T18:08:10.643250Z","shell.execute_reply.started":"2022-07-31T18:08:10.638119Z","shell.execute_reply":"2022-07-31T18:08:10.642467Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can observe that number of passengers in class 3 are highest. But very less passengers survived in class 3<br>\nAnd from class 1, highest number of passengers survived ","metadata":{}},{"cell_type":"markdown","source":"### Survival rate for sex and Pclass","metadata":{}},{"cell_type":"code","source":"pd.crosstab([data.Sex,data.Survived],data.Pclass,margins=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:08:40.141257Z","iopub.execute_input":"2022-07-31T18:08:40.141863Z","iopub.status.idle":"2022-07-31T18:08:40.188169Z","shell.execute_reply.started":"2022-07-31T18:08:40.141826Z","shell.execute_reply":"2022-07-31T18:08:40.186949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.factorplot('Pclass','Survived',hue='Sex',data=data)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:08:48.231846Z","iopub.execute_input":"2022-07-31T18:08:48.232255Z","iopub.status.idle":"2022-07-31T18:08:48.790512Z","shell.execute_reply.started":"2022-07-31T18:08:48.232224Z","shell.execute_reply":"2022-07-31T18:08:48.789665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can clearly see that survival rate of female is higher than male and almost all female from class 1 survived.<br>\nFor male, class 1 male have high survival rate","metadata":{}},{"cell_type":"markdown","source":"### Feature Anaysis : Age","metadata":{}},{"cell_type":"markdown","source":"As we know Age has missing values. So we'll fill the missing values by understanding the Age More","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=[15,7])\nsns.set_theme(style=\"whitegrid\")\nsns.violinplot(data=data, x='Pclass',y=\"Age\",hue='Survived',split=True,inner=\"quart\",linewidth=1)\nsns.despine(left=True)\n\n# reference link : https://seaborn.pydata.org/examples/grouped_violinplots.html\n# followed the link and just updated the data as per project requirement :)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:09:18.704394Z","iopub.execute_input":"2022-07-31T18:09:18.704793Z","iopub.status.idle":"2022-07-31T18:09:18.996081Z","shell.execute_reply.started":"2022-07-31T18:09:18.704762Z","shell.execute_reply":"2022-07-31T18:09:18.995063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*plotly code*","metadata":{}},{"cell_type":"code","source":"# fig = go.Figure(data=[go.Violin(x = data['Pclass'][data['Survived']==0],\n#                                y = data['Age'][data['Survived']==0],\n#                                 box_visible=True,\n#                                 name='Not Survived',\n#                                 side='negative',\n#                                 marker=dict(color='red')\n#                                ),\n#                       go.Violin(x = data['Pclass'][data['Survived']==1],\n#                                y = data['Age'][data['Survived']==1],\n#                                 box_visible=True,\n#                                 name='Survived',\n#                                 side='positive',\n#                                 marker=dict(color='blue')\n#                                )\n#                      ]\n#                )\n# fig.update_layout(violinmode='overlay',\n#                  title_text='Pclass and Age vs Survived')\n# fig.update_xaxes(title_text='Pclass')\n# fig.update_yaxes(title_text='Age')\n# fig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:09:37.109090Z","iopub.execute_input":"2022-07-31T18:09:37.109954Z","iopub.status.idle":"2022-07-31T18:09:37.115170Z","shell.execute_reply.started":"2022-07-31T18:09:37.109913Z","shell.execute_reply":"2022-07-31T18:09:37.114093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can observe that, as class increases the number of children and their survival rate is also increasing<br>\nSurvival rate is decreasing as the age goes increasing","metadata":{}},{"cell_type":"markdown","source":"Since there are a lot of passengers with different age, we don't want to fill missing age with mean or meadian to all missing values<br>If we do that, then there is a high chance that we might assign mean/median age to 4 year old or 60 year old","metadata":{}},{"cell_type":"markdown","source":"<a id=mid_step_four_one></a>\n## 4.1 Fill missing value (Age)","metadata":{}},{"cell_type":"markdown","source":"We will use Name feature to understand if we can get any idea about age ","metadata":{}},{"cell_type":"code","source":"data.Name.values[:20]","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:10:17.958044Z","iopub.execute_input":"2022-07-31T18:10:17.958393Z","iopub.status.idle":"2022-07-31T18:10:17.965317Z","shell.execute_reply.started":"2022-07-31T18:10:17.958364Z","shell.execute_reply":"2022-07-31T18:10:17.964599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are salutations to the name like Mr. Mrs. Miss. We'll extract them from Name","metadata":{}},{"cell_type":"code","source":"a=[]\nimport re\nfor i in data:\n    data['Initial'] = data.Name.str.extract('([A-Za-z]+)\\.')","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:10:35.185487Z","iopub.execute_input":"2022-07-31T18:10:35.185915Z","iopub.status.idle":"2022-07-31T18:10:35.242259Z","shell.execute_reply.started":"2022-07-31T18:10:35.185881Z","shell.execute_reply":"2022-07-31T18:10:35.241470Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:10:40.984371Z","iopub.execute_input":"2022-07-31T18:10:40.984778Z","iopub.status.idle":"2022-07-31T18:10:41.001855Z","shell.execute_reply.started":"2022-07-31T18:10:40.984744Z","shell.execute_reply":"2022-07-31T18:10:41.001007Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.Initial.value_counts().index","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:10:47.825682Z","iopub.execute_input":"2022-07-31T18:10:47.826075Z","iopub.status.idle":"2022-07-31T18:10:47.834739Z","shell.execute_reply.started":"2022-07-31T18:10:47.826041Z","shell.execute_reply":"2022-07-31T18:10:47.833921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have Mr,Miss,Mrs,Master in our data. We'll keep them and rest of the salutations we'll replace with apropriate salutation<br>Mr -> Dr, Major, Don, Sir, Capt<br>\nMiss -> Mlle, Mme, Ms<br>\nMrs -> Lady, Countess<br>\nOther -> Jonkheer, Col, Rev","metadata":{}},{"cell_type":"code","source":"data['Initial'].replace(['Mlle','Mme' ,'Ms'  ,'Dr','Major','Lady','Countess','Jonkheer','Col'  ,'Rev'  ,'Capt','Sir','Don','Dona'],\n                        ['Miss','Miss','Miss','Mr','Mr'   ,'Mrs'  ,'Mrs'    ,'Other'   ,'Other','Other','Mr'  ,'Mr' ,'Mr','Mrs'],\n                       inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:11:01.084125Z","iopub.execute_input":"2022-07-31T18:11:01.084559Z","iopub.status.idle":"2022-07-31T18:11:01.093771Z","shell.execute_reply.started":"2022-07-31T18:11:01.084522Z","shell.execute_reply":"2022-07-31T18:11:01.092908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.groupby('Initial')['Age'].mean()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:11:08.947689Z","iopub.execute_input":"2022-07-31T18:11:08.948081Z","iopub.status.idle":"2022-07-31T18:11:08.959282Z","shell.execute_reply.started":"2022-07-31T18:11:08.948048Z","shell.execute_reply":"2022-07-31T18:11:08.958385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we have the better Idea about the Age. Lets fill missing values with mean for appropriate salutations","metadata":{}},{"cell_type":"code","source":"data.loc[(data.Age.isnull())&(data.Initial=='Mr'),'Age']=32.73\ndata.loc[(data.Age.isnull())&(data.Initial=='Mrs'),'Age']=36\ndata.loc[(data.Age.isnull())&(data.Initial=='Master'),'Age']=4.5\ndata.loc[(data.Age.isnull())&(data.Initial=='Miss'),'Age']=21.8\ndata.loc[(data.Age.isnull())&(data.Initial=='Other'),'Age']=45.8","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:11:26.635928Z","iopub.execute_input":"2022-07-31T18:11:26.636327Z","iopub.status.idle":"2022-07-31T18:11:26.650393Z","shell.execute_reply.started":"2022-07-31T18:11:26.636296Z","shell.execute_reply":"2022-07-31T18:11:26.649564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.Age.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:11:31.855776Z","iopub.execute_input":"2022-07-31T18:11:31.856163Z","iopub.status.idle":"2022-07-31T18:11:31.864294Z","shell.execute_reply.started":"2022-07-31T18:11:31.856133Z","shell.execute_reply":"2022-07-31T18:11:31.862967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.factorplot('Pclass','Survived',hue='Initial',data=data)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:11:37.433433Z","iopub.execute_input":"2022-07-31T18:11:37.433801Z","iopub.status.idle":"2022-07-31T18:11:38.452662Z","shell.execute_reply.started":"2022-07-31T18:11:37.433772Z","shell.execute_reply":"2022-07-31T18:11:38.451457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we can observe that, Mrs, Miss and Master from class 1 and 2, almost all are survived.","metadata":{}},{"cell_type":"markdown","source":"### Feature Analysis : Embarked","metadata":{}},{"cell_type":"code","source":"pd.crosstab([data.Embarked,data.Pclass],[data.Sex,data.Survived],margins=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:12:02.252156Z","iopub.execute_input":"2022-07-31T18:12:02.252535Z","iopub.status.idle":"2022-07-31T18:12:02.305157Z","shell.execute_reply.started":"2022-07-31T18:12:02.252506Z","shell.execute_reply":"2022-07-31T18:12:02.304400Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.factorplot('Pclass','Survived',hue='Sex',col='Embarked',data=data)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:12:09.118974Z","iopub.execute_input":"2022-07-31T18:12:09.119690Z","iopub.status.idle":"2022-07-31T18:12:10.514234Z","shell.execute_reply.started":"2022-07-31T18:12:09.119646Z","shell.execute_reply":"2022-07-31T18:12:10.513169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.factorplot('Embarked','Survived',data=data)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:12:17.930458Z","iopub.execute_input":"2022-07-31T18:12:17.930835Z","iopub.status.idle":"2022-07-31T18:12:18.373051Z","shell.execute_reply.started":"2022-07-31T18:12:17.930804Z","shell.execute_reply":"2022-07-31T18:12:18.371792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can observe that,passengers C Embarked have high survival rate ","metadata":{}},{"cell_type":"code","source":"data.groupby(by='Embarked')['Survived'].count()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:12:34.464663Z","iopub.execute_input":"2022-07-31T18:12:34.465069Z","iopub.status.idle":"2022-07-31T18:12:34.475123Z","shell.execute_reply.started":"2022-07-31T18:12:34.465039Z","shell.execute_reply":"2022-07-31T18:12:34.474090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=[10,5])\na = plt.bar(data.groupby(by='Embarked')['Survived'].count().sort_values(ascending=False).index,\n        data.groupby(by='Embarked')['Survived'].count().sort_values(ascending=False))\nplt.bar_label(a,fmt='%.0f')\nplt.title(\"Passengers per Embarked\",fontsize=16)\nplt.xlabel('Embarked')\nplt.ylabel('Passengers')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:12:41.026140Z","iopub.execute_input":"2022-07-31T18:12:41.026547Z","iopub.status.idle":"2022-07-31T18:12:41.182417Z","shell.execute_reply.started":"2022-07-31T18:12:41.026513Z","shell.execute_reply":"2022-07-31T18:12:41.181410Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*plotly code*","metadata":{}},{"cell_type":"code","source":"# fig = go.Figure(data=(go.Bar(x=['S','C','Q'],\n#                              y=data.groupby(by='Embarked')['Survived'].count().sort_values(ascending=False).values,\n#                             text=data.groupby(by='Embarked')['Survived'].count().sort_values(ascending=False).values)))\n# fig.update_layout(title_text=\"No. of Passengers Onboarded\")\n# fig.update_xaxes(title_text='Embarked')\n# fig.update_yaxes(title_text='Passenger Count')\n# fig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:12:54.361364Z","iopub.execute_input":"2022-07-31T18:12:54.361784Z","iopub.status.idle":"2022-07-31T18:12:54.365781Z","shell.execute_reply.started":"2022-07-31T18:12:54.361750Z","shell.execute_reply":"2022-07-31T18:12:54.365039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Observation : Most of the passengers onboarded from S port","metadata":{}},{"cell_type":"markdown","source":"<a id=\"mid_step_four_two\"></a>\n## 4.2 Fill missing values from Embarked","metadata":{}},{"cell_type":"code","source":"data['Embarked'].fillna('S',inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:13:19.573016Z","iopub.execute_input":"2022-07-31T18:13:19.573395Z","iopub.status.idle":"2022-07-31T18:13:19.578585Z","shell.execute_reply.started":"2022-07-31T18:13:19.573364Z","shell.execute_reply":"2022-07-31T18:13:19.577694Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.Embarked.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:13:24.284870Z","iopub.execute_input":"2022-07-31T18:13:24.286087Z","iopub.status.idle":"2022-07-31T18:13:24.293841Z","shell.execute_reply.started":"2022-07-31T18:13:24.286039Z","shell.execute_reply":"2022-07-31T18:13:24.293026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Feature Analysis : SibSip","metadata":{}},{"cell_type":"code","source":"pd.crosstab(data.SibSp,data.Survived,margins=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:13:39.421649Z","iopub.execute_input":"2022-07-31T18:13:39.422015Z","iopub.status.idle":"2022-07-31T18:13:39.466012Z","shell.execute_reply.started":"2022-07-31T18:13:39.421987Z","shell.execute_reply":"2022-07-31T18:13:39.464841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"label = [str(i) for i in pd.crosstab(data.SibSp,data.Survived).index]\nlabel","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:13:46.168310Z","iopub.execute_input":"2022-07-31T18:13:46.168711Z","iopub.status.idle":"2022-07-31T18:13:46.188157Z","shell.execute_reply.started":"2022-07-31T18:13:46.168680Z","shell.execute_reply":"2022-07-31T18:13:46.186904Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels = [str(i) for i in pd.crosstab(data.SibSp,data.Survived).index]\nNot_survived = pd.crosstab(data.SibSp,data.Survived)[0].values\nsurvived = pd.crosstab(data.SibSp,data.Survived)[1].values\n\nx = np.arange(len(labels))  # the label locations\nwidth = 0.35  # the width of the bars\n\nfig, ax = plt.subplots(figsize=(12,8))\nrects1 = ax.bar(x - width/2, Not_survived, width, label='Not Survived',color=color[3])\nrects2 = ax.bar(x + width/2, survived, width, label='Survived',color=color[2])\n\n# Add some text for labels, title and custom x-axis tick labels, etc.\nax.set_ylabel('Passengers')\nax.set_title('Sibsp vs Survived',fontsize=16)\nax.set_xticks(x, labels)\nax.legend()\n\nax.bar_label(rects1, padding=3)\nax.bar_label(rects2, padding=3)\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:13:52.474080Z","iopub.execute_input":"2022-07-31T18:13:52.474494Z","iopub.status.idle":"2022-07-31T18:13:52.779876Z","shell.execute_reply.started":"2022-07-31T18:13:52.474456Z","shell.execute_reply":"2022-07-31T18:13:52.778795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*pllotly code*","metadata":{}},{"cell_type":"code","source":"# fig = go.Figure(data=(go.Bar(x=data[data.Survived==1]['SibSp'].value_counts().index,\n#                             y=data[data.Survived==1]['SibSp'].value_counts().values,\n#                             name='Survived'),\n#                      go.Bar(x=data[data.Survived==0]['SibSp'].value_counts().index,\n#                             y=data[data.Survived==0]['SibSp'].value_counts().values,\n#                            name='Not Survived')))\n# fig.update_layout(title_text='SibSp VS Survived')\n# fig.update_xaxes(title_text='Family Size')\n# fig.update_yaxes(title_text='Passenger Count')\n# fig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:14:18.427273Z","iopub.execute_input":"2022-07-31T18:14:18.427676Z","iopub.status.idle":"2022-07-31T18:14:18.432653Z","shell.execute_reply.started":"2022-07-31T18:14:18.427644Z","shell.execute_reply":"2022-07-31T18:14:18.431482Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Observation : If family size is more than 1, the survival chances are extremely low","metadata":{}},{"cell_type":"markdown","source":"### Feature Analysis : Parch","metadata":{}},{"cell_type":"code","source":"pd.crosstab([data.Parch,data.Survived],data.Pclass,margins=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:15:08.630739Z","iopub.execute_input":"2022-07-31T18:15:08.631146Z","iopub.status.idle":"2022-07-31T18:15:08.682544Z","shell.execute_reply.started":"2022-07-31T18:15:08.631111Z","shell.execute_reply":"2022-07-31T18:15:08.681463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that large family size is in class 3 ","metadata":{}},{"cell_type":"code","source":"sns.factorplot('Parch','Survived',hue='Embarked',col='Pclass',data=data)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:15:25.977251Z","iopub.execute_input":"2022-07-31T18:15:25.977696Z","iopub.status.idle":"2022-07-31T18:15:27.996495Z","shell.execute_reply.started":"2022-07-31T18:15:25.977655Z","shell.execute_reply":"2022-07-31T18:15:27.995434Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Observations:<br>\nPassengers with their parents onboard have greater chance of survival<br>\nFamily size more than 3 are from class 3 and their survival rate is very low<br>\nFamily from class 1 and 2 with size more than 1 and less than 4 which onboarded from port S have high survival rate<br>\nIn class 3, Family size more than 1 and less than 4 which onboarded from port C have high survival rate","metadata":{}},{"cell_type":"markdown","source":"**Feature Analysis : Fare**","metadata":{}},{"cell_type":"code","source":"lst = [data[data['Pclass']==i]['Fare'].values for i in range(1,4)]","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:15:51.505851Z","iopub.execute_input":"2022-07-31T18:15:51.506248Z","iopub.status.idle":"2022-07-31T18:15:51.515300Z","shell.execute_reply.started":"2022-07-31T18:15:51.506214Z","shell.execute_reply":"2022-07-31T18:15:51.514038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig,ax = plt.subplots(figsize=[12,8])\nax.hist(x=lst,stacked=True,bins=70)\nax.legend([1,2,3],title='Pclass')\nax.set_title('Fare per passenger class',fontsize=16)\nax.set_xlabel('Fare')\nax.set_ylabel('Passenger Counts')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:15:56.257527Z","iopub.execute_input":"2022-07-31T18:15:56.257890Z","iopub.status.idle":"2022-07-31T18:15:56.974185Z","shell.execute_reply.started":"2022-07-31T18:15:56.257860Z","shell.execute_reply":"2022-07-31T18:15:56.973133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*plotly code*","metadata":{}},{"cell_type":"code","source":"# fig = px.histogram(data,x='Fare',color='Pclass')\n# fig.update_traces(opacity=0.60)\n# fig.update_layout(\n#     title_text='Fare in class', # title of plot\n#     xaxis_title_text='Fare', # xaxis label\n#     yaxis_title_text='Count', # yaxis label\n#     bargap=0.2, # gap between bars of adjacent location coordinates\n#     bargroupgap=0.1 # gap between bars of the same location coordinates\n# )\n# fig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:16:12.207379Z","iopub.execute_input":"2022-07-31T18:16:12.208576Z","iopub.status.idle":"2022-07-31T18:16:12.213427Z","shell.execute_reply.started":"2022-07-31T18:16:12.208531Z","shell.execute_reply":"2022-07-31T18:16:12.212072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Observation:<br>\nWe can observe that class 1 fare is very high followed by class 2 and then class 3 ","metadata":{}},{"cell_type":"markdown","source":"<a id=\"step_five\"></a>\n## 5 Feature Engineering","metadata":{}},{"cell_type":"markdown","source":"**New Feature  <br>Family Size<br>Alone**","metadata":{}},{"cell_type":"code","source":"data['Family_Size'] = data['Parch']+data['SibSp']\ndata['Alone'] = 0\ndata.loc[data.Family_Size==0,'Alone']=1","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:16:49.466562Z","iopub.execute_input":"2022-07-31T18:16:49.466929Z","iopub.status.idle":"2022-07-31T18:16:49.476194Z","shell.execute_reply.started":"2022-07-31T18:16:49.466899Z","shell.execute_reply":"2022-07-31T18:16:49.474844Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"step_six\"></a>\n## 6 Feature Correlation","metadata":{}},{"cell_type":"code","source":"sns.heatmap(data.corr(),annot=True,linewidths=0.2)\nfig = plt.gcf()\nfig.set_size_inches(15,10)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:17:04.845262Z","iopub.execute_input":"2022-07-31T18:17:04.845936Z","iopub.status.idle":"2022-07-31T18:17:05.561865Z","shell.execute_reply.started":"2022-07-31T18:17:04.845892Z","shell.execute_reply":"2022-07-31T18:17:05.560607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Observation:<br><br>\nPositive Correlation : <br><br>Family_size and SibSp and Parch >0.75<br>SibSp and Parch 0.41<br>Fare and Parch 0.22<br> Fare and Survived 0.26<br><br>\nNegative Correlation : <br><br>Fare and Pclass -0.55<br>Age and Pclass<br>Age and SibSp<br>\n","metadata":{}},{"cell_type":"markdown","source":"<a id=\"step_seven\"></a>\n## 7 Post processing and Data Preparation for prediction","metadata":{}},{"cell_type":"markdown","source":"**Scale down continuous features**<br>\nAge and Fare","metadata":{}},{"cell_type":"code","source":"scaler = StandardScaler()\nscale_cols = ['Age','Fare']\nfor col in scale_cols:\n    data[col] = scaler.fit_transform(data[[col]])","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:17:39.249857Z","iopub.execute_input":"2022-07-31T18:17:39.250226Z","iopub.status.idle":"2022-07-31T18:17:39.263605Z","shell.execute_reply.started":"2022-07-31T18:17:39.250197Z","shell.execute_reply":"2022-07-31T18:17:39.262569Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Age and Fare can be converted into age and fare band (categorical features)","metadata":{}},{"cell_type":"markdown","source":"**Conversion of continuous features into catogorical fetures**<br>\nAge and Fare","metadata":{}},{"cell_type":"code","source":"# #Age coversion to Age band\n# data['Age_band']=0\n# data.loc[data['Age']<=16,'Age_band']=0\n# data.loc[(data['Age']>16)&(data['Age']<=32),'Age_band']=1\n# data.loc[(data['Age']>32)&(data['Age']<=48),'Age_band']=2\n# data.loc[(data['Age']>48)&(data['Age']<=64),'Age_band']=3\n# data.loc[data['Age']>64,'Age_band']=4\n# data.head(2)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:18:09.250951Z","iopub.execute_input":"2022-07-31T18:18:09.251329Z","iopub.status.idle":"2022-07-31T18:18:09.255715Z","shell.execute_reply.started":"2022-07-31T18:18:09.251299Z","shell.execute_reply":"2022-07-31T18:18:09.254937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Fare\n# a = pd.qcut(data['Fare'],4)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:18:14.411873Z","iopub.execute_input":"2022-07-31T18:18:14.412265Z","iopub.status.idle":"2022-07-31T18:18:14.416514Z","shell.execute_reply.started":"2022-07-31T18:18:14.412217Z","shell.execute_reply":"2022-07-31T18:18:14.415217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# a.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:18:19.330811Z","iopub.execute_input":"2022-07-31T18:18:19.331229Z","iopub.status.idle":"2022-07-31T18:18:19.336064Z","shell.execute_reply.started":"2022-07-31T18:18:19.331179Z","shell.execute_reply":"2022-07-31T18:18:19.335166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# data['Fare_cat']=0\n# data.loc[data['Fare']<=7.91,'Fare_cat']=0\n# data.loc[(data['Fare']>7.91)&(data['Fare']<=14.454),'Fare_cat']=1\n# data.loc[(data['Fare']>14.454)&(data['Fare']<=31),'Fare_cat']=2\n# data.loc[(data['Fare']>31)&(data['Fare']<=513),'Fare_cat']=3","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:18:24.380226Z","iopub.execute_input":"2022-07-31T18:18:24.380803Z","iopub.status.idle":"2022-07-31T18:18:24.383999Z","shell.execute_reply.started":"2022-07-31T18:18:24.380771Z","shell.execute_reply":"2022-07-31T18:18:24.383285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Categorical Nominal Data Encoding**<br>\nSex<br>\nEmbarked<br>\nInitial<br>","metadata":{}},{"cell_type":"code","source":"data['Sex'].replace(['male','female'],[0,1],inplace=True)\ndata['Embarked'].replace(['S','C','Q'],[0,1,2],inplace=True)\ndata['Initial'].replace(['Mr','Mrs','Miss','Master','Other'],[0,1,2,3,4],inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:18:39.724523Z","iopub.execute_input":"2022-07-31T18:18:39.724911Z","iopub.status.idle":"2022-07-31T18:18:39.738635Z","shell.execute_reply.started":"2022-07-31T18:18:39.724877Z","shell.execute_reply":"2022-07-31T18:18:39.737403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Using Label Encoder**","metadata":{}},{"cell_type":"code","source":"# le = preprocessing.LabelEncoder()\n# cols = ['Sex','Embarked','Initial']\n# for col in cols:\n#     data[col] = le.fit_transform(data[[col]])","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:19:10.000839Z","iopub.execute_input":"2022-07-31T18:19:10.001330Z","iopub.status.idle":"2022-07-31T18:19:10.006553Z","shell.execute_reply.started":"2022-07-31T18:19:10.001278Z","shell.execute_reply":"2022-07-31T18:19:10.005304Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<br><br>Now, let's see our dataset","metadata":{}},{"cell_type":"code","source":"data.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:19:24.549801Z","iopub.execute_input":"2022-07-31T18:19:24.550174Z","iopub.status.idle":"2022-07-31T18:19:24.570999Z","shell.execute_reply.started":"2022-07-31T18:19:24.550144Z","shell.execute_reply":"2022-07-31T18:19:24.569844Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## We don't need following features :\n**Passenger ID**<br>\n**Name**<br>\n**Age** *As we have created Age_band from the Age we can drop it*<br>\n**Ticket**<br>\n**Fare** *As we have created Fare_cat from the Fare we can drop it*<br>\n**Cabin** *Maximum Nan Values and many passengers have multiple cabins so it is of no use*","metadata":{}},{"cell_type":"code","source":"data.drop(['Name','Ticket','Cabin','PassengerId','SibSp','Parch'],axis=1,inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:19:45.395914Z","iopub.execute_input":"2022-07-31T18:19:45.396331Z","iopub.status.idle":"2022-07-31T18:19:45.403548Z","shell.execute_reply.started":"2022-07-31T18:19:45.396293Z","shell.execute_reply":"2022-07-31T18:19:45.402374Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:19:50.177387Z","iopub.execute_input":"2022-07-31T18:19:50.177777Z","iopub.status.idle":"2022-07-31T18:19:50.192107Z","shell.execute_reply.started":"2022-07-31T18:19:50.177749Z","shell.execute_reply":"2022-07-31T18:19:50.190749Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.heatmap(data.corr(),annot=True,linewidths=0.2)\nfig = plt.gcf()\nfig.set_size_inches(15,10)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:19:56.801295Z","iopub.execute_input":"2022-07-31T18:19:56.801685Z","iopub.status.idle":"2022-07-31T18:19:57.661635Z","shell.execute_reply.started":"2022-07-31T18:19:56.801653Z","shell.execute_reply":"2022-07-31T18:19:57.660572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Good to go","metadata":{}},{"cell_type":"markdown","source":"<a id=\"step_eight\"></a>\n## 8 Prediction","metadata":{}},{"cell_type":"markdown","source":"We will develop base model of : \n\n    1.Logistic Regression   ref = https://scikit-learn.org/stable/modules/generated/sklearn.linear_model.LogisticRegression.html\n    2.Decision Tree   ref = http://scikit-learn.org/stable/modules/generated/sklearn.tree.DecisionTreeClassifier.html\n    3.Random Forest   ref = https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestClassifier.html\n    \nand later we'll apply Grid Search CV and Random Search CV for Hyperparameter Tuning\n    ","metadata":{}},{"cell_type":"code","source":"from sklearn.linear_model import LogisticRegression\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.ensemble import RandomForestClassifier\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.model_selection import cross_val_score\n\nfrom sklearn import metrics\nfrom sklearn.metrics import classification_report,ConfusionMatrixDisplay,plot_confusion_matrix, confusion_matrix, precision_score,recall_score,accuracy_score,f1_score","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:21:01.197867Z","iopub.execute_input":"2022-07-31T18:21:01.198281Z","iopub.status.idle":"2022-07-31T18:21:01.468977Z","shell.execute_reply.started":"2022-07-31T18:21:01.198245Z","shell.execute_reply":"2022-07-31T18:21:01.467760Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"accuracy = []\nprecision = []\nrecall = []\nscore_f1 = []\nmodel_name = ['Logistic regression',\"Decision Tree\",\"Random forest\"]","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:21:05.865674Z","iopub.execute_input":"2022-07-31T18:21:05.866030Z","iopub.status.idle":"2022-07-31T18:21:05.871399Z","shell.execute_reply.started":"2022-07-31T18:21:05.866001Z","shell.execute_reply":"2022-07-31T18:21:05.870246Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Split data in dependent and independent variables <br>\nX {independent variables}<br>\ny {target/dependent variables}","metadata":{}},{"cell_type":"code","source":"X = data.iloc[:,1:].values\ny = data.iloc[:,0].values","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:21:21.336260Z","iopub.execute_input":"2022-07-31T18:21:21.336790Z","iopub.status.idle":"2022-07-31T18:21:21.344084Z","shell.execute_reply.started":"2022-07-31T18:21:21.336744Z","shell.execute_reply":"2022-07-31T18:21:21.342914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"X,y can be splitted in following way as well ","metadata":{}},{"cell_type":"code","source":"# X = data.drop('Survived',axis=1)\n# y = data['Survived']","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:21:38.050553Z","iopub.execute_input":"2022-07-31T18:21:38.050902Z","iopub.status.idle":"2022-07-31T18:21:38.054811Z","shell.execute_reply.started":"2022-07-31T18:21:38.050874Z","shell.execute_reply":"2022-07-31T18:21:38.053792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" Split the data","metadata":{}},{"cell_type":"code","source":"X_train,X_test,y_train,y_test = train_test_split(X,y,test_size=0.25,random_state=0)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:22:04.600356Z","iopub.execute_input":"2022-07-31T18:22:04.600777Z","iopub.status.idle":"2022-07-31T18:22:04.607247Z","shell.execute_reply.started":"2022-07-31T18:22:04.600746Z","shell.execute_reply":"2022-07-31T18:22:04.606302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"step_eight_one\"></a>\n## 1. Logistic Regression","metadata":{}},{"cell_type":"code","source":"logistic_base_model = LogisticRegression()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:22:18.998051Z","iopub.execute_input":"2022-07-31T18:22:18.998424Z","iopub.status.idle":"2022-07-31T18:22:19.002951Z","shell.execute_reply.started":"2022-07-31T18:22:18.998380Z","shell.execute_reply":"2022-07-31T18:22:19.002120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"logistic_base_model.fit(X_train,y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:22:24.516847Z","iopub.execute_input":"2022-07-31T18:22:24.517247Z","iopub.status.idle":"2022-07-31T18:22:24.535099Z","shell.execute_reply.started":"2022-07-31T18:22:24.517211Z","shell.execute_reply":"2022-07-31T18:22:24.534083Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Prediction and evaliation based on X_test data","metadata":{}},{"cell_type":"code","source":"y_pred_logistic_base_model = logistic_base_model.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:23:03.760254Z","iopub.execute_input":"2022-07-31T18:23:03.760626Z","iopub.status.idle":"2022-07-31T18:23:03.765065Z","shell.execute_reply.started":"2022-07-31T18:23:03.760597Z","shell.execute_reply":"2022-07-31T18:23:03.764367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cm = pd.DataFrame(confusion_matrix(y_test, y_pred_logistic_base_model))\ncm.index = ['Actual Died','Actual Survived']\ncm.columns = ['Predicted Died','Predicted Survived']\nprint(cm)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:23:08.924246Z","iopub.execute_input":"2022-07-31T18:23:08.924612Z","iopub.status.idle":"2022-07-31T18:23:08.935555Z","shell.execute_reply.started":"2022-07-31T18:23:08.924582Z","shell.execute_reply":"2022-07-31T18:23:08.934305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Using inbuilt confusion metrics plotting**","metadata":{}},{"cell_type":"code","source":"sns.set_theme(style=\"white\")\nplot_confusion_matrix(logistic_base_model,\n                     X_test,\n                     y_test)\n                     #display_labels = ['Survived (positive)','Not Survived (negative)'])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:23:24.542446Z","iopub.execute_input":"2022-07-31T18:23:24.542840Z","iopub.status.idle":"2022-07-31T18:23:24.703269Z","shell.execute_reply.started":"2022-07-31T18:23:24.542808Z","shell.execute_reply":"2022-07-31T18:23:24.702494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(' Accuracy : %.3f'%metrics.accuracy_score(y_pred_logistic_base_model,y_test))\nprint('Precision : %.3f' % precision_score(y_test, y_pred_logistic_base_model))\nprint('   Recall : %.3f' % recall_score(y_test, y_pred_logistic_base_model))\nprint(' F1 Score : %.3f' % f1_score(y_test, y_pred_logistic_base_model))\n\naccuracy.append(metrics.accuracy_score(y_pred_logistic_base_model,y_test))\nprecision.append(precision_score(y_test, y_pred_logistic_base_model))\nrecall.append(recall_score(y_test, y_pred_logistic_base_model))\nscore_f1.append(f1_score(y_test, y_pred_logistic_base_model))","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:23:35.044871Z","iopub.execute_input":"2022-07-31T18:23:35.045223Z","iopub.status.idle":"2022-07-31T18:23:35.061083Z","shell.execute_reply.started":"2022-07-31T18:23:35.045194Z","shell.execute_reply":"2022-07-31T18:23:35.059926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"logistic_base_model = LogisticRegression()\nscores = cross_val_score(logistic_base_model, X_train, y_train, cv=10)\nprint(\"%0.2f accuracy with a standard deviation of %0.2f\" % (scores.mean(), scores.std()))","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:23:40.711968Z","iopub.execute_input":"2022-07-31T18:23:40.712373Z","iopub.status.idle":"2022-07-31T18:23:40.808909Z","shell.execute_reply.started":"2022-07-31T18:23:40.712337Z","shell.execute_reply":"2022-07-31T18:23:40.807798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"step_eight_two\"></a>\n## 2 Decision Tree","metadata":{}},{"cell_type":"code","source":"decision_tree_base_model = DecisionTreeClassifier()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:24:01.700393Z","iopub.execute_input":"2022-07-31T18:24:01.700820Z","iopub.status.idle":"2022-07-31T18:24:01.706206Z","shell.execute_reply.started":"2022-07-31T18:24:01.700784Z","shell.execute_reply":"2022-07-31T18:24:01.705058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"decision_tree_base_model.fit(X_train,y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:24:06.970158Z","iopub.execute_input":"2022-07-31T18:24:06.970565Z","iopub.status.idle":"2022-07-31T18:24:06.982201Z","shell.execute_reply.started":"2022-07-31T18:24:06.970529Z","shell.execute_reply":"2022-07-31T18:24:06.981094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred_decision_tree_base_model = decision_tree_base_model.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:24:12.606982Z","iopub.execute_input":"2022-07-31T18:24:12.607389Z","iopub.status.idle":"2022-07-31T18:24:12.612796Z","shell.execute_reply.started":"2022-07-31T18:24:12.607350Z","shell.execute_reply":"2022-07-31T18:24:12.611999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cm = pd.DataFrame(confusion_matrix(y_test, y_pred_decision_tree_base_model))\ncm.index = ['Actual Died','Actual Survived']\ncm.columns = ['Predicted Died','Predicted Survived']\nprint(cm)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:24:17.285457Z","iopub.execute_input":"2022-07-31T18:24:17.285843Z","iopub.status.idle":"2022-07-31T18:24:17.297133Z","shell.execute_reply.started":"2022-07-31T18:24:17.285809Z","shell.execute_reply":"2022-07-31T18:24:17.295964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set_theme(style=\"white\")\nplot_confusion_matrix(decision_tree_base_model,\n                     X_test,\n                     y_test)\n                     #display_labels = ['Survived (positive)','Not Survived (negative)'])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:24:24.995947Z","iopub.execute_input":"2022-07-31T18:24:24.996333Z","iopub.status.idle":"2022-07-31T18:24:25.219865Z","shell.execute_reply.started":"2022-07-31T18:24:24.996302Z","shell.execute_reply":"2022-07-31T18:24:25.218652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(' Accuracy : %.3f'%metrics.accuracy_score(y_pred_decision_tree_base_model,y_test))\nprint('Precision : %.3f' % precision_score(y_test, y_pred_decision_tree_base_model))\nprint('   Recall : %.3f' % recall_score(y_test, y_pred_decision_tree_base_model))\nprint(' F1 Score : %.3f' % f1_score(y_test, y_pred_decision_tree_base_model))\naccuracy.append(metrics.accuracy_score(y_pred_decision_tree_base_model,y_test))\nprecision.append(precision_score(y_test, y_pred_decision_tree_base_model))\nrecall.append(recall_score(y_test, y_pred_decision_tree_base_model))\nscore_f1.append(f1_score(y_test, y_pred_decision_tree_base_model))","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:24:31.210545Z","iopub.execute_input":"2022-07-31T18:24:31.211364Z","iopub.status.idle":"2022-07-31T18:24:31.226249Z","shell.execute_reply.started":"2022-07-31T18:24:31.211331Z","shell.execute_reply":"2022-07-31T18:24:31.225009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"decision_tree_base_model = DecisionTreeClassifier()\nscores = cross_val_score(decision_tree_base_model, X_train, y_train, cv=10)\nprint(\"%0.2f accuracy with a standard deviation of %0.2f\" % (scores.mean(), scores.std()))","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:24:43.541847Z","iopub.execute_input":"2022-07-31T18:24:43.542602Z","iopub.status.idle":"2022-07-31T18:24:43.576732Z","shell.execute_reply.started":"2022-07-31T18:24:43.542553Z","shell.execute_reply":"2022-07-31T18:24:43.575385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"step_eight_three\"></a>\n## 3.Random Forest","metadata":{}},{"cell_type":"code","source":"random_forest_base_model = RandomForestClassifier()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:25:10.480905Z","iopub.execute_input":"2022-07-31T18:25:10.481256Z","iopub.status.idle":"2022-07-31T18:25:10.485659Z","shell.execute_reply.started":"2022-07-31T18:25:10.481227Z","shell.execute_reply":"2022-07-31T18:25:10.484636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"random_forest_base_model.fit(X_train,y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:25:26.032409Z","iopub.execute_input":"2022-07-31T18:25:26.032855Z","iopub.status.idle":"2022-07-31T18:25:26.251750Z","shell.execute_reply.started":"2022-07-31T18:25:26.032822Z","shell.execute_reply":"2022-07-31T18:25:26.250731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred_random_forest_base_model = random_forest_base_model.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:25:38.954286Z","iopub.execute_input":"2022-07-31T18:25:38.954672Z","iopub.status.idle":"2022-07-31T18:25:38.978124Z","shell.execute_reply.started":"2022-07-31T18:25:38.954641Z","shell.execute_reply":"2022-07-31T18:25:38.976979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cm = pd.DataFrame(confusion_matrix(y_test, y_pred_random_forest_base_model))\ncm.index = ['Actual Died','Actual Survived']\ncm.columns = ['Predicted Died','Predicted Survived']\nprint(cm)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:25:45.914224Z","iopub.execute_input":"2022-07-31T18:25:45.914631Z","iopub.status.idle":"2022-07-31T18:25:45.924797Z","shell.execute_reply.started":"2022-07-31T18:25:45.914597Z","shell.execute_reply":"2022-07-31T18:25:45.923959Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set_theme(style=\"white\")\nplot_confusion_matrix(random_forest_base_model,\n                     X_test,\n                     y_test)\n                     #display_labels = ['Survived (positive)','Not Survived (negative)'])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:25:50.730360Z","iopub.execute_input":"2022-07-31T18:25:50.730752Z","iopub.status.idle":"2022-07-31T18:25:50.915790Z","shell.execute_reply.started":"2022-07-31T18:25:50.730720Z","shell.execute_reply":"2022-07-31T18:25:50.914611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(' Accuracy : %.3f'%metrics.accuracy_score(y_pred_decision_tree_base_model,y_test))\nprint('Precision : %.3f' % precision_score(y_test, y_pred_decision_tree_base_model))\nprint('   Recall : %.3f' % recall_score(y_test, y_pred_decision_tree_base_model))\nprint(' F1 Score : %.3f' % f1_score(y_test, y_pred_decision_tree_base_model))\naccuracy.append(metrics.accuracy_score(y_pred_random_forest_base_model,y_test))\nprecision.append(precision_score(y_test, y_pred_random_forest_base_model))\nrecall.append(recall_score(y_test, y_pred_random_forest_base_model))\nscore_f1.append(f1_score(y_test, y_pred_random_forest_base_model))","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:25:56.933728Z","iopub.execute_input":"2022-07-31T18:25:56.934148Z","iopub.status.idle":"2022-07-31T18:25:56.953011Z","shell.execute_reply.started":"2022-07-31T18:25:56.934114Z","shell.execute_reply":"2022-07-31T18:25:56.951824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"random_forest_base_model = RandomForestClassifier()\nscores = cross_val_score(random_forest_base_model, X_train, y_train, cv=10)\nprint(\"%0.2f accuracy with a standard deviation of %0.2f\" % (scores.mean(), scores.std()))","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:26:02.496124Z","iopub.execute_input":"2022-07-31T18:26:02.496528Z","iopub.status.idle":"2022-07-31T18:26:04.648652Z","shell.execute_reply.started":"2022-07-31T18:26:02.496493Z","shell.execute_reply":"2022-07-31T18:26:04.647461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"results = pd.DataFrame({'accuracy':accuracy,'precision':precision,'recall':recall,'f1_score':score_f1},index=model_name)\nresults","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:26:07.086841Z","iopub.execute_input":"2022-07-31T18:26:07.088028Z","iopub.status.idle":"2022-07-31T18:26:07.103183Z","shell.execute_reply.started":"2022-07-31T18:26:07.087976Z","shell.execute_reply":"2022-07-31T18:26:07.102109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"step_nine\"></a>\n## Classification Report","metadata":{}},{"cell_type":"code","source":"print(\"Logistic Regression\\n\\n\",classification_report(y_test,y_pred_logistic_base_model))","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:26:23.937010Z","iopub.execute_input":"2022-07-31T18:26:23.937401Z","iopub.status.idle":"2022-07-31T18:26:23.949986Z","shell.execute_reply.started":"2022-07-31T18:26:23.937366Z","shell.execute_reply":"2022-07-31T18:26:23.948900Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Decision Treen\\n\\n\",classification_report(y_test,y_pred_decision_tree_base_model))","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:26:35.585692Z","iopub.execute_input":"2022-07-31T18:26:35.586185Z","iopub.status.idle":"2022-07-31T18:26:35.599396Z","shell.execute_reply.started":"2022-07-31T18:26:35.586140Z","shell.execute_reply":"2022-07-31T18:26:35.598501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Random Forest\\n\\n\",classification_report(y_test,y_pred_random_forest_base_model))","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:26:43.533356Z","iopub.execute_input":"2022-07-31T18:26:43.533765Z","iopub.status.idle":"2022-07-31T18:26:43.545786Z","shell.execute_reply.started":"2022-07-31T18:26:43.533734Z","shell.execute_reply":"2022-07-31T18:26:43.544808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"step_ten\"></a>\n## grid search CV","metadata":{}},{"cell_type":"markdown","source":"### Let's apply grid search CV","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV\naccuracy = []\nprecision = []\nrecall = []\nscore_f1 = []\nmodel_name = ['Logistic regression',\"Decision Tree\",\"Random forest\"]\nlogistic_base_model = LogisticRegression()\ndecision_tree_base_model = DecisionTreeClassifier()\nrandom_forest_base_model = RandomForestClassifier()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:27:14.653984Z","iopub.execute_input":"2022-07-31T18:27:14.654403Z","iopub.status.idle":"2022-07-31T18:27:14.660827Z","shell.execute_reply.started":"2022-07-31T18:27:14.654371Z","shell.execute_reply":"2022-07-31T18:27:14.659702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"logistic_params = {'penalty': ['l1','l2', 'elasticnet'],\n                   'C': [0.001,0.01,0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8,0.9,1],\n                  }\nrandom_forest_params = {'n_estimators': [100, 200, 300],\n                        'max_features': ['auto', 'sqrt'],\n                        'max_depth': [10, 60, 110, None],\n                        'min_samples_split': [2, 5],\n                        'min_samples_leaf': [1, 2, 3],\n                        'bootstrap': [True, False]}\ndecision_tree_params = {'criterion':['gini', 'entropy', 'log_loss'],\n                       'splitter':['best', 'random'],\n                        'max_depth':['None',2,4,8,10],\n                        'min_samples_split':[2,4,6,8],\n                        'min_samples_leaf':[1,2],\n                        'max_features':['None','auto', 'sqrt', 'log2']\n                       }","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:27:22.314547Z","iopub.execute_input":"2022-07-31T18:27:22.314948Z","iopub.status.idle":"2022-07-31T18:27:22.324134Z","shell.execute_reply.started":"2022-07-31T18:27:22.314912Z","shell.execute_reply":"2022-07-31T18:27:22.322979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Grid1 = GridSearchCV(logistic_base_model,param_grid=logistic_params,cv=5,n_jobs=-1,scoring='accuracy')\nGrid2 = GridSearchCV(decision_tree_base_model,param_grid=decision_tree_params,cv=5,n_jobs=-1,scoring='accuracy')\nGrid3 = GridSearchCV(random_forest_base_model,param_grid=random_forest_params,cv=5,n_jobs=-1,scoring='accuracy')","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:27:28.131837Z","iopub.execute_input":"2022-07-31T18:27:28.132207Z","iopub.status.idle":"2022-07-31T18:27:28.139146Z","shell.execute_reply.started":"2022-07-31T18:27:28.132176Z","shell.execute_reply":"2022-07-31T18:27:28.137885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import time\nstart = time.time()\nprint('Grid Search Started')\nprint('Logistic Regression Started')\nGrid1.fit(X_train,y_train)\nprint('Decision Tree Started')\nGrid2.fit(X_train,y_train)\nprint('Random Forest Started')\nGrid3.fit(X_train,y_train)\nend = time.time()\nprint('Done !')\nprint('Time Required : ',end-start,' Seconds')","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:27:41.526840Z","iopub.execute_input":"2022-07-31T18:27:41.527219Z","iopub.status.idle":"2022-07-31T18:31:14.935648Z","shell.execute_reply.started":"2022-07-31T18:27:41.527189Z","shell.execute_reply":"2022-07-31T18:31:14.934317Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Estimators of Logistic Regression')\nprint(Grid1.best_estimator_,'\\n')\nprint('Estimators of Decision Tree')\nprint(Grid2.best_estimator_,'\\n')\nprint('Estimators of Random Forest')\nprint(Grid3.best_estimator_)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:31:14.937990Z","iopub.execute_input":"2022-07-31T18:31:14.938334Z","iopub.status.idle":"2022-07-31T18:31:14.946148Z","shell.execute_reply.started":"2022-07-31T18:31:14.938300Z","shell.execute_reply":"2022-07-31T18:31:14.945074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Score Logistic Regression')\nprint(Grid1.best_score_,'\\n')\nprint('Score Decision Tree')\nprint(Grid2.best_score_,'\\n')\nprint('Score Random Forest')\nprint(Grid3.best_score_)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:31:14.947840Z","iopub.execute_input":"2022-07-31T18:31:14.948564Z","iopub.status.idle":"2022-07-31T18:31:14.958029Z","shell.execute_reply.started":"2022-07-31T18:31:14.948520Z","shell.execute_reply":"2022-07-31T18:31:14.956871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Parameters of Logistic Regression')\nprint(Grid1.best_params_,'\\n')\nprint('Parameters of Decision Tree')\nprint(Grid2.best_params_,'\\n')\nprint('Parameters of Random Forest')\nprint(Grid3.best_params_)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:31:14.961534Z","iopub.execute_input":"2022-07-31T18:31:14.961957Z","iopub.status.idle":"2022-07-31T18:31:14.969017Z","shell.execute_reply.started":"2022-07-31T18:31:14.961913Z","shell.execute_reply":"2022-07-31T18:31:14.968208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"All three models showed almost same result<br>\nLet's build the model using best params","metadata":{}},{"cell_type":"code","source":"lr_model = LogisticRegression(C= 0.2, penalty= 'l2')\ndt_model = DecisionTreeClassifier(criterion= 'entropy', max_depth= 4, max_features= 'log2', min_samples_leaf= 1,\n                                  min_samples_split= 8, splitter= 'best')\nrf_model = RandomForestClassifier(bootstrap= True, max_depth= 60, max_features= 'auto', min_samples_leaf= 3,\n                                  min_samples_split= 5, n_estimators= 100)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:31:14.970365Z","iopub.execute_input":"2022-07-31T18:31:14.970976Z","iopub.status.idle":"2022-07-31T18:31:14.982048Z","shell.execute_reply.started":"2022-07-31T18:31:14.970946Z","shell.execute_reply":"2022-07-31T18:31:14.981305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lr_model.fit(X_train,y_train)\ndt_model.fit(X_train,y_train)\nrf_model.fit(X_train,y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:31:14.983764Z","iopub.execute_input":"2022-07-31T18:31:14.984562Z","iopub.status.idle":"2022-07-31T18:31:15.200422Z","shell.execute_reply.started":"2022-07-31T18:31:14.984518Z","shell.execute_reply":"2022-07-31T18:31:15.199206Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"step_eleven\"></a>\n## Save model","metadata":{}},{"cell_type":"code","source":"#import pickle\n#pickle.dump(lr_model,open('lr_titanic_model.pkl','wb'))\n#pickle.dump(dt_model,open('dt_titanic_model.pkl','wb'))\n#pickle.dump(rf_model,open('lr_titanic_model.pkl','wb'))","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:31:15.202607Z","iopub.execute_input":"2022-07-31T18:31:15.203441Z","iopub.status.idle":"2022-07-31T18:31:15.208572Z","shell.execute_reply.started":"2022-07-31T18:31:15.203380Z","shell.execute_reply":"2022-07-31T18:31:15.207338Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Thank You","metadata":{}},{"cell_type":"code","source":"def pipeline(df):\n    data = df.copy()\n    data.fillna(0,inplace=True)\n    data['Sex'].replace(['male','female'],[0,1],inplace=True)\n    data['Embarked'].replace(['S','C','Q'],[0,1,2],inplace=True)\n    for i in data:\n        data['Initial'] = data.Name.str.extract('([A-Za-z]+)\\.')\n    data['Initial'].replace(['Mlle','Mme' ,'Ms'  ,'Dr','Major','Lady','Countess','Jonkheer','Col'  ,'Rev'  ,'Capt','Sir','Don','Dona'],\n                        ['Miss','Miss','Miss','Mr','Mr'   ,'Mrs'  ,'Mrs'    ,'Other'   ,'Other','Other','Mr'  ,'Mr' ,'Mr','Mrs'],\n                       inplace=True)\n    data['Initial'].replace(['Mr','Mrs','Miss','Master','Other'],[0,1,2,3,4],inplace=True)\n    data['Family_Size'] = data['Parch']+data['SibSp']\n    data['Alone'] = 0\n    data.loc[data.Family_Size==0,'Alone']=1\n    scale_cols = ['Age','Fare']\n    for col in scale_cols:\n        data[col] = scaler.transform(data[[col]])\n    data.drop(['Name','Ticket','Cabin','PassengerId','SibSp','Parch'],axis=1,inplace=True)\n    return data","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:51:53.518603Z","iopub.execute_input":"2022-07-31T18:51:53.519018Z","iopub.status.idle":"2022-07-31T18:51:53.531824Z","shell.execute_reply.started":"2022-07-31T18:51:53.518983Z","shell.execute_reply":"2022-07-31T18:51:53.530407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"a = pd.read_csv('../input/titanic/test.csv')\na.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:51:58.524167Z","iopub.execute_input":"2022-07-31T18:51:58.524552Z","iopub.status.idle":"2022-07-31T18:51:58.546115Z","shell.execute_reply.started":"2022-07-31T18:51:58.524520Z","shell.execute_reply":"2022-07-31T18:51:58.545424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_data = pipeline(a)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:52:01.502838Z","iopub.execute_input":"2022-07-31T18:52:01.503798Z","iopub.status.idle":"2022-07-31T18:52:01.546336Z","shell.execute_reply.started":"2022-07-31T18:52:01.503758Z","shell.execute_reply":"2022-07-31T18:52:01.545438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:52:04.392956Z","iopub.execute_input":"2022-07-31T18:52:04.393351Z","iopub.status.idle":"2022-07-31T18:52:04.405818Z","shell.execute_reply.started":"2022-07-31T18:52:04.393317Z","shell.execute_reply":"2022-07-31T18:52:04.404749Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred1 = lr_model.predict(X_data)\npred2 = dt_model.predict(X_data)\npred3 = rf_model.predict(X_data)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T18:52:12.336322Z","iopub.execute_input":"2022-07-31T18:52:12.336769Z","iopub.status.idle":"2022-07-31T18:52:12.366977Z","shell.execute_reply.started":"2022-07-31T18:52:12.336734Z","shell.execute_reply":"2022-07-31T18:52:12.365832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"result1 = pd.DataFrame({'PassengerId':a.PassengerId,'Survived':pred1})\nresult2 = pd.DataFrame({'PassengerId':a.PassengerId,'Survived':pred2})\nresult3 = pd.DataFrame({'PassengerId':a.PassengerId,'Survived':pred3})","metadata":{"execution":{"iopub.status.busy":"2022-07-31T19:07:24.179589Z","iopub.execute_input":"2022-07-31T19:07:24.179978Z","iopub.status.idle":"2022-07-31T19:07:24.187887Z","shell.execute_reply.started":"2022-07-31T19:07:24.179937Z","shell.execute_reply":"2022-07-31T19:07:24.186770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"result3.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T19:07:35.717084Z","iopub.execute_input":"2022-07-31T19:07:35.717492Z","iopub.status.idle":"2022-07-31T19:07:35.728717Z","shell.execute_reply.started":"2022-07-31T19:07:35.717458Z","shell.execute_reply":"2022-07-31T19:07:35.727243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"result3.to_csv('submission.csv',index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T19:07:41.261005Z","iopub.execute_input":"2022-07-31T19:07:41.261370Z","iopub.status.idle":"2022-07-31T19:07:41.267688Z","shell.execute_reply.started":"2022-07-31T19:07:41.261342Z","shell.execute_reply":"2022-07-31T19:07:41.266678Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}