{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"###### Introduction\nThe sinking of Titanic is one of the most notorious shipwrecks in the history. In 1912, during her voyage, the Titanic sank after colliding with an iceberg,killing 1502 out of 2224 passangers and crew.\n\n<font color = 'blue'>\n      \nContent: \n      \n1. [Load and Check Data](#1)\n2. [Variable Description](#2)\n    * [Univariate Variable Analysis](#3)\n        * [Categorical Variable Analysis](#4)\n        * [Numerical Variable Analysis](#5)\n3. [Basic Data Analysis](#6)\n4. [Outlier Detection](#7)\n5. [Missing Values](#8)\n    * [ Find Missing Values](#9)\n    * [ Fill Missing Values](#10)\n  \n6. [Visualization](#11)\n     * [Correlation Between Sibsp -- Parch -- Age -- Fare -- Survived](#12)\n     * [SibSp -- Survived](#13)\n     * [Parch -- Survived](#14)\n     * [Pclass -- Survived](#15)\n     * [Age -- Survived](#16)\n     * [Pclass -- Survived -- Age](#17)\n     * [Embarked -- Sex -- Pclass -- Survived](#18)\n     * [Embarked -- Sex -- Fare -- Survived](#19)\n     * [Fill Missing: Age Feature](#20)\n    ","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nplt.style.use(\"seaborn-whitegrid\")\n\nimport seaborn as sns\n\nfrom collections import Counter\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-08-07T15:53:54.359014Z","iopub.execute_input":"2022-08-07T15:53:54.360136Z","iopub.status.idle":"2022-08-07T15:53:54.377180Z","shell.execute_reply.started":"2022-08-07T15:53:54.360071Z","shell.execute_reply":"2022-08-07T15:53:54.375682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"1\"></a><br>\n## Load and Check Data","metadata":{}},{"cell_type":"code","source":"#upload dataset\ntrain_df = pd.read_csv(\"/kaggle/input/titanic/train.csv\")\ntest_df = pd.read_csv(\"/kaggle/input/titanic/test.csv\")\ntest_PassangerId = test_df[\"PassengerId\"]","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:54.379369Z","iopub.execute_input":"2022-08-07T15:53:54.380892Z","iopub.status.idle":"2022-08-07T15:53:54.403038Z","shell.execute_reply.started":"2022-08-07T15:53:54.380842Z","shell.execute_reply":"2022-08-07T15:53:54.401518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#columns of dataframe\ntrain_df.columns","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:54.404978Z","iopub.execute_input":"2022-08-07T15:53:54.405944Z","iopub.status.idle":"2022-08-07T15:53:54.415031Z","shell.execute_reply.started":"2022-08-07T15:53:54.405874Z","shell.execute_reply":"2022-08-07T15:53:54.413629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#first 5 data of dataframe to check out what we're dealing with.\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:54.417125Z","iopub.execute_input":"2022-08-07T15:53:54.417468Z","iopub.status.idle":"2022-08-07T15:53:54.446260Z","shell.execute_reply.started":"2022-08-07T15:53:54.417435Z","shell.execute_reply":"2022-08-07T15:53:54.444619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#some statistical values of dataframe\ntrain_df.describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:54.448723Z","iopub.execute_input":"2022-08-07T15:53:54.449243Z","iopub.status.idle":"2022-08-07T15:53:54.489471Z","shell.execute_reply.started":"2022-08-07T15:53:54.449182Z","shell.execute_reply":"2022-08-07T15:53:54.487805Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"2\"></a><br>\n## Variable Description\n1. \nPassangerId: Id number to each passanger\n1. \nSurvived: Passanger survive(1) or died(0)\n1. \nPclass: Passanger Class\n1. \nName: Passanger Name\n1. \nSex: Gender of Passanger\n1. \nAge: Age of Passanger\n1. \nSibSp: Number of Siblings/Spouses\n1. \nParch: Number of Parents/Children\n1. \nTicket: Number of ticket\n1. \nFare: Amount of money spent on ticket\n1. \nCabin: Cabin Category\n1. \nEmbarked: Port where passanger embarked (C = Cherbourg, Q= Queenstown, S= Southampton)\n","metadata":{}},{"cell_type":"code","source":"#information of dataframe\ntrain_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:54.493707Z","iopub.execute_input":"2022-08-07T15:53:54.494266Z","iopub.status.idle":"2022-08-07T15:53:54.512782Z","shell.execute_reply.started":"2022-08-07T15:53:54.494228Z","shell.execute_reply":"2022-08-07T15:53:54.511091Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* \nfloat64(2): Fare and Age\n* \nint64(5): PassengerId,Survived,Pclass,SibSp,Parch,\n* \nobject(5): Name,Sex,Ticket,Cabin,Embarked","metadata":{}},{"cell_type":"markdown","source":"<a id =\"3\"></a><br>\n# Univariate Variable Analysis\n* Categorical Variable: Survived,Sex,Pclass,Embarked,Cabin,Name,Ticket,SibSp and Parch\n* Numerical Variable: PassangerId,Age,Fare","metadata":{}},{"cell_type":"markdown","source":"<a id =\"4\"></a><br>\n# Categorical Variable","metadata":{}},{"cell_type":"code","source":"def ba_plot(variable):\n    \"\"\"\n    input:variable ex: \"Sex\"\n    output: bar plot & value count\n    \"\"\"\n    #get feature\n    var = train_df[variable]\n    #count number of categorical variable(value/sample)\n    varValue = var.value_counts()\n    \n    #visualize\n    \n    plt.figure(figsize = (9,3))\n    plt.bar(varValue.index,varValue)\n    plt.xticks(varValue.index,varValue.index.values)\n    plt.ylabel(\"Frequency\")\n    plt.title(variable)\n    plt.show()\n    print(\"{}:\\n {}\".format(variable,varValue))","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:54.515205Z","iopub.execute_input":"2022-08-07T15:53:54.516162Z","iopub.status.idle":"2022-08-07T15:53:54.525294Z","shell.execute_reply.started":"2022-08-07T15:53:54.516113Z","shell.execute_reply":"2022-08-07T15:53:54.524313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"category1 = [\"Survived\",\"Sex\",\"Pclass\",\"Embarked\",\"SibSp\",\"Parch\"]\nfor c in category1:\n    ba_plot(c)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:54.527356Z","iopub.execute_input":"2022-08-07T15:53:54.528317Z","iopub.status.idle":"2022-08-07T15:53:55.796970Z","shell.execute_reply.started":"2022-08-07T15:53:54.528265Z","shell.execute_reply":"2022-08-07T15:53:55.795480Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"5\"></a><br>\n# Numerical Variable ","metadata":{}},{"cell_type":"code","source":"def plot_hist(variable):\n    \n    plt.figure(figsize = (9,3))\n    \n    plt.hist(train_df[variable],bins = 50)\n    plt.xlabel(variable)\n    plt.ylabel(\"Frequency\")\n    plt.title(\"{} distribution with histogram\".format(variable))\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:55.798718Z","iopub.execute_input":"2022-08-07T15:53:55.799892Z","iopub.status.idle":"2022-08-07T15:53:55.808334Z","shell.execute_reply.started":"2022-08-07T15:53:55.799842Z","shell.execute_reply":"2022-08-07T15:53:55.806637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"numericVar =  [\"PassengerId\",\"Age\",\"Fare\"]\n\nfor c in numericVar:\n    plot_hist(c)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:55.809839Z","iopub.execute_input":"2022-08-07T15:53:55.810508Z","iopub.status.idle":"2022-08-07T15:53:56.625158Z","shell.execute_reply.started":"2022-08-07T15:53:55.810473Z","shell.execute_reply":"2022-08-07T15:53:56.623680Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"6\"></a><br>\n# Basic Data Analysis\n* Pclass - Survived\n* Sex - Survived\n* SibSp - Survived\n* Parch - Survived","metadata":{}},{"cell_type":"code","source":"# Pclass vs Survived\n# train_df.groupby([\"Survived\",\"Pclass\"]).size()\ntrain_df[[\"Pclass\",\"Survived\"]].groupby([\"Pclass\"],as_index = False).mean().sort_values(by = \"Survived\",ascending = False)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:56.626872Z","iopub.execute_input":"2022-08-07T15:53:56.627401Z","iopub.status.idle":"2022-08-07T15:53:56.645617Z","shell.execute_reply.started":"2022-08-07T15:53:56.627354Z","shell.execute_reply":"2022-08-07T15:53:56.644655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Sex vs Survived\n# train_df.groupby([\"Survived\",\"Sex\"]).size()\ntrain_df[[\"Sex\",\"Survived\"]].groupby([\"Sex\"],as_index = False).mean().sort_values(by = \"Survived\",ascending = False)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:56.646963Z","iopub.execute_input":"2022-08-07T15:53:56.648031Z","iopub.status.idle":"2022-08-07T15:53:56.665187Z","shell.execute_reply.started":"2022-08-07T15:53:56.647997Z","shell.execute_reply":"2022-08-07T15:53:56.664211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# SibSp vs Survived\n# train_df.groupby([\"Survived\",\"SibSp\"]).size()\ntrain_df[[\"SibSp\",\"Survived\"]].groupby([\"SibSp\"],as_index = False).mean().sort_values(by = \"Survived\",ascending = False)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:56.666766Z","iopub.execute_input":"2022-08-07T15:53:56.667113Z","iopub.status.idle":"2022-08-07T15:53:56.684853Z","shell.execute_reply.started":"2022-08-07T15:53:56.667080Z","shell.execute_reply":"2022-08-07T15:53:56.683241Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Parch vs Survived\n# train_df.groupby([\"Survived\",\"Sex\"]).size()\ntrain_df[[\"Parch\",\"Survived\"]].groupby([\"Parch\"],as_index = False).mean().sort_values(by = \"Survived\",ascending = False)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:56.686655Z","iopub.execute_input":"2022-08-07T15:53:56.687105Z","iopub.status.idle":"2022-08-07T15:53:56.704554Z","shell.execute_reply.started":"2022-08-07T15:53:56.687067Z","shell.execute_reply":"2022-08-07T15:53:56.703177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"7\"></a><br>\n# Outlier Detection","metadata":{}},{"cell_type":"code","source":"def detect_outliers(df,features):\n    \n    outlier_indices = []\n    \n    for c in features:\n        \n        # 1st quartile\n        \n        Q1 = np.percentile(df[c],25)\n        \n        # 3rd quartile\n        \n        Q3 = np.percentile(df[c],50)\n        \n        # IQR\n        \n        IQR = Q3 - Q1\n        \n        # Outlier step\n        outlier_step = IQR *1.5\n        \n        # Detect outlier and their indeces\n        \n        outlier_list_col = df[(df[c] < Q1 - outlier_step) | (df[c] > Q3 + outlier_step)].index\n        \n        # store indeces\n        outlier_indices.extend(outlier_list_col)\n    \n    outlier_indices = Counter(outlier_indices)\n    multiple_outliers = list(i for i, v in outlier_indices.items() if v > 2)\n    \n    return multiple_outliers\n        ","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:56.706138Z","iopub.execute_input":"2022-08-07T15:53:56.707040Z","iopub.status.idle":"2022-08-07T15:53:56.715163Z","shell.execute_reply.started":"2022-08-07T15:53:56.707002Z","shell.execute_reply":"2022-08-07T15:53:56.713506Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.loc[detect_outliers(train_df,[\"Age\",\"SibSp\",\"Parch\",\"Fare\"])]","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:56.716982Z","iopub.execute_input":"2022-08-07T15:53:56.718136Z","iopub.status.idle":"2022-08-07T15:53:56.761883Z","shell.execute_reply.started":"2022-08-07T15:53:56.718101Z","shell.execute_reply":"2022-08-07T15:53:56.760245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#drop outliers\ntrain_df = train_df.drop(detect_outliers(train_df,[\"Age\",\"SibSp\",\"Parch\",\"Fare\"]),axis = 0).reset_index(drop = True)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:56.763460Z","iopub.execute_input":"2022-08-07T15:53:56.763904Z","iopub.status.idle":"2022-08-07T15:53:56.779370Z","shell.execute_reply.started":"2022-08-07T15:53:56.763871Z","shell.execute_reply":"2022-08-07T15:53:56.777969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"8\"></a><br>\n# Missing Values\n* Find Missing Values\n* Fill Missing Values","metadata":{}},{"cell_type":"code","source":"train_df_len = len(train_df)\ntrain_df = pd.concat([train_df,test_df],axis = 0).reset_index(drop = True)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:56.790744Z","iopub.execute_input":"2022-08-07T15:53:56.791117Z","iopub.status.idle":"2022-08-07T15:53:56.800774Z","shell.execute_reply.started":"2022-08-07T15:53:56.791074Z","shell.execute_reply":"2022-08-07T15:53:56.799835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:56.802205Z","iopub.execute_input":"2022-08-07T15:53:56.802530Z","iopub.status.idle":"2022-08-07T15:53:56.823545Z","shell.execute_reply.started":"2022-08-07T15:53:56.802498Z","shell.execute_reply":"2022-08-07T15:53:56.822027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.columns[train_df.isnull().any()]","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:56.825779Z","iopub.execute_input":"2022-08-07T15:53:56.826271Z","iopub.status.idle":"2022-08-07T15:53:56.839494Z","shell.execute_reply.started":"2022-08-07T15:53:56.826234Z","shell.execute_reply":"2022-08-07T15:53:56.838316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:56.841253Z","iopub.execute_input":"2022-08-07T15:53:56.842497Z","iopub.status.idle":"2022-08-07T15:53:56.855405Z","shell.execute_reply.started":"2022-08-07T15:53:56.842429Z","shell.execute_reply":"2022-08-07T15:53:56.853863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =\"10\"></a><br>\n## Fill Missing Values\n* Embarked has 2 missing values\n* Fare has only 1 missing value","metadata":{}},{"cell_type":"code","source":"train_df[train_df[\"Embarked\"].isnull()]","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:56.857095Z","iopub.execute_input":"2022-08-07T15:53:56.858528Z","iopub.status.idle":"2022-08-07T15:53:56.879105Z","shell.execute_reply.started":"2022-08-07T15:53:56.858493Z","shell.execute_reply":"2022-08-07T15:53:56.877850Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":" #estime Embarked using Fare column\ntrain_df.boxplot(column = \"Fare\", by = \"Embarked\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:56.880422Z","iopub.execute_input":"2022-08-07T15:53:56.880767Z","iopub.status.idle":"2022-08-07T15:53:57.055366Z","shell.execute_reply.started":"2022-08-07T15:53:56.880735Z","shell.execute_reply":"2022-08-07T15:53:57.053974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#filling NaN values with 'C' Embarked\ntrain_df[\"Embarked\"] = train_df[\"Embarked\"].fillna(\"C\")\ntrain_df[train_df[\"Embarked\"].isnull()]","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:57.058244Z","iopub.execute_input":"2022-08-07T15:53:57.059171Z","iopub.status.idle":"2022-08-07T15:53:57.075059Z","shell.execute_reply.started":"2022-08-07T15:53:57.059109Z","shell.execute_reply":"2022-08-07T15:53:57.073605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[train_df[\"Fare\"].isnull()]","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:57.076937Z","iopub.execute_input":"2022-08-07T15:53:57.078350Z","iopub.status.idle":"2022-08-07T15:53:57.099515Z","shell.execute_reply.started":"2022-08-07T15:53:57.078300Z","shell.execute_reply":"2022-08-07T15:53:57.098176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#filling NaN values with using mean of Fare which also paid by 3. class cruise ship\ntrain_df[\"Fare\"] =train_df[\"Fare\"].fillna(np.mean(train_df[train_df[\"Pclass\"] == 3][\"Fare\"]))\ntrain_df[train_df[\"Fare\"].isnull()]","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:57.102143Z","iopub.execute_input":"2022-08-07T15:53:57.102740Z","iopub.status.idle":"2022-08-07T15:53:57.121949Z","shell.execute_reply.started":"2022-08-07T15:53:57.102691Z","shell.execute_reply":"2022-08-07T15:53:57.120996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualization\n<a id =\"11\"></a><br>","metadata":{}},{"cell_type":"markdown","source":"## Correlation Between Sibsp -- Parch -- Age -- Fare -- Survived\n<a id =\"12\"></a><br>","metadata":{}},{"cell_type":"code","source":"sns.heatmap(train_df[[\"SibSp\", \"Parch\", \"Age\", \"Fare\", \"Survived\"]].corr(),annot = True,fmt = \".2f\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:57.123238Z","iopub.execute_input":"2022-08-07T15:53:57.124330Z","iopub.status.idle":"2022-08-07T15:53:57.461727Z","shell.execute_reply.started":"2022-08-07T15:53:57.124286Z","shell.execute_reply":"2022-08-07T15:53:57.460389Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Fate feature seems to have corralation with survived feature (0.26)","metadata":{}},{"cell_type":"markdown","source":"## SibSp -- Survived\n<a id = \"13\"></a><br>","metadata":{}},{"cell_type":"code","source":"df = train_df[train_df[\"SibSp\"] == 4]\ndf","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:57.463294Z","iopub.execute_input":"2022-08-07T15:53:57.463635Z","iopub.status.idle":"2022-08-07T15:53:57.487140Z","shell.execute_reply.started":"2022-08-07T15:53:57.463601Z","shell.execute_reply":"2022-08-07T15:53:57.485469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.histplot(data = train_df, x = \"SibSp\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:57.488495Z","iopub.execute_input":"2022-08-07T15:53:57.489371Z","iopub.status.idle":"2022-08-07T15:53:57.822439Z","shell.execute_reply.started":"2022-08-07T15:53:57.489321Z","shell.execute_reply":"2022-08-07T15:53:57.821071Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"g = sns.factorplot(x = \"SibSp\",y = \"Survived\", data = train_df, kind =  \"bar\",size = 6)\ng.set_ylabels(\"Survived Probability\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:57.824022Z","iopub.execute_input":"2022-08-07T15:53:57.824443Z","iopub.status.idle":"2022-08-07T15:53:58.302537Z","shell.execute_reply.started":"2022-08-07T15:53:57.824410Z","shell.execute_reply":"2022-08-07T15:53:58.301619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Having a lot of SibSp has less chance to survive\n* if sibps = 0 or 1 or 2, passenger has more chance to survive (if SibSp = 4 ,it seems has certain survived probability.However, there are 5 passengers and 4 of them are anknown in terms of survived feature)\n* We can consider a new feature describing these categories","metadata":{}},{"cell_type":"markdown","source":"## Parch -- Survived\n<a id = \"14\"></a><br>","metadata":{}},{"cell_type":"code","source":"g = sns.factorplot(x = \"Parch\", y = \"Survived\", kind = \"bar\", data = train_df, size = 6)\ng.set_ylabels(\"Survived Probability\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:58.303641Z","iopub.execute_input":"2022-08-07T15:53:58.304388Z","iopub.status.idle":"2022-08-07T15:53:58.699361Z","shell.execute_reply.started":"2022-08-07T15:53:58.304355Z","shell.execute_reply":"2022-08-07T15:53:58.698335Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* SibSp and parch can be used for new feature extraction with th =3\n* Small families have more chance to survive\n* There is std in survival of passanger with parch = 3","metadata":{}},{"cell_type":"markdown","source":"## Pclass -- Survived\n<a id = \"15\"></a><br>","metadata":{}},{"cell_type":"code","source":"g = sns.factorplot(x = \"Pclass\", y = \"Survived\", data = train_df, kind = \"bar\", size = 6)\ng.set_ylabels(\"Survived Probability\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:58.700936Z","iopub.execute_input":"2022-08-07T15:53:58.701828Z","iopub.status.idle":"2022-08-07T15:53:59.559430Z","shell.execute_reply.started":"2022-08-07T15:53:58.701712Z","shell.execute_reply":"2022-08-07T15:53:59.558127Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Age -- Survived\n<a id = \"16\"></a><br>","metadata":{}},{"cell_type":"code","source":"g = sns.FacetGrid(train_df,col = \"Survived\")\ng.map(sns.distplot,\"Age\",bins = 25)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:53:59.561350Z","iopub.execute_input":"2022-08-07T15:53:59.562787Z","iopub.status.idle":"2022-08-07T15:54:00.048253Z","shell.execute_reply.started":"2022-08-07T15:53:59.562747Z","shell.execute_reply":"2022-08-07T15:54:00.046849Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Age <= 10 has high survival rate.\n* Old passengers (80) survived rate is high.\n* Large number of 20 years old did not survive\n* Most pasenfers are in 15-35 age range.\n* Use age feature to fill NaN values in Age feature.","metadata":{}},{"cell_type":"markdown","source":"## Pclass -- Survived -- Age\n<a id = \"17\"></a><br>","metadata":{}},{"cell_type":"code","source":"g = sns.FacetGrid(train_df,col = \"Survived\",row = \"Pclass\",size= 3)\ng.map(plt.hist,\"Age\",bins = 25)\ng.add_legend()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:54:00.050041Z","iopub.execute_input":"2022-08-07T15:54:00.051154Z","iopub.status.idle":"2022-08-07T15:54:01.730754Z","shell.execute_reply.started":"2022-08-07T15:54:00.051114Z","shell.execute_reply":"2022-08-07T15:54:01.729465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Pclass is important feature for model traning.","metadata":{}},{"cell_type":"markdown","source":"## Embarked -- Sex -- Fare -- Survived\n<a id = \"18\"></a><br>","metadata":{}},{"cell_type":"code","source":"g = sns.FacetGrid(train_df,row = \"Embarked\",col = \"Survived\",size = 3)\ng.map(sns.barplot,\"Sex\",\"Fare\")\ng.add_legend()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:54:01.732083Z","iopub.execute_input":"2022-08-07T15:54:01.732422Z","iopub.status.idle":"2022-08-07T15:54:03.179939Z","shell.execute_reply.started":"2022-08-07T15:54:01.732390Z","shell.execute_reply":"2022-08-07T15:54:03.178687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Female passangers have much survival rate then males\n* Passangers who pay higher fare have better survival rate. Fare can be used as categorical for training\n","metadata":{}},{"cell_type":"markdown","source":"## Embarked -- Sex -- Pclass -- Survived\n<a id = \"19\"></a><br>","metadata":{}},{"cell_type":"code","source":"g = sns.FacetGrid(train_df,row = \"Embarked\",size = 3)\ng.map(sns.pointplot,\"Pclass\",\"Survived\",\"Sex\")\ng.add_legend()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:54:03.181427Z","iopub.execute_input":"2022-08-07T15:54:03.181766Z","iopub.status.idle":"2022-08-07T15:54:04.591024Z","shell.execute_reply.started":"2022-08-07T15:54:03.181730Z","shell.execute_reply":"2022-08-07T15:54:04.589500Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Males have better survival rate in Pclass (3) in C\n* Embarked and sex can be used for training.","metadata":{}},{"cell_type":"markdown","source":"## Fill Missing: Age Feature\n<a id = \"20\"></a><br>","metadata":{}},{"cell_type":"code","source":"train_df[train_df[\"Age\"].isnull()]","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:54:04.592735Z","iopub.execute_input":"2022-08-07T15:54:04.593260Z","iopub.status.idle":"2022-08-07T15:54:04.625719Z","shell.execute_reply.started":"2022-08-07T15:54:04.593223Z","shell.execute_reply":"2022-08-07T15:54:04.624756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.factorplot(x = \"Sex\",y = \"Age\",data = train_df,kind = \"box\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:54:04.627055Z","iopub.execute_input":"2022-08-07T15:54:04.627620Z","iopub.status.idle":"2022-08-07T15:54:04.971424Z","shell.execute_reply.started":"2022-08-07T15:54:04.627585Z","shell.execute_reply":"2022-08-07T15:54:04.970016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Sex is not informative for age prediction. Age distribution seems to be same.","metadata":{}},{"cell_type":"code","source":"sns.factorplot(x = \"Sex\",y = \"Age\", hue = \"Pclass\",data = train_df,kind = \"box\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:54:04.973480Z","iopub.execute_input":"2022-08-07T15:54:04.973878Z","iopub.status.idle":"2022-08-07T15:54:05.388656Z","shell.execute_reply.started":"2022-08-07T15:54:04.973844Z","shell.execute_reply":"2022-08-07T15:54:05.387301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* 1st class passengers are older than 2nd, and 2nd is older than 3rd clas.","metadata":{}},{"cell_type":"code","source":"sns.factorplot(x = \"Parch\",y = \"Age\",data = train_df,kind = \"box\")\nsns.factorplot(x = \"SibSp\",y = \"Age\",data = train_df,kind = \"box\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:54:05.390506Z","iopub.execute_input":"2022-08-07T15:54:05.391056Z","iopub.status.idle":"2022-08-07T15:54:06.036533Z","shell.execute_reply.started":"2022-08-07T15:54:05.391005Z","shell.execute_reply":"2022-08-07T15:54:06.035473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[\"Sex\"]","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:54:06.037975Z","iopub.execute_input":"2022-08-07T15:54:06.038528Z","iopub.status.idle":"2022-08-07T15:54:06.047412Z","shell.execute_reply.started":"2022-08-07T15:54:06.038495Z","shell.execute_reply":"2022-08-07T15:54:06.045797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[\"Sex\"] = [1 if i == \"male\" else 0 for i in train_df[\"Sex\"]]","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:54:06.049156Z","iopub.execute_input":"2022-08-07T15:54:06.049558Z","iopub.status.idle":"2022-08-07T15:54:06.060715Z","shell.execute_reply.started":"2022-08-07T15:54:06.049525Z","shell.execute_reply":"2022-08-07T15:54:06.059473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.heatmap(train_df[[\"Age\",\"Sex\",\"SibSp\",\"Parch\",\"Pclass\"]].corr(),annot = True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:54:06.062200Z","iopub.execute_input":"2022-08-07T15:54:06.062789Z","iopub.status.idle":"2022-08-07T15:54:06.362375Z","shell.execute_reply.started":"2022-08-07T15:54:06.062757Z","shell.execute_reply":"2022-08-07T15:54:06.360842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Age is not correlated with sex but it is correlated with SibSp, Parch, Pclass","metadata":{}},{"cell_type":"code","source":"index_nan_age = list(train_df[\"Age\"][train_df[\"Age\"].isnull()].index)\n\nfor i in index_nan_age:\n    \n    age_pred = train_df[\"Age\"][((train_df[\"SibSp\"] == train_df.iloc[i][\"SibSp\"]) &(train_df[\"Parch\"] == train_df.iloc[i][\"Parch\"])& (train_df[\"Pclass\"] == train_df.iloc[i][\"Pclass\"]))].median()\n    age_med = train_df[\"Age\"].median()\n    if not np.isnan(age_pred):\n        train_df[\"Age\"].iloc[i] = age_pred\n    else:\n        train_df[\"Age\"].iloc[i] = age_med\n        ","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:54:06.364237Z","iopub.execute_input":"2022-08-07T15:54:06.364593Z","iopub.status.idle":"2022-08-07T15:54:07.020694Z","shell.execute_reply.started":"2022-08-07T15:54:06.364560Z","shell.execute_reply":"2022-08-07T15:54:07.019356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[train_df[\"Age\"].isnull()]","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:54:07.022228Z","iopub.execute_input":"2022-08-07T15:54:07.022578Z","iopub.status.idle":"2022-08-07T15:54:07.036595Z","shell.execute_reply.started":"2022-08-07T15:54:07.022545Z","shell.execute_reply":"2022-08-07T15:54:07.035205Z"},"trusted":true},"execution_count":null,"outputs":[]}]}