{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# INTRODUCTION\nThe sinking of Titanic is one of the most notorius shipwrecks in the history. In 1912, during her voyage, the Titanic sank after colliding with an iceberg. Killing 1502 out of 2224 passengers and crew.\n\n<font color=\"orange\">\nContent:\n\n1.[Load and Check Data](#1) <br>\n2.[Variable Description](#2) <br>\n*     [Univariate Variable Analysis](#3) <br>\n *          [Categorical Variable Analysis](#4) <br>\n *         [Numerical Variable Analysis](#5) <br>\n3.[Basic Data Analysis](#6)<br>\n4.[Outlier Detection](#7)<br>\n5.[Missing Value](#8)\n* [Find Missing Value](#9)   \n* [Fill Missing Value](#10)    ","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nplt.style.use(\"seaborn-paper\")\n\nimport seaborn as sns\n\nfrom collections import Counter\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-08-06T14:57:31.571701Z","iopub.execute_input":"2022-08-06T14:57:31.572328Z","iopub.status.idle":"2022-08-06T14:57:32.094751Z","shell.execute_reply.started":"2022-08-06T14:57:31.572265Z","shell.execute_reply":"2022-08-06T14:57:32.093650Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=1> </a> <br>\n## 1. Load and Check Data","metadata":{}},{"cell_type":"code","source":"train_df= pd.read_csv(\"/kaggle/input/titanic/train.csv\")\ntest_df= pd.read_csv(\"/kaggle/input/titanic/test.csv\")\n\ntest_PassengerId= test_df[\"PassengerId\"]","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:57:32.100223Z","iopub.execute_input":"2022-08-06T14:57:32.100526Z","iopub.status.idle":"2022-08-06T14:57:32.115982Z","shell.execute_reply.started":"2022-08-06T14:57:32.100497Z","shell.execute_reply":"2022-08-06T14:57:32.115021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.columns","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:57:32.117451Z","iopub.execute_input":"2022-08-06T14:57:32.118304Z","iopub.status.idle":"2022-08-06T14:57:32.127814Z","shell.execute_reply.started":"2022-08-06T14:57:32.118270Z","shell.execute_reply":"2022-08-06T14:57:32.126595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:57:32.131765Z","iopub.execute_input":"2022-08-06T14:57:32.133176Z","iopub.status.idle":"2022-08-06T14:57:32.152370Z","shell.execute_reply.started":"2022-08-06T14:57:32.133143Z","shell.execute_reply":"2022-08-06T14:57:32.151145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:57:32.153401Z","iopub.execute_input":"2022-08-06T14:57:32.154297Z","iopub.status.idle":"2022-08-06T14:57:32.187847Z","shell.execute_reply.started":"2022-08-06T14:57:32.154252Z","shell.execute_reply":"2022-08-06T14:57:32.187011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=2> </a> <br>\n## 2. Variable Description\n\n1.PassengerId : Unique id number to each passenger <br>\n2.Survived : Passenger survive(1) or died(0) <br>\n3.Pclass : Passenger Class<br> \n4.Name : Passenger Name<br>\n5.Sex : Gender of Passenger<br>\n6.Age : Age of Passenger<br>\n7.SibSp : Number of Siblings/Spouses<br>\n8.Parch : Number of Parents/children<br>\n9.Ticket : Ticket number <br>\n10.Fare : Amount of money spent on ticket <br>\n11.Cabin : Cabin Category <br>\n12.Embarked : Port where passenger embarked( C=Cherbourg, Q= Queenstown, S= Southampton) <br>\n","metadata":{}},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:57:32.189506Z","iopub.execute_input":"2022-08-06T14:57:32.190225Z","iopub.status.idle":"2022-08-06T14:57:32.204941Z","shell.execute_reply.started":"2022-08-06T14:57:32.190180Z","shell.execute_reply":"2022-08-06T14:57:32.203671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* float64 (2):  Fare and Age\n* int64(5): Pclass, Survived, PassengerId, SibSp, Parch\n* object(5): Name, Sex, Ticket, Cabin, Embarked\n","metadata":{}},{"cell_type":"markdown","source":"<a id=3> </a> <br>\n### Univariate Variable Analysis\n* Categorical Variable: Survived, Sex, Pclass, Embarked, Cabin, Name, Ticket,Sibsp and Parch\n* Numerical Variable: Age, PassengerId, Fare","metadata":{}},{"cell_type":"markdown","source":"<a id=4> </a> <br>\n**Categorical Variable Analysis**\n","metadata":{}},{"cell_type":"code","source":"def bar_plot(variable):\n    \"\"\"\n    input: variable ex:\"Sex\"\n    output: bar plot & value count\n    \"\"\"\n    # get feature\n    var= train_df[variable]\n    # count number of categorical variable(value/sample)\n    varValue= var.value_counts()\n    #visualize\n    plt.figure(figsize=(9,3))\n    plt.bar(varValue.index, varValue)\n    plt.xticks(varValue.index, varValue.index.values)\n    plt.ylabel(\"Frequency\")\n    plt.title(variable)\n    plt.show()\n    print(\"{}: \\n {}\". format(variable,varValue))","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:57:32.206851Z","iopub.execute_input":"2022-08-06T14:57:32.207317Z","iopub.status.idle":"2022-08-06T14:57:32.214012Z","shell.execute_reply.started":"2022-08-06T14:57:32.207276Z","shell.execute_reply":"2022-08-06T14:57:32.213256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"category1= [\"Survived\",\"Sex\",\"Pclass\",\"Embarked\",\"SibSp\",\"Parch\"]\nfor i in category1:\n    bar_plot(i)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:57:32.215092Z","iopub.execute_input":"2022-08-06T14:57:32.215403Z","iopub.status.idle":"2022-08-06T14:57:33.226637Z","shell.execute_reply.started":"2022-08-06T14:57:32.215374Z","shell.execute_reply":"2022-08-06T14:57:33.225819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"category2= [\"Cabin\",\"Name\",\"Ticket\"]\nfor c in category2:\n    print(\"{} \\n\" .format(train_df[c].value_counts()))","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:57:33.228137Z","iopub.execute_input":"2022-08-06T14:57:33.229461Z","iopub.status.idle":"2022-08-06T14:57:33.242173Z","shell.execute_reply.started":"2022-08-06T14:57:33.229414Z","shell.execute_reply":"2022-08-06T14:57:33.240976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=5> </a> <br>\n**Numerical Variable Analysis**","metadata":{}},{"cell_type":"code","source":"def plot_hist(variable):\n    plt.figure(figsize=(9,3))\n    plt.hist(train_df[variable], bins=50)\n    plt.xlabel(variable)\n    plt.ylabel(\"Frequency\")\n    plt.title(variable + \"distribution with hist\")\n    plt.show","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:57:33.245087Z","iopub.execute_input":"2022-08-06T14:57:33.245428Z","iopub.status.idle":"2022-08-06T14:57:33.251324Z","shell.execute_reply.started":"2022-08-06T14:57:33.245397Z","shell.execute_reply":"2022-08-06T14:57:33.250211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"numericVar=[\"Fare\",\"Age\",\"PassengerId\"]\nfor n in numericVar:\n    plot_hist(n)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:57:33.252718Z","iopub.execute_input":"2022-08-06T14:57:33.253040Z","iopub.status.idle":"2022-08-06T14:57:34.060277Z","shell.execute_reply.started":"2022-08-06T14:57:33.253010Z","shell.execute_reply":"2022-08-06T14:57:34.059144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=6> </a> <br>\n## 3.Basic Data Analysis\n\n* Pclass - Survived\n* Sex - Survived\n* SibSp - Survived\n* Parch - Survived\n","metadata":{}},{"cell_type":"markdown","source":"* **Pclass vs Survived**","metadata":{}},{"cell_type":"code","source":"train_df[[\"Pclass\",\"Survived\"]].groupby([\"Pclass\"], as_index= False).mean().sort_values(by=\"Survived\", ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:57:34.062074Z","iopub.execute_input":"2022-08-06T14:57:34.062670Z","iopub.status.idle":"2022-08-06T14:57:34.076873Z","shell.execute_reply.started":"2022-08-06T14:57:34.062635Z","shell.execute_reply":"2022-08-06T14:57:34.075890Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* **Sex vs Survived**","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:34:36.788682Z","iopub.execute_input":"2022-08-06T14:34:36.789597Z","iopub.status.idle":"2022-08-06T14:34:36.795063Z","shell.execute_reply.started":"2022-08-06T14:34:36.789547Z","shell.execute_reply":"2022-08-06T14:34:36.793757Z"}}},{"cell_type":"code","source":"train_df[[\"Sex\",\"Survived\"]].groupby([\"Sex\"], as_index= False).mean().sort_values(by=\"Survived\", ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:57:34.080221Z","iopub.execute_input":"2022-08-06T14:57:34.080526Z","iopub.status.idle":"2022-08-06T14:57:34.095649Z","shell.execute_reply.started":"2022-08-06T14:57:34.080499Z","shell.execute_reply":"2022-08-06T14:57:34.094321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* **SibSp vs Survived**","metadata":{}},{"cell_type":"code","source":"train_df[[\"SibSp\",\"Survived\"]].groupby([\"SibSp\"], as_index= False).mean().sort_values(by=\"Survived\", ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:57:34.097132Z","iopub.execute_input":"2022-08-06T14:57:34.097989Z","iopub.status.idle":"2022-08-06T14:57:34.115079Z","shell.execute_reply.started":"2022-08-06T14:57:34.097930Z","shell.execute_reply":"2022-08-06T14:57:34.113759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* **Parch vs Survived**","metadata":{}},{"cell_type":"code","source":"train_df[[\"Parch\",\"Survived\"]].groupby([\"Parch\"], as_index= False).mean().sort_values(by=\"Survived\", ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:57:34.117055Z","iopub.execute_input":"2022-08-06T14:57:34.117890Z","iopub.status.idle":"2022-08-06T14:57:34.134038Z","shell.execute_reply.started":"2022-08-06T14:57:34.117847Z","shell.execute_reply":"2022-08-06T14:57:34.132393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=7> </a> <br>\n## 4.Outlier Detection","metadata":{}},{"cell_type":"code","source":"def detect_outliers(df,features):\n    outlier_indices=[]\n    for c in features:\n        #1st quartile\n        Q1= np.percentile(df[c],25)\n        #3rd quartile\n        Q3= np.percentile(df[c],75)\n        #IQR\n        IQR= Q3-Q1\n        #Outlier Step\n        outlier_step= IQR *1.5\n        #Detect Outlier and Their Indeces\n        outlier_list_col= df[(df[c]<Q1 - outlier_step) | (df[c]>Q3 + outlier_step)].index\n        #Store Indeces\n        outlier_indices.extend(outlier_list_col)\n    outlier_indices= Counter(outlier_indices) \n    multiple_outliers= list(i for i, v in outlier_indices.items() if v>2)\n    return multiple_outliers","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:57:34.135385Z","iopub.execute_input":"2022-08-06T14:57:34.136700Z","iopub.status.idle":"2022-08-06T14:57:34.145536Z","shell.execute_reply.started":"2022-08-06T14:57:34.136653Z","shell.execute_reply":"2022-08-06T14:57:34.144435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.loc[detect_outliers(train_df,[\"Age\",\"SibSp\",\"Parch\",\"Fare\"])]","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:57:34.146772Z","iopub.execute_input":"2022-08-06T14:57:34.147190Z","iopub.status.idle":"2022-08-06T14:57:34.178654Z","shell.execute_reply.started":"2022-08-06T14:57:34.147157Z","shell.execute_reply":"2022-08-06T14:57:34.177551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#drop outliers\ntrain_df=train_df.drop(detect_outliers(train_df,[\"Age\",\"SibSp\",\"Parch\",\"Fare\"]),axis=0).reset_index(drop=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:57:34.180349Z","iopub.execute_input":"2022-08-06T14:57:34.181301Z","iopub.status.idle":"2022-08-06T14:57:34.193347Z","shell.execute_reply.started":"2022-08-06T14:57:34.181256Z","shell.execute_reply":"2022-08-06T14:57:34.192256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=8> </a> <br>\n## 5. Missing Value\n* Find Missing Value   \n* Fill Missing Value","metadata":{}},{"cell_type":"code","source":"train_df_len=len(train_df)\ntrain_df=pd.concat([train_df,test_df],axis=0).reset_index(drop=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:57:34.195052Z","iopub.execute_input":"2022-08-06T14:57:34.195373Z","iopub.status.idle":"2022-08-06T14:57:34.205474Z","shell.execute_reply.started":"2022-08-06T14:57:34.195345Z","shell.execute_reply":"2022-08-06T14:57:34.204480Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=9> </a> <br>\n **Find Missing Value**","metadata":{}},{"cell_type":"code","source":"train_df.columns[train_df.isnull().any()]","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:58:10.973477Z","iopub.execute_input":"2022-08-06T14:58:10.973908Z","iopub.status.idle":"2022-08-06T14:58:10.984166Z","shell.execute_reply.started":"2022-08-06T14:58:10.973872Z","shell.execute_reply":"2022-08-06T14:58:10.982881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:58:37.442849Z","iopub.execute_input":"2022-08-06T14:58:37.443421Z","iopub.status.idle":"2022-08-06T14:58:37.453580Z","shell.execute_reply.started":"2022-08-06T14:58:37.443372Z","shell.execute_reply":"2022-08-06T14:58:37.452783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=10> </a> <br>\n **Fill Missing Value**\n* Embarked has 2 missing value\n* Fare has only 1","metadata":{}},{"cell_type":"code","source":"train_df[train_df[\"Embarked\"].isnull()]","metadata":{"execution":{"iopub.status.busy":"2022-08-06T15:00:22.090308Z","iopub.execute_input":"2022-08-06T15:00:22.091650Z","iopub.status.idle":"2022-08-06T15:00:22.112279Z","shell.execute_reply.started":"2022-08-06T15:00:22.091605Z","shell.execute_reply":"2022-08-06T15:00:22.111213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.boxplot(column=\"Fare\",by=\"Embarked\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T15:01:14.816160Z","iopub.execute_input":"2022-08-06T15:01:14.816702Z","iopub.status.idle":"2022-08-06T15:01:15.055617Z","shell.execute_reply.started":"2022-08-06T15:01:14.816652Z","shell.execute_reply":"2022-08-06T15:01:15.054494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[\"Embarked\"]= train_df[\"Embarked\"].fillna(\"C\")\ntrain_df[train_df[\"Embarked\"].isnull()]","metadata":{"execution":{"iopub.status.busy":"2022-08-06T15:02:47.381848Z","iopub.execute_input":"2022-08-06T15:02:47.382289Z","iopub.status.idle":"2022-08-06T15:02:47.396178Z","shell.execute_reply.started":"2022-08-06T15:02:47.382251Z","shell.execute_reply":"2022-08-06T15:02:47.395117Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[train_df[\"Fare\"].isnull()]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[\"Fare\"]= train_df[\"Fare\"].fillna(np.mean(train_df[train_df[\"Pclass\"]==3][\"Fare\"]))","metadata":{"execution":{"iopub.status.busy":"2022-08-06T15:04:43.581678Z","iopub.execute_input":"2022-08-06T15:04:43.582228Z","iopub.status.idle":"2022-08-06T15:04:43.591530Z","shell.execute_reply.started":"2022-08-06T15:04:43.582178Z","shell.execute_reply":"2022-08-06T15:04:43.590328Z"},"trusted":true},"execution_count":null,"outputs":[]}]}