{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<center><img src=\"https://i.pinimg.com/originals/7d/0d/9b/7d0d9b4c0214f0a9f074623e2974b500.gif\"></center>\n<h1><center>  <div style=\"background-color:skyblue;border-radius:10px; padding: 10px;\">Introduction 🚢</div></center></h1>\n\n>On April 15, 1912, during her maiden voyage, the widely considered “unsinkable” RMS Titanic sank after colliding with an iceberg. Unfortunately, there weren’t enough lifeboats for everyone onboard, resulting in the death of 1502 out of 2224 passengers and crew.\n\n>Many of us have come across this datasets in the start of Data Science journey, and this is a classic example to practice the skills of EDA and Machine Learning. This notebook is especially more useful for beginners in Data Science. It gives step by step guide for EDA and ML, after reading this notebook, I am sure you will be more comfortable in EDA and ML and you will be able to present your data with better visuals.","metadata":{}},{"cell_type":"markdown","source":"#### Loading the libraries","metadata":{}},{"cell_type":"code","source":"# For data handling\nimport numpy as np\nimport pandas as pd\n\n# For visvalization\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nimport plotly.express as px\nimport plotly.graph_objects as go\n\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-03-14T05:10:35.972747Z","iopub.execute_input":"2022-03-14T05:10:35.973222Z","iopub.status.idle":"2022-03-14T05:10:38.202328Z","shell.execute_reply.started":"2022-03-14T05:10:35.973093Z","shell.execute_reply":"2022-03-14T05:10:38.201124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Reading the files\ndf = pd.read_csv('../input/titanic/train.csv')\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:10:38.204286Z","iopub.execute_input":"2022-03-14T05:10:38.204698Z","iopub.status.idle":"2022-03-14T05:10:38.250692Z","shell.execute_reply.started":"2022-03-14T05:10:38.204657Z","shell.execute_reply":"2022-03-14T05:10:38.249710Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Number of rows and columns\ndf.shape","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:10:38.253454Z","iopub.execute_input":"2022-03-14T05:10:38.253853Z","iopub.status.idle":"2022-03-14T05:10:38.259780Z","shell.execute_reply.started":"2022-03-14T05:10:38.253821Z","shell.execute_reply":"2022-03-14T05:10:38.258647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3><center>  <div style=\"background-color:skyblue;border-radius:10px; padding: 10px;\">Data Cleaning 🚿</div></center></h3>\n\n> Data may have missing values, wrong data types, outliers, etc. So, before you do EDA on your data, please check for these things and if they are present, you have to clean the data before moving further otherwise you can get misleading and incorrect visuals. So let's fold the sleeves and get ready for cleaning!","metadata":{}},{"cell_type":"code","source":"# Dropping the unnecessary columns as they won't contribute for training ML model\ndf = df.drop(['PassengerId', 'Ticket', 'Name'], axis = 1)","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:10:38.261963Z","iopub.execute_input":"2022-03-14T05:10:38.262372Z","iopub.status.idle":"2022-03-14T05:10:38.271400Z","shell.execute_reply.started":"2022-03-14T05:10:38.262330Z","shell.execute_reply":"2022-03-14T05:10:38.270689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Checkign data types\ndf.info()","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:10:38.272420Z","iopub.execute_input":"2022-03-14T05:10:38.272847Z","iopub.status.idle":"2022-03-14T05:10:38.295135Z","shell.execute_reply.started":"2022-03-14T05:10:38.272803Z","shell.execute_reply":"2022-03-14T05:10:38.294169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Categorical features**: Sex, Embarked, and Pclass. \n\n**Numeric features**: Age, Fare. Discrete: SibSp, Parch.\n\n> Changing the data types of `Pclass` from numeric to category","metadata":{}},{"cell_type":"code","source":"# Changing the data types\ncat_col = ['Pclass']\ndf[cat_col] = df[cat_col].astype('object')","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:10:38.296340Z","iopub.execute_input":"2022-03-14T05:10:38.296622Z","iopub.status.idle":"2022-03-14T05:10:38.302421Z","shell.execute_reply.started":"2022-03-14T05:10:38.296594Z","shell.execute_reply":"2022-03-14T05:10:38.301729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Missing values\ndf.isnull().mean()*100","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:10:38.303456Z","iopub.execute_input":"2022-03-14T05:10:38.303903Z","iopub.status.idle":"2022-03-14T05:10:38.332678Z","shell.execute_reply.started":"2022-03-14T05:10:38.303845Z","shell.execute_reply":"2022-03-14T05:10:38.331975Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> If a colum has 70%-80% missing data, it is advisable to remove that column. `Cabin` has 77% missing data, so we will remove that **column**. `Age` and `Embarked` has <20% missing data, so we shall impute the missing values with mean or mode.\n\n> 📌 Some models can give you an error if missing values are present and some will show incorrect results, so it is important to remove them.","metadata":{}},{"cell_type":"code","source":"# Removing Cabin column\ndf=df.drop(['Cabin'], axis=1)\n\n# Imputing missing values\ndf['Age'] = df['Age'].fillna(df['Age'].median())\n\ndf['Embarked'] = df['Embarked'].fillna(df['Embarked'].mode()[0])","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:10:38.335142Z","iopub.execute_input":"2022-03-14T05:10:38.335603Z","iopub.status.idle":"2022-03-14T05:10:38.347378Z","shell.execute_reply.started":"2022-03-14T05:10:38.335562Z","shell.execute_reply":"2022-03-14T05:10:38.346600Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1><center>  <div style=\"background-color:skyblue;border-radius:10px; padding: 10px;\">Exploratory Data Analysis 📊</div></center></h1>\n\n> Before doing any kind of EDA, you should have a purpose of doing it. It should be clear what you want otherwise you will produce beautiful charts with no use. You should ask yourself, what you want to see and then choose the appropriate plot and features. So, I have mentioned the questions, which I thought before plotting the charts.\n\n> Note that, for a perticular question, there can be multiple charts. I have choosen the chart which I think was more appropriate. Now let's do some EDA.","metadata":{"execution":{"iopub.status.busy":"2021-06-10T05:57:46.053334Z","iopub.execute_input":"2021-06-10T05:57:46.053646Z","iopub.status.idle":"2021-06-10T05:57:46.069143Z","shell.execute_reply.started":"2021-06-10T05:57:46.053621Z","shell.execute_reply":"2021-06-10T05:57:46.067901Z"}}},{"cell_type":"code","source":"# Function for understanding the distribution of the columns\ndef quick_plot(df):\n\n    categorical_col = df.select_dtypes('object').columns.to_list()\n    \n    numeric_col = df.select_dtypes(['int', 'float']).columns.to_list()\n\n    if len(numeric_col)> len(categorical_col):\n        max_length = len(numeric_col)\n    else:\n        max_length = len(categorical_col)\n\n    fig, (ax1, ax2) = plt.subplots(2, max_length, figsize = (20, 10))\n\n    # Plotting for categorical columns\n    for axis, col in zip(ax1.ravel(), categorical_col):\n        axis = sns.countplot(x = col, data = df, ax = axis)\n        axis.set_ylabel(None)\n        axis.set_xlabel(col, size = 16)\n        \n        for i in axis.patches:    \n            axis.text(x = i.get_x() + i.get_width()/2, y = i.get_height()/2,\n                    s = f\"{round(i.get_height(),1)}\", \n                    ha = 'center', size = 16, weight = 'bold', rotation = 0, color = 'white')\n\n    # Plotting for numeric columns\n    for axis, col in zip(ax2.ravel(), numeric_col):\n        axis = sns.histplot(x = col, data = df, multiple = 'stack', ax = axis)\n        axis.set_ylabel(None)\n        axis.set_xlabel(col, size = 16)\n    \n        \n    fig.text(0.5, 1, 'The distribution of Data', ha = 'center', fontsize = 20);\n    plt.tight_layout()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-03-14T05:10:38.349208Z","iopub.execute_input":"2022-03-14T05:10:38.349614Z","iopub.status.idle":"2022-03-14T05:10:38.360411Z","shell.execute_reply.started":"2022-03-14T05:10:38.349576Z","shell.execute_reply":"2022-03-14T05:10:38.359362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Quickly understanding the distribution of the columns\nplt.style.use('seaborn')\nquick_plot(df)","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:10:38.362112Z","iopub.execute_input":"2022-03-14T05:10:38.362774Z","iopub.status.idle":"2022-03-14T05:10:40.249916Z","shell.execute_reply.started":"2022-03-14T05:10:38.362730Z","shell.execute_reply":"2022-03-14T05:10:40.248886Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 🔎Observation:\n> Passengers of class 3 are huge in number.\n\n> Very few passenger embarked on Q (Queenstown) port.\n\n> The distribution of `Fare` is skew.","metadata":{}},{"cell_type":"markdown","source":"<a id = \"1\"></a><br>\n<h3><center>  <div style=\"background-color:skyblue;border-radius:10px; padding: 10px;\">Distribution of Age and Fare 💰</div></center></h3>\n\n#### Question: How the features `age` and `fare` are distributed, are there any outliers ❓\n\n<h5><div class=\"alert alert-block alert-info\"> Beginner tip: It is applicable to numeric feature and used for understanding the data distribution and also for detecting outliers! </div></h5>","metadata":{}},{"cell_type":"code","source":"plt.style.use('seaborn')\nfig, (ax1, ax2) = plt.subplots(1, 2, figsize = (10, 6))\n\nax1 = sns.boxplot(y = 'Age', data = df, ax = ax1)\nax2 = sns.boxplot(y = 'Fare', data = df, ax = ax2)\n\nax1.set_ylabel(None)\nax1.set_xlabel(None)\nax1.set_title('Age', size = 16)\n\nax2.set_ylabel(None)\nax2.set_xlabel(None)\nax2.set_title('Fare', size = 16);","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-03-14T05:10:40.251375Z","iopub.execute_input":"2022-03-14T05:10:40.251970Z","iopub.status.idle":"2022-03-14T05:10:40.479712Z","shell.execute_reply.started":"2022-03-14T05:10:40.251924Z","shell.execute_reply":"2022-03-14T05:10:40.478920Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":">We have many values that are more than 75 percentile. The data is an outlier or not is the decision of the field expert. We can also tell for some of the data like `Age` - it can't be more than 100 or less than 0, etc. but in this case, I will remove these values while doing ML.","metadata":{}},{"cell_type":"markdown","source":"<a id = \"2\"></a><br>\n<h3><center>  <div style=\"background-color:skyblue;border-radius:10px; padding: 10px;\">Relation between Age and Fare 💰</div></center></h3>\n\n#### Question: Is there any relation between 'age' and 'fare'❓ \n\n<h5><div class=\"alert alert-block alert-info\"> Beginner tip: It is applicable when you have two numeric features. It's used for visvalizing relation between numeric features.</div></h5> ","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(1,1, figsize = (10, 6), constrained_layout = True)\nax = sns.regplot(x = 'Age', y = 'Fare', data = df)\n\nax.set_ylabel('Fare', fontsize = 14)\nax.set_xlabel('Age', fontsize = 14)\nplt.title('Age Vs Fare', fontsize = 16)\n\ncorrelation = np.corrcoef(df['Age'], df['Fare'])[0][1]\n\nax.text(x = 60, y = 180,\n        s = f\"correlation : {round(correlation,3)}\",\n        ha = 'center', size = 12, rotation = 0, color = 'black',\n        bbox=dict(boxstyle=\"round,pad=0.5\", fc='skyblue', ec=\"skyblue\", lw=2));","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-03-14T05:10:40.480743Z","iopub.execute_input":"2022-03-14T05:10:40.481162Z","iopub.status.idle":"2022-03-14T05:10:40.884033Z","shell.execute_reply.started":"2022-03-14T05:10:40.481131Z","shell.execute_reply":"2022-03-14T05:10:40.882995Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 🔍 Observations\n> The correlation between `Age` and `Fare` is 0.097, which is close to 0, so there is no relation between them. It means, we can't guess how much fare a passenger has paid just by looking at his/her age.\n\n#### 📝 Note\n\n> The correlation score:\n> - 1 : Strongly and positively correlated (one increases, other also increases and vice versa)\n> - 0 : No correlation\n> - -1 : Strongly and negetively correlated (one increases, other also decreases and vice versa)","metadata":{}},{"cell_type":"markdown","source":"<a id = \"3\"></a><br>\n<h3><center>  <div style=\"background-color:skyblue;border-radius:10px; padding: 10px;\">Analysis of Categorical Features 🔬</div></center></h3>\n\n> We have three categorical features`Pclass`, `Sex`, and `Embarked`. Not counting `Survived` as it is the target. Let's analyse the categorical features one by one based on the target.\n\n> For the analysis, I will be using PowerBI, as it is simple and efficient!","metadata":{}},{"cell_type":"markdown","source":"<h4><center>  <div style=\"background-color:pink;border-radius:10px; padding: 10px;\">Pclass (Ticket Class)</div></center></h4>\n\nThere are three classes 1st, 2nd, and 3rd.\n\n>- 1 = 1st\n>- 2 = 2nd\n>- 3 = 3rd","metadata":{}},{"cell_type":"markdown","source":"<center><img src = https://i.imgur.com/PXspAW3.jpg></center>","metadata":{}},{"cell_type":"markdown","source":"#### 🔎 Observations\n> As seen from previous plot also, Passengers of class 3 are huge in number. The reason can be cheap cost of the ticket.\n\n> 76% of the passengers form the 3rd class didn't survive while 63% of passengers form the 1st class survived.\n\n> The 1st class had mostly elders, may be because they can afford the ticket price and 3rd class had mostly young people.\n\n> Most of the rich 🤑 and old 👴👵 passengers survived.","metadata":{}},{"cell_type":"markdown","source":"<h4><center>  <div style=\"background-color:pink;border-radius:10px; padding: 10px;\">Gender 🤵💃</div></center></h4>","metadata":{}},{"cell_type":"markdown","source":"<center><img src = https://i.imgur.com/Zh1Ok1r.jpg></center>","metadata":{}},{"cell_type":"markdown","source":"#### 🔎 Observations\n> 81% men didn't survive while 74% women survived.\n\n> On average, female passengers paid higher fare than male passengers.","metadata":{}},{"cell_type":"markdown","source":"<h4><center>  <div style=\"background-color:pink;border-radius:10px; padding: 10px;\">Port of Embarkment 🚢</div></center></h4>\n\n> Passengers embarked from three different ports : C = Cherbourg, Q = Queenstown, S = Southampton","metadata":{}},{"cell_type":"markdown","source":"<center><img src = https://i.imgur.com/LSLpYqR.jpg></center>","metadata":{}},{"cell_type":"markdown","source":"#### 🔎 Observations\n> As seen previously also, very few passenger embarked on Q (Queenstown) port.\n\n> More than half of the passengers didn't survive who embarked from Queenstown (Q) and Southampton (S).\n\n> Passengers embarking from Cherbourg (C) paid more average fare, and also more than 50% of them survived.","metadata":{}},{"cell_type":"markdown","source":"<h4><center>  <div style=\"background-color:pink;border-radius:10px; padding: 10px;\">Comparison between Pclass, Gender 🤵💃, and Port of Embarkment 🚢</div></center></h4>\n\n<h5><center>  <div style=\"background-color:lightgreen;border-radius:10px; padding: 10px;\">1st Class Passengers who Survived</div></center></h5>","metadata":{}},{"cell_type":"markdown","source":"<center><img src = https://i.imgur.com/HGjag5k.jpg></center>","metadata":{}},{"cell_type":"markdown","source":"#### 🔎 Observations\n> Out of 55% of the passengers embarked from  Cherbourg (C) and survived, 35% belongs to 1st class. As the fare of 1st class is high so as the case with passengers embarking from Cherbourg (C), most of them belongs to 1st class.\n\n> Very few passengers from 1st class who survived, embarked from Queenstown (Q).","metadata":{}},{"cell_type":"markdown","source":"<h5><center>  <div style=\"background-color:lightgreen;border-radius:10px; padding: 10px;\">Survived passengers embarking from Southampton (S)</div></center></h5>","metadata":{}},{"cell_type":"markdown","source":"<center><img src = https://i.imgur.com/5yXaN27.jpg></center>","metadata":{}},{"cell_type":"markdown","source":"#### 🔎 Observations\n> Out of all female 74% survived and 45% of these embarked from Southampton (S).\n\n> 47% of 2nd class passengers survived and 41% of them embarked from Southampton (S).\n\n> Out of all male 19% survived and 13% of these embarked from Southampton (S). So, only 6% of the male who embarked from Cherbourg (C) and Queenstown (Q), survived. ","metadata":{}},{"cell_type":"markdown","source":"<h5><center>  <div style=\"background-color:lightgreen;border-radius:10px; padding: 10px;\">Class 3 passengers who didn't survive</div></center></h5>","metadata":{}},{"cell_type":"markdown","source":"<center><img src = https://i.imgur.com/1l8cw6b.jpg></center>","metadata":{}},{"cell_type":"markdown","source":"#### 🔎 Observations\n> Out of 61% passengers who didn't survived and embarked from Queenstown (Q), 58% belongs to class 3.\n\n> Mojority of the female passengers who didn't survived belong to class 3.","metadata":{}},{"cell_type":"markdown","source":"<h3><center>  <div style=\"background-color:skyblue;border-radius:10px; padding: 10px;\">Analysis of Numeric Features 🔬</div></center></h3>","metadata":{}},{"cell_type":"code","source":"fig, (ax1,ax2) = plt.subplots(1,2, figsize = (15, 6), constrained_layout = True)\nax1 = sns.boxenplot(x = 'Survived', y = 'Age', data = df, ax = ax1)\nax2 = sns.histplot(x = 'Age', data = df, hue = 'Survived', multiple = 'stack', kde = True, ax = ax2)\n\nax1.set_xlabel(None)\nax1.set_ylabel('Age', fontsize = 14)\nax1.set_title('Passenger Age and their Survival', fontsize = 16)\nax1.set_xticklabels(['Not Survived', 'Survived'], size = 14)\n\nax2.set_xlabel('Age', fontsize = 14)\nax2.set_ylabel('Number of Passengers', fontsize = 14)\nax2.set_title('Distribution of Age and Passenger Survival', fontsize = 16);","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-03-14T05:10:40.885461Z","iopub.execute_input":"2022-03-14T05:10:40.886084Z","iopub.status.idle":"2022-03-14T05:10:41.755498Z","shell.execute_reply.started":"2022-03-14T05:10:40.886037Z","shell.execute_reply":"2022-03-14T05:10:41.754389Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 🔎 Observations\n> There is not much relation between age and survival of the passengers. This also becomes clear from the kde plot, as the nature of both, the blue line (Non Survived) and gree line (Survived) is similar.","metadata":{}},{"cell_type":"code","source":"fig, (ax1,ax2) = plt.subplots(1,2, figsize = (15, 6), constrained_layout = True)\nax1 = sns.boxenplot(x = 'Survived', y = 'Fare', data = df, ax = ax1)\nax2 = sns.histplot(x = 'Fare', data = df, hue = 'Survived', multiple = 'stack', kde = True, ax = ax2)\n\nax1.set_xlabel(None)\nax1.set_ylabel('Fare', fontsize = 14)\nax1.set_title('Passenger Age and their Survival', fontsize = 16)\nax1.set_xticklabels(['Not Survived', 'Survived'], size = 14)\n\nax2.set_xscale('log')\nax2.set_xlabel('Age (Log scale)', fontsize = 14)\nax2.set_ylabel('Number of Passengers', fontsize = 14)\nax2.set_title('Distribution of Fare and Passenger Survival', fontsize = 16);","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-03-14T05:10:41.756738Z","iopub.execute_input":"2022-03-14T05:10:41.757044Z","iopub.status.idle":"2022-03-14T05:10:43.332040Z","shell.execute_reply.started":"2022-03-14T05:10:41.757017Z","shell.execute_reply":"2022-03-14T05:10:43.330829Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 🔎 Observations\n> There is some relation between `Fare` and `Survived` of the passengers. This also becomes clear from the kde plot, as the nature of both, the blue line (Non Survived) and gree line (Survived) is different.","metadata":{"_kg_hide-input":true}},{"cell_type":"markdown","source":"<h3><center>  <div style=\"background-color:skyblue;border-radius:10px; padding: 10px;\">Feature Enginner ⚙</div></center></h3>","metadata":{"execution":{"iopub.status.busy":"2021-07-29T11:53:54.739296Z","iopub.execute_input":"2021-07-29T11:53:54.739714Z","iopub.status.idle":"2021-07-29T11:53:54.751261Z","shell.execute_reply.started":"2021-07-29T11:53:54.739679Z","shell.execute_reply":"2021-07-29T11:53:54.749637Z"}}},{"cell_type":"markdown","source":"> The `Fare` feature had skew distribution, so I will log transform it.","metadata":{}},{"cell_type":"code","source":"df['Fare'] = df['Fare'].transform(np.log)\ndf['Fare'].min()","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:10:43.333527Z","iopub.execute_input":"2022-03-14T05:10:43.333825Z","iopub.status.idle":"2022-03-14T05:10:43.342426Z","shell.execute_reply.started":"2022-03-14T05:10:43.333796Z","shell.execute_reply":"2022-03-14T05:10:43.341431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# As new Fare contains -infinity, I will replace it will next minimum fare\nleast_value = df['Fare'].sort_values().unique()[1]\n\ndf['Fare'] = df['Fare'].replace(-np.inf, least_value)","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:10:43.343831Z","iopub.execute_input":"2022-03-14T05:10:43.344226Z","iopub.status.idle":"2022-03-14T05:10:43.352655Z","shell.execute_reply.started":"2022-03-14T05:10:43.344192Z","shell.execute_reply":"2022-03-14T05:10:43.351583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Creating bins for age\ndf['Age_Bins'] = pd.cut(df['Age'], [0, 10, 25, 40, 60,100], \n                        labels = ['<10', '10 to 25', '26 to 40', '41 to 60','60+'], include_lowest = True)","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:10:43.354148Z","iopub.execute_input":"2022-03-14T05:10:43.354452Z","iopub.status.idle":"2022-03-14T05:10:43.368197Z","shell.execute_reply.started":"2022-03-14T05:10:43.354423Z","shell.execute_reply":"2022-03-14T05:10:43.367108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3><center>  <div style=\"background-color:skyblue;border-radius:10px; padding: 10px;\">Making data ready for training</div></center></h3>","metadata":{}},{"cell_type":"code","source":"# Creating X (features) and y (target)\nX = df.drop(['Survived'], axis = 1)\ny = df['Survived'].astype('int')\n\n# One hot encoding for categorical features\nX_new = pd.get_dummies(X, drop_first = True)","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:10:43.369735Z","iopub.execute_input":"2022-03-14T05:10:43.370127Z","iopub.status.idle":"2022-03-14T05:10:43.393974Z","shell.execute_reply.started":"2022-03-14T05:10:43.370094Z","shell.execute_reply":"2022-03-14T05:10:43.392838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_new.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:10:43.395519Z","iopub.execute_input":"2022-03-14T05:10:43.395844Z","iopub.status.idle":"2022-03-14T05:10:43.413750Z","shell.execute_reply.started":"2022-03-14T05:10:43.395810Z","shell.execute_reply":"2022-03-14T05:10:43.412913Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.decomposition import PCA\nfrom sklearn.preprocessing import StandardScaler","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:12:06.976545Z","iopub.execute_input":"2022-03-14T05:12:06.976977Z","iopub.status.idle":"2022-03-14T05:12:06.981232Z","shell.execute_reply.started":"2022-03-14T05:12:06.976929Z","shell.execute_reply":"2022-03-14T05:12:06.980374Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scalar = StandardScaler()\nscaled_data = scalar.fit_transform(X_new)","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:12:28.035241Z","iopub.execute_input":"2022-03-14T05:12:28.035610Z","iopub.status.idle":"2022-03-14T05:12:28.047611Z","shell.execute_reply.started":"2022-03-14T05:12:28.035580Z","shell.execute_reply":"2022-03-14T05:12:28.046337Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# PCA\npca = PCA(0.9, random_state=99)\npca_df = pd.DataFrame(pca.fit_transform(scaled_data))\npca_df.index = X_new.index","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:13:35.057065Z","iopub.execute_input":"2022-03-14T05:13:35.057437Z","iopub.status.idle":"2022-03-14T05:13:35.080934Z","shell.execute_reply.started":"2022-03-14T05:13:35.057407Z","shell.execute_reply":"2022-03-14T05:13:35.079418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Inverting the PCA\nrestored_df = pd.DataFrame(pca.inverse_transform(pca_df), index=pca_df.index)","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:14:51.668803Z","iopub.execute_input":"2022-03-14T05:14:51.669216Z","iopub.status.idle":"2022-03-14T05:14:51.674509Z","shell.execute_reply.started":"2022-03-14T05:14:51.669184Z","shell.execute_reply":"2022-03-14T05:14:51.673249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Loss per feature\nloss_vertical = np.sum((np.array(X_new) - np.array(restored_df)) ** 2, axis=0)\n\nfeature_loss = pd.DataFrame(index=X_new.columns)\nfeature_loss['Loss'] = loss_vertical\n\nfeature_loss.sort_values(by='Loss')","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:16:18.050360Z","iopub.execute_input":"2022-03-14T05:16:18.050985Z","iopub.status.idle":"2022-03-14T05:16:18.064283Z","shell.execute_reply.started":"2022-03-14T05:16:18.050945Z","shell.execute_reply":"2022-03-14T05:16:18.063280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3><center>  <div style=\"background-color:skyblue;border-radius:10px; padding: 10px;\">Class Imbalance</div></center></h3>","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"code","source":"fig, ax = plt.subplots(1,1, figsize = (8, 6), constrained_layout = True)\nax = sns.countplot(x = 'Survived', data = df)\n\nax.set_ylabel('Number of Passengers', fontsize = 14)\nax.set_xlabel('')\nax.set_xticklabels(['Not Survived', 'Survived'], fontsize = 14)\nplt.title('Count of Passengers Survived', fontsize = 16)\n\nfor i in ax.patches:\n    ax.text(x = i.get_x() + i.get_width()/2, y = i.get_height()/4, \n            s = f\"{round(i.get_height()/len(df)*100,1)}%\", \n            ha = 'center', size = 50, weight = 'bold', rotation = 90, color = 'white')","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-03-14T05:10:43.415498Z","iopub.execute_input":"2022-03-14T05:10:43.415932Z","iopub.status.idle":"2022-03-14T05:10:43.630748Z","shell.execute_reply.started":"2022-03-14T05:10:43.415898Z","shell.execute_reply":"2022-03-14T05:10:43.629989Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Only 38% passengers suvived, so there is some class imbalance and so it has to be solved. \nSMOTE\nADSYN","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"code","source":"from imblearn.over_sampling import SMOTE, ADASYN\nadasyn = ADASYN()\nprint('Before sampling', y.value_counts())\n\nX_resample, y_resample = adasyn.fit_resample(X_new, y)\n\nprint('After sampling', y_resample.value_counts())","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-03-14T05:10:43.631997Z","iopub.execute_input":"2022-03-14T05:10:43.632498Z","iopub.status.idle":"2022-03-14T05:10:44.191074Z","shell.execute_reply.started":"2022-03-14T05:10:43.632464Z","shell.execute_reply":"2022-03-14T05:10:44.189996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> We got similar numbers of 0's and 1's.","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"code","source":"# Train - Test split\nfrom sklearn.model_selection import train_test_split\nX_train, X_test, y_train, y_test = train_test_split(X_resample, y_resample, train_size = 0.8,\n                                                    random_state = 99, stratify = y_resample)","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:10:44.192481Z","iopub.execute_input":"2022-03-14T05:10:44.193080Z","iopub.status.idle":"2022-03-14T05:10:44.202400Z","shell.execute_reply.started":"2022-03-14T05:10:44.193035Z","shell.execute_reply":"2022-03-14T05:10:44.201561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Scaling Data\n> Feature scaling is essential for machine learning algorithms that calculate distances between data. Since the range of values of raw data varies widely, in some machine learning algorithms, objective functions do not work correctly without normalization.\n\n> Though scaling is not required for tree based algorithms like Decision trees and Random Forest.","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import StandardScaler\nscalar = StandardScaler()\n\nX_train[['Age','Fare']] = scalar.fit_transform(X_train[['Age','Fare']])\nX_test[['Age','Fare']] = scalar.transform(X_test[['Age','Fare']])","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:10:44.206059Z","iopub.execute_input":"2022-03-14T05:10:44.206629Z","iopub.status.idle":"2022-03-14T05:10:44.221904Z","shell.execute_reply.started":"2022-03-14T05:10:44.206593Z","shell.execute_reply":"2022-03-14T05:10:44.221039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1><center>  <div style=\"background-color:skyblue;border-radius:10px; padding: 10px;\">Creating Machine Learning Models </div></center></h1>\n\n> First we will create simple model for baseline and then move to more complex models to improve further\n\n> For evaluation purpose I am using AUC ROC score which is mostly used for classification problems\n\n**AUC ROC Curve 📈**\n\n> Area Under Curve (AUC) Receiver Operating Characteristic (ROC) Curve The Area Under the Curve (AUC) is the measure of the ability of a classifier to distinguish between classes and is used as a summary of the ROC curve. The higher the AUC, the better the performance of the model at distinguishing between the positive and negative classes.","metadata":{}},{"cell_type":"markdown","source":"<h4><center>  <div style=\"background-color:orange;border-radius:10px; padding: 10px;\">Logistic Regression</div></center></h4>","metadata":{}},{"cell_type":"code","source":"from sklearn.linear_model import LogisticRegression\nfrom sklearn.metrics import roc_auc_score, confusion_matrix\n\nlog_model = LogisticRegression().fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:10:44.223538Z","iopub.execute_input":"2022-03-14T05:10:44.224018Z","iopub.status.idle":"2022-03-14T05:10:44.267620Z","shell.execute_reply.started":"2022-03-14T05:10:44.223968Z","shell.execute_reply":"2022-03-14T05:10:44.266426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Train AUC-ROC score for Logistic Regression is', roc_auc_score(y_train, log_model.predict_proba(X_train)[:, 1]))\nprint('Test AUC-ROC score for Logistic Regression is', roc_auc_score(y_test, log_model.predict_proba(X_test)[:, 1]))","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:10:44.269268Z","iopub.execute_input":"2022-03-14T05:10:44.269911Z","iopub.status.idle":"2022-03-14T05:10:44.290795Z","shell.execute_reply.started":"2022-03-14T05:10:44.269847Z","shell.execute_reply":"2022-03-14T05:10:44.289661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plotting the ROC Curve and finding optimal threshold\n# Function for plotting ROC curve, confurion matrix and finding optimal threshold\nfrom sklearn.metrics import roc_auc_score, plot_roc_curve, accuracy_score, recall_score, confusion_matrix\n\ndef my_roc_curve(model, accuracy_weight, title, X_train, y_train):\n    from sklearn import metrics\n    # Plotting AUC ROC curve\n    y_scores = model.predict(X_train)\n    fpr, tpr, thresholds = metrics.roc_curve(y_train, y_scores)\n    \n    fig, (ax1,ax2) = plt.subplots(1,2, figsize = (12, 6), constrained_layout = True)\n    ax1.plot(fpr, tpr, color='skyblue', label='ROC')\n    ax1.plot([0, 1], [0, 1], color='pink', linestyle='--')\n    ax1.text(x = 0.8, y = 0.3,\n            s = f\"AUC : {round(metrics.roc_auc_score(y_train, y_scores),2)}\",\n            ha = 'center', size = 12, rotation = 0, color = 'black',\n            bbox=dict(boxstyle=\"round,pad=0.5\", fc='skyblue', ec=\"skyblue\", lw=2));\n\n    ax1.set_xlabel('False Positive Rate', fontsize = 12)\n    ax1.set_ylabel('True Positive Rate', fontsize = 12)\n    ax1.set_title(f'ROC curve',  fontsize=16, y=1.05)\n\n    # Plotting optimal Confusion Matrix\n    from sklearn import metrics\n    probability = model.predict_proba(X_train)\n    matrix = pd.DataFrame()\n    \n    base_accuracy = metrics.accuracy_score(y_train, model.predict(X_train))\n    base_recall = metrics.recall_score(y_train, model.predict(X_train))\n    best_score = 0.6*base_accuracy + 0.4*base_recall\n    best_thresold = 0.5\n    \n    # Finding optimul thresold to maximize > accuracy_weight*threshold_accuracy + (1-accuracy_weight)*threshold_recall\n    for threshold in np.linspace(0, 1, 100):\n        y_predict = (probability>=threshold).astype(int)[:,1]\n        threshold_accuracy = metrics.accuracy_score(y_train, y_predict)\n        threshold_recall = metrics.recall_score(y_train, y_predict)\n        weighted_score = accuracy_weight*threshold_accuracy + (1-accuracy_weight)*threshold_recall\n        \n        if weighted_score>best_score:\n            best_thresold = threshold\n            best_score = weighted_score\n    \n    y_predict = (probability>=best_thresold).astype(int)[:,1]\n    matrix = metrics.confusion_matrix(y_train, y_predict)\n\n    ax2 = sns.heatmap(matrix, annot=True, fmt = '.0f', cbar=False, cmap='Blues',\n                        linewidths=3, square=True, ax = ax2, annot_kws={\"fontsize\":20})\n    ax2.set_title(f\"Confusion Matrics | Threshold : {round(best_thresold, 2)}\", fontsize=16, y=1.05);\n    ax2.set_xlabel('Predicted', fontsize=12)\n    ax2.set_ylabel('Actual', fontsize=12)\n    ax2.set_xticklabels([0,1], fontsize=12 )\n    ax2.set_yticklabels([0,1], fontsize=12, rotation=0)\n    \n    print('Optimal Threshold for Accuracy and Recall is : ', round(best_thresold, 2))\n    print(f\"Train Accuracy for  {title}: \", round(metrics.accuracy_score(y_train, y_predict),2), \n          f\"| Train Recall for {title}: \", round(metrics.recall_score(y_train, y_predict),2))\n    \n    plt.suptitle(f'{title}', fontsize=19, y=1.05)\n    \n    # For test data\n    probability = model.predict_proba(X_test)\n    y_predict = (probability>=best_thresold).astype(int)[:,1]\n    print(f\"Test Accuracy for {title}: \", round(metrics.accuracy_score(y_test, y_predict),2),\n         f\"| Test Recall for {title}: \", round(metrics.recall_score(y_test, y_predict),2))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-03-14T05:10:44.292699Z","iopub.execute_input":"2022-03-14T05:10:44.293440Z","iopub.status.idle":"2022-03-14T05:10:44.327177Z","shell.execute_reply.started":"2022-03-14T05:10:44.293392Z","shell.execute_reply":"2022-03-14T05:10:44.325908Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h4><center>  <div style=\"background-color:orange;border-radius:10px; padding: 10px;\">Decision Tree 🌳</div></center></h4>","metadata":{}},{"cell_type":"code","source":"from sklearn.tree import DecisionTreeClassifier\nfrom sklearn.model_selection import GridSearchCV\n\ntree = DecisionTreeClassifier()\nparameters = {'max_depth':[3,5,7], 'min_samples_leaf':[5,10,15]}\n\ntree_clf = GridSearchCV(tree, parameters, cv=5)\ntree_clf.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:10:44.329085Z","iopub.execute_input":"2022-03-14T05:10:44.329850Z","iopub.status.idle":"2022-03-14T05:10:44.679685Z","shell.execute_reply.started":"2022-03-14T05:10:44.329795Z","shell.execute_reply":"2022-03-14T05:10:44.678977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Train AUC-ROC score for Decision Tree is', roc_auc_score(y_train, tree_clf.predict_proba(X_train)[:, 1]))\nprint('Test AUC-ROC score for Decision Tree is', roc_auc_score(y_test, tree_clf.predict_proba(X_test)[:, 1]))","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:10:44.680951Z","iopub.execute_input":"2022-03-14T05:10:44.681389Z","iopub.status.idle":"2022-03-14T05:10:44.695373Z","shell.execute_reply.started":"2022-03-14T05:10:44.681343Z","shell.execute_reply":"2022-03-14T05:10:44.694400Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h4><center>  <div style=\"background-color:orange;border-radius:10px; padding: 10px;\">CatBoost 🐱‍👓</div></center></h4>","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"code","source":"from catboost import CatBoostClassifier\n\ncat_model = CatBoostClassifier(verbose = 100)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-03-14T05:10:44.696934Z","iopub.execute_input":"2022-03-14T05:10:44.697602Z","iopub.status.idle":"2022-03-14T05:10:44.858177Z","shell.execute_reply.started":"2022-03-14T05:10:44.697563Z","shell.execute_reply":"2022-03-14T05:10:44.857154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Creating validation dataset\nX_train_new, X_val, y_train_new, y_val = train_test_split(X_train, y_train, train_size = 0.8,\n                                                  random_state = 99, stratify = y_train)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-03-14T05:10:44.859305Z","iopub.execute_input":"2022-03-14T05:10:44.859578Z","iopub.status.idle":"2022-03-14T05:10:44.868518Z","shell.execute_reply.started":"2022-03-14T05:10:44.859552Z","shell.execute_reply":"2022-03-14T05:10:44.867454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_model.fit(X_train_new, y_train_new, eval_set= (X_val, y_val), early_stopping_rounds = 50, use_best_model = True)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-03-14T05:10:44.871308Z","iopub.execute_input":"2022-03-14T05:10:44.871625Z","iopub.status.idle":"2022-03-14T05:10:45.638342Z","shell.execute_reply.started":"2022-03-14T05:10:44.871596Z","shell.execute_reply":"2022-03-14T05:10:45.637353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Train AUC-ROC score for CatBoost is', roc_auc_score(y_train, cat_model.predict_proba(X_train)[:, 1]))\nprint('Test AUC-ROC score for CatBoost is', roc_auc_score(y_test, cat_model.predict_proba(X_test)[:, 1]))","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:10:45.639634Z","iopub.execute_input":"2022-03-14T05:10:45.639969Z","iopub.status.idle":"2022-03-14T05:10:45.655628Z","shell.execute_reply.started":"2022-03-14T05:10:45.639939Z","shell.execute_reply":"2022-03-14T05:10:45.654649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_model.feature_importances_","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:17:38.253705Z","iopub.execute_input":"2022-03-14T05:17:38.254431Z","iopub.status.idle":"2022-03-14T05:17:38.262239Z","shell.execute_reply.started":"2022-03-14T05:17:38.254373Z","shell.execute_reply":"2022-03-14T05:17:38.261012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"feature_imp = pd.DataFrame(index=X_val.columns)\nfeature_imp['Importance'] = cat_model.feature_importances_\n\nfeature_imp.sort_values(by='Importance', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:19:15.788812Z","iopub.execute_input":"2022-03-14T05:19:15.789384Z","iopub.status.idle":"2022-03-14T05:19:15.801713Z","shell.execute_reply.started":"2022-03-14T05:19:15.789349Z","shell.execute_reply":"2022-03-14T05:19:15.800766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"feature_loss.sort_values(by='Loss')","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:19:46.103070Z","iopub.execute_input":"2022-03-14T05:19:46.103469Z","iopub.status.idle":"2022-03-14T05:19:46.114360Z","shell.execute_reply.started":"2022-03-14T05:19:46.103436Z","shell.execute_reply":"2022-03-14T05:19:46.113054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h4><center>  <div style=\"background-color:orange;border-radius:10px; padding: 10px;\">Stacking 📚</div></center></h4>","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"code","source":"from sklearn.ensemble import StackingClassifier\nfrom sklearn.svm import SVC\nfrom sklearn.naive_bayes import GaussianNB\nfrom sklearn.pipeline import Pipeline","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-03-14T05:10:45.657051Z","iopub.execute_input":"2022-03-14T05:10:45.657473Z","iopub.status.idle":"2022-03-14T05:10:45.664860Z","shell.execute_reply.started":"2022-03-14T05:10:45.657439Z","shell.execute_reply":"2022-03-14T05:10:45.663817Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"estimators = [\n        ('tree',  DecisionTreeClassifier(max_depth=9, min_samples_leaf=10, random_state=99)),\n        ('svr', Pipeline([('scalar', StandardScaler()), ('svm', SVC(random_state = 99))])),\n        ('gauss', GaussianNB())]","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-03-14T05:10:45.666224Z","iopub.execute_input":"2022-03-14T05:10:45.666534Z","iopub.status.idle":"2022-03-14T05:10:45.673580Z","shell.execute_reply.started":"2022-03-14T05:10:45.666507Z","shell.execute_reply":"2022-03-14T05:10:45.672779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"staked_model = StackingClassifier(estimators, cv = 5, \n                                   final_estimator = cat_model)\nstaked_model.fit(X_train, y_train)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-03-14T05:10:45.674650Z","iopub.execute_input":"2022-03-14T05:10:45.674996Z","iopub.status.idle":"2022-03-14T05:10:47.382439Z","shell.execute_reply.started":"2022-03-14T05:10:45.674962Z","shell.execute_reply":"2022-03-14T05:10:47.381317Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Train AUC-ROC score for CatBoost is', roc_auc_score(y_train, staked_model.predict_proba(X_train)[:, 1]))\nprint('Test AUC-ROC score for CatBoost is', roc_auc_score(y_test, staked_model.predict_proba(X_test)[:, 1]))","metadata":{"execution":{"iopub.status.busy":"2022-03-14T05:10:47.383821Z","iopub.execute_input":"2022-03-14T05:10:47.384164Z","iopub.status.idle":"2022-03-14T05:10:47.451201Z","shell.execute_reply.started":"2022-03-14T05:10:47.384132Z","shell.execute_reply":"2022-03-14T05:10:47.450015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Let's try AutoML methods, to see if we can increase this further","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"markdown","source":"<h2><center>  <div style=\"background-color:pink;border-radius:10px; padding: 10px;\">AutoML 🤖</div></center></h2>\n\n<center> <img src= \"https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcTi48tkmw7TT50iGq2uJRuvRxg6Cz_aUdkkIpGZ5bVWpsugEobka8wWUFdfhqKKKupgR1Q&usqp=CAU\" ></center>\n\n> H2O is an open source library, it quite facinating and easy to use and I will be using this library for AutoML. So, let's get started !! 🚗","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"code","source":"# Importing h2o library\nimport h2o\nfrom h2o.automl import H2OAutoML","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-03-14T05:10:47.452492Z","iopub.execute_input":"2022-03-14T05:10:47.452784Z","iopub.status.idle":"2022-03-14T05:10:47.579629Z","shell.execute_reply.started":"2022-03-14T05:10:47.452756Z","shell.execute_reply":"2022-03-14T05:10:47.578737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Preprocessing for H2O\n> For training using h2o, we have to create a h2o frame, which is like a pandas dataframe. We have two options for doing this:\n>- read and preprocess the data using h2o frame\n>- read and preprocess the data using pandas and convert it to h2o frame\n\n> As I am more confirtable in handling the data with pandas so I am choosing the second method but if you feel confirtable in handling data with h2o frame, you can very well do that. We have already preprocessed the data using pycaret library so we will continue from that. From pycaret setup we get, X_train, y_train, X_test and y_test as pandas dataframe.\n\n> I didn't find any function to feed pandas dataframe to h2o, so first I convert these df to .csv and then read the .csv using h2o.import_file, to convert them into h2o frame.","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"code","source":"# Combining X and y\ndf_train = pd.concat([X_train_new, y_train_new], axis =1)\ndf_test = pd.concat([X_test, y_test], axis =1)\ndf_val = pd.concat([X_val, y_val], axis =1)\n\n# Saving as csv files\ndf_train.to_csv('train.csv', index = False)\ndf_test.to_csv('test.csv', index = False)\ndf_val.to_csv('val.csv', index = False)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-03-14T05:10:47.581044Z","iopub.execute_input":"2022-03-14T05:10:47.581480Z","iopub.status.idle":"2022-03-14T05:10:47.610570Z","shell.execute_reply.started":"2022-03-14T05:10:47.581436Z","shell.execute_reply":"2022-03-14T05:10:47.609288Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Initializing h2o\nh2o.init()","metadata":{"_kg_hide-output":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-03-14T05:10:47.611933Z","iopub.execute_input":"2022-03-14T05:10:47.612244Z","iopub.status.idle":"2022-03-14T05:10:54.378971Z","shell.execute_reply.started":"2022-03-14T05:10:47.612214Z","shell.execute_reply":"2022-03-14T05:10:54.378115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Reding the data using h2o frames\ntrain = h2o.import_file('./train.csv')\ntest = h2o.import_file('./test.csv')\nval = h2o.import_file('./val.csv')","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-03-14T05:10:54.382148Z","iopub.execute_input":"2022-03-14T05:10:54.382497Z","iopub.status.idle":"2022-03-14T05:10:55.699027Z","shell.execute_reply.started":"2022-03-14T05:10:54.382461Z","shell.execute_reply":"2022-03-14T05:10:55.697851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Identifing predictors and response\nx = train.columns\ny = \"Survived\"\nx.remove(y)\n\n# For binary classification, response should be a factor\ntrain[y] = train[y].asfactor()\ntest[y] = test[y].asfactor()\nval[y] = val[y].asfactor()","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-03-14T05:10:55.700659Z","iopub.execute_input":"2022-03-14T05:10:55.701351Z","iopub.status.idle":"2022-03-14T05:10:55.708929Z","shell.execute_reply.started":"2022-03-14T05:10:55.701303Z","shell.execute_reply":"2022-03-14T05:10:55.708039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Run AutoML for 20 base models (limited to 1 hour max runtime by default)\naml = H2OAutoML(max_models=10, seed=1)\naml.train(x=x, y=y, training_frame=train, validation_frame = val)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-03-14T05:10:55.710426Z","iopub.execute_input":"2022-03-14T05:10:55.711131Z","iopub.status.idle":"2022-03-14T05:11:29.026571Z","shell.execute_reply.started":"2022-03-14T05:10:55.711084Z","shell.execute_reply":"2022-03-14T05:11:29.025322Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# View the AutoML Leaderboard\nlb = aml.leaderboard\nlb.head(rows=10)  # Print all rows instead of default (10 rows)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-03-14T05:11:29.028312Z","iopub.execute_input":"2022-03-14T05:11:29.028728Z","iopub.status.idle":"2022-03-14T05:11:29.139221Z","shell.execute_reply.started":"2022-03-14T05:11:29.028686Z","shell.execute_reply":"2022-03-14T05:11:29.138273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> With H2O AutoML also we got AUC of 90%, which is  closer to than what we obtained with CatBoost for test set.","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"code","source":"perf = aml.leader.model_performance(test)\nperf","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-03-14T05:11:29.140767Z","iopub.execute_input":"2022-03-14T05:11:29.141181Z","iopub.status.idle":"2022-03-14T05:11:29.515272Z","shell.execute_reply.started":"2022-03-14T05:11:29.141140Z","shell.execute_reply":"2022-03-14T05:11:29.514264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Conclusion:\n\n> The data analysis shows that the rich, old and woman mostly survived\n\n> The ML model using H2O.ai gives AUC ROC score of 0.91 and accuracy of 88%","metadata":{}},{"cell_type":"markdown","source":"<h3><center>  <div style=\"background-color:pink;border-radius:10px; padding: 10px;\">If you like, don't forget to upvote!</div></center></h3>","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}}]}