{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n\nimport numpy as np \nimport pandas as pd \nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\n# finding data set locale\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-18T23:13:46.891471Z","iopub.execute_input":"2022-07-18T23:13:46.892105Z","iopub.status.idle":"2022-07-18T23:13:46.900960Z","shell.execute_reply.started":"2022-07-18T23:13:46.892066Z","shell.execute_reply":"2022-07-18T23:13:46.900050Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = pd.read_csv(\"/kaggle/input/titanic/train.csv\")\ntest_df = pd.read_csv(\"/kaggle/input/titanic/test.csv\")\ncombined = [train_df, test_df] ","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:46.946803Z","iopub.execute_input":"2022-07-18T23:13:46.947118Z","iopub.status.idle":"2022-07-18T23:13:46.961636Z","shell.execute_reply.started":"2022-07-18T23:13:46.947087Z","shell.execute_reply":"2022-07-18T23:13:46.960763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Deep Neural Network Implementation for Predicting Titanic Passenger Survivorship\n#### Written by Jonathan Ma - 06/17/2022","metadata":{}},{"cell_type":"markdown","source":"Welcome to my data analysis and modeling notebook for predicting survivorship of passengers on the Titanic. Briefly, let's discuss some important facts about the Titanic.\n\n- The Titanic was a British passenger liner which struck an iceberg on April 15, 1912, and sunk to the ocean floor. \n- Aboard the Titanic were lifeboats with an estimated capacity of 1,178 people — massively insufficient for the estimated 2,224 passengers and crew aboard.\n- When deciding who should be given priority for boarding the lifeboats, **women and infants were given priority.**\n- Tickets for the Titanic's maiden voyage were stratified into three classes: **first, second, and third class, with increasing corresponding price.**\n\n#### Let's begin by previewing our data, and defining our variables.","metadata":{}},{"cell_type":"code","source":"train_df.tail(5)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:47.051290Z","iopub.execute_input":"2022-07-18T23:13:47.051919Z","iopub.status.idle":"2022-07-18T23:13:47.068326Z","shell.execute_reply.started":"2022-07-18T23:13:47.051884Z","shell.execute_reply":"2022-07-18T23:13:47.067310Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Dictionary\n| Variable | Definition | Key |\n| --- | --- | --- |\n| survival | Survival | 0 = No, 1 = Yes |\n| pclass | Ticket class | 1 = 1st, 2 = 2nd, 3 = 3rd |\n| sex | Sex |  |\n| Age | Age in years |  |\n| sibsp | # of siblings / spouses aboard the Titanic |  |\n| parch | # of parents / children aboard the Titanic |  |\n| ticket | Ticket number |  | \n| fare | Passenger fare |  |\n| cabin | Cabin number |  |\n| embarked | Port of Embarkation | C = Cherbourg, Q = Queenstown, S = Southampton |\n\n\n# Classifying Data\nWe will begin our data analysis by classifying our variables\n#### Categorical Data\n- Ordinal Data: pclass\n- Nominal Data: embarked, survived, sex, name\n\n\n#### Numerical Data\n- Discrete Data: age, sibsp, parch\n- Continuous Data: fare\n\n\n#### Mixed Data\n- cabin: Latin letter followed by a number\n- ticket: Mix of numbers and Latin strings followed by numbers\n\nNow that we have an idea of what our variables mean, let's move on to identifying the state of our dataset.\n    ","metadata":{}},{"cell_type":"code","source":"train_df.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:47.134400Z","iopub.execute_input":"2022-07-18T23:13:47.135086Z","iopub.status.idle":"2022-07-18T23:13:47.166364Z","shell.execute_reply.started":"2022-07-18T23:13:47.135054Z","shell.execute_reply":"2022-07-18T23:13:47.165270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.describe(include='O')","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:47.203518Z","iopub.execute_input":"2022-07-18T23:13:47.203787Z","iopub.status.idle":"2022-07-18T23:13:47.225316Z","shell.execute_reply.started":"2022-07-18T23:13:47.203762Z","shell.execute_reply":"2022-07-18T23:13:47.224473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looks like we are missing some age and embarkement location labels. Additionally, it looks like our cabin number variable is unusable — we are missing too many values. We will deal with these null values when we clean our data, but for now, let's find some correlations in our data.","metadata":{}},{"cell_type":"markdown","source":"### Analysis","metadata":{}},{"cell_type":"markdown","source":"We can posit that a higher passenger class led to higher survival rates.","metadata":{}},{"cell_type":"code","source":"train_df[[\"Pclass\", \"Survived\"]].groupby([\"Pclass\"], as_index=False).mean()","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:47.271723Z","iopub.execute_input":"2022-07-18T23:13:47.272058Z","iopub.status.idle":"2022-07-18T23:13:47.287426Z","shell.execute_reply.started":"2022-07-18T23:13:47.272029Z","shell.execute_reply":"2022-07-18T23:13:47.286319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And of course, females were more likely to survive the sinking, being given priority boarding to the lifeboats.","metadata":{}},{"cell_type":"code","source":"train_df[[\"Sex\",\"Survived\"]].groupby([\"Sex\"], as_index=False).mean()","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:47.339298Z","iopub.execute_input":"2022-07-18T23:13:47.339641Z","iopub.status.idle":"2022-07-18T23:13:47.354652Z","shell.execute_reply.started":"2022-07-18T23:13:47.339611Z","shell.execute_reply":"2022-07-18T23:13:47.353231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's see if ticket fare played a role in survival rate by using a histogram to visualize our data.","metadata":{}},{"cell_type":"code","source":"fare_histogram = sns.FacetGrid(train_df, col = \"Survived\")\nfare_histogram.map(plt.hist, \"Fare\", bins=25)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:47.404229Z","iopub.execute_input":"2022-07-18T23:13:47.404778Z","iopub.status.idle":"2022-07-18T23:13:47.833230Z","shell.execute_reply.started":"2022-07-18T23:13:47.404747Z","shell.execute_reply":"2022-07-18T23:13:47.832317Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can do the same with passenger age.","metadata":{}},{"cell_type":"code","source":"age_histogram = sns.FacetGrid(train_df, col = \"Survived\")\nage_histogram.map(plt.hist, \"Age\", bins = 20)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:47.835600Z","iopub.execute_input":"2022-07-18T23:13:47.836326Z","iopub.status.idle":"2022-07-18T23:13:48.238721Z","shell.execute_reply.started":"2022-07-18T23:13:47.836285Z","shell.execute_reply":"2022-07-18T23:13:48.237831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally, as port of embarkment reveal geographical differences of wealth, we can assume that it may also have had a direct impact on survival rate.","metadata":{}},{"cell_type":"code","source":"train_df[[\"Embarked\", \"Survived\"]].groupby([\"Embarked\"], as_index=False).mean()","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:48.240029Z","iopub.execute_input":"2022-07-18T23:13:48.240458Z","iopub.status.idle":"2022-07-18T23:13:48.257533Z","shell.execute_reply.started":"2022-07-18T23:13:48.240419Z","shell.execute_reply":"2022-07-18T23:13:48.256730Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"With our preliminary data analysis, we have that:\n\n\n- Gender played a vital role in survivorship in the Titanic (74% of females survived vs. ~19% of males)\n- Passenger class also heavily affected survivorship (63% of First class passengers survived vs. 24% in Third class)\n- A lower fare points towards to a lower survival rate.\n- Infants in particular had a greater chance of survival on the Titanic.\n- Port of embarkment affected survivorship.\n\nLet's work on condensing our variables into usable features for our model.\n","metadata":{}},{"cell_type":"markdown","source":"# Feature Engineering\n\n### Titles\nDuring the early 20th century, certain titles were a rare commodity amongst commmon folk. Typically, titles such as Dr. or Sir. were reserved for higher class members of society. We will take advantage of the fact that our data set contains titles, and create a new input feature named \"Title.\" Luckily, this will only involve some quick globbing, as the titles all end with periods.","metadata":{}},{"cell_type":"code","source":"for i in combined:\n    i[\"Title\"] = i[\"Name\"].str.extract(' ([A-Za-z]+)\\.', expand=False)\npd.crosstab(train_df[\"Title\"], train_df[\"Sex\"])","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:48.260111Z","iopub.execute_input":"2022-07-18T23:13:48.260599Z","iopub.status.idle":"2022-07-18T23:13:48.284251Z","shell.execute_reply.started":"2022-07-18T23:13:48.260563Z","shell.execute_reply":"2022-07-18T23:13:48.283375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can immediately see that the majority of these titles fall under commoner titles (Miss., Mr., Mrs., etc.), and we have a good number of ~~Jedi~~ \"Masters\" aboard the Titanic. We will go ahead and group together the rarer titles in a category called \"Exotic\". We will also clean up the data by anglicizing the titles.","metadata":{}},{"cell_type":"code","source":"for i in combined:\n    i[\"Title\"] = i[\"Title\"].replace([\"Lady\", \"Countess\", \"Capt\", \"Col\",\n                                     \"Don\", \"Dr\", \"Major\", \"Rev\", \"Sir\",\n                                     \"Jonkheer\", \"Dona\"], \"Exotic\")\n    i[\"Title\"] = i[\"Title\"].replace(\"Mlle\", \"Miss\")\n    i[\"Title\"] = i[\"Title\"].replace(\"Ms\", \"Miss\")\n    i[\"Title\"] = i[\"Title\"].replace(\"Mme\", \"Mrs\")\n    \ntrain_df[[\"Title\", \"Survived\"]].groupby([\"Title\"], as_index=False).mean()","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:48.285583Z","iopub.execute_input":"2022-07-18T23:13:48.285917Z","iopub.status.idle":"2022-07-18T23:13:48.308609Z","shell.execute_reply.started":"2022-07-18T23:13:48.285883Z","shell.execute_reply":"2022-07-18T23:13:48.307560Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As expected, the feminine titles had higher survival rates, and the more \"High Society\" titles had higher survival rates than commoner titles. We will convert these categorical titles to ordinal titles by survival rate.","metadata":{}},{"cell_type":"code","source":"title_map = {\"Mr\": 1, \"Exotic\": 2, \"Master\": 3, \"Miss\": 4, \"Mrs\": 5}\nfor i in combined:\n    i[\"Title\"] = i[\"Title\"].map(title_map)\n    i[\"Title\"] = i[\"Title\"].fillna(0)\ntrain_df.tail()","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:48.310035Z","iopub.execute_input":"2022-07-18T23:13:48.310376Z","iopub.status.idle":"2022-07-18T23:13:48.329496Z","shell.execute_reply.started":"2022-07-18T23:13:48.310341Z","shell.execute_reply":"2022-07-18T23:13:48.328627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The data looks a little unwieldy right now, so let's drop some unneeded variables. We don't need name anymore, as we've extracted the only meaningful part of the variable. As stated previously, Cabin number is utterly incomplete, so we will drop that variable. Ticket number should not contribute meaningfully to our model, as it's a unique identifier. Finally, passenger ID is used purely for submission purposes, so let's drop that for now.","metadata":{}},{"cell_type":"code","source":"train_df = train_df.drop([\"Name\", \"PassengerId\", \"Ticket\", \"Cabin\"], axis=1)\ntest_df = test_df.drop([\"Name\", \"Ticket\", \"Cabin\"], axis=1)\ncombined = [train_df, test_df]\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:48.331027Z","iopub.execute_input":"2022-07-18T23:13:48.331445Z","iopub.status.idle":"2022-07-18T23:13:48.352129Z","shell.execute_reply.started":"2022-07-18T23:13:48.331409Z","shell.execute_reply":"2022-07-18T23:13:48.351470Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Much better. Let's also map Sex onto a boolean function of Female or not Female.","metadata":{}},{"cell_type":"code","source":"for i in combined:\n    i[\"Sex\"] = i[\"Sex\"].map({\"female\": 1, \"male\":0}).astype('int')\n    \ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:48.353062Z","iopub.execute_input":"2022-07-18T23:13:48.353306Z","iopub.status.idle":"2022-07-18T23:13:48.372671Z","shell.execute_reply.started":"2022-07-18T23:13:48.353283Z","shell.execute_reply":"2022-07-18T23:13:48.371848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Completing Null Entries\nEarlier, we noticed that some of the Age values are missing. Let's generate some random numbers between the mean and the standard deviation to quickly fill those missing entries.","metadata":{}},{"cell_type":"code","source":"age_mean = int(train_df[\"Age\"].mean())\nage_std = int(train_df[\"Age\"].std())\n\nfor i in combined:\n    i.loc[i.Age.isnull(), 'Age'] = np.random.randint(\n        age_std, age_mean)\n    i[\"Age\"] = i[\"Age\"].astype('int')\n\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:48.373863Z","iopub.execute_input":"2022-07-18T23:13:48.374287Z","iopub.status.idle":"2022-07-18T23:13:48.393833Z","shell.execute_reply.started":"2022-07-18T23:13:48.374248Z","shell.execute_reply":"2022-07-18T23:13:48.393137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Onto our Sibling/Spouse count and Parents/Children count variables. These variables are a mess. Let's condense them into one variable: family size. Then, let's see if we can spot some correlation between family size and survival rate.","metadata":{}},{"cell_type":"code","source":"for i in combined:\n    i[\"FamilySize\"] = i[\"SibSp\"] + i[\"Parch\"] + 1\n\ntrain_df[[\"FamilySize\", \"Survived\"]].groupby([\"FamilySize\"], as_index=False).mean().sort_values(by=\"Survived\", ascending=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:48.396235Z","iopub.execute_input":"2022-07-18T23:13:48.397065Z","iopub.status.idle":"2022-07-18T23:13:48.413806Z","shell.execute_reply.started":"2022-07-18T23:13:48.397032Z","shell.execute_reply":"2022-07-18T23:13:48.412856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It looks like generally there is little correlation between survivorship and family size. Let's dig a little deeper and see if there is correlation between passengers who were alone and passengers who boarded the Titanic with a family.","metadata":{}},{"cell_type":"code","source":"for i in combined:\n    i[\"Alone\"] = 0\n    i.loc[i[\"FamilySize\"] == 1, \"Alone\"] = 1\n\ntrain_df[[\"Alone\", \"Survived\"]].groupby([\"Alone\"], as_index=False).mean()","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:48.415213Z","iopub.execute_input":"2022-07-18T23:13:48.415626Z","iopub.status.idle":"2022-07-18T23:13:48.433007Z","shell.execute_reply.started":"2022-07-18T23:13:48.415589Z","shell.execute_reply":"2022-07-18T23:13:48.432176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here, we can see that there is some correlation between boarding the Titanic independently and survival. We will keep this final feature and drop the Sibling Count and Parent Count variables.","metadata":{}},{"cell_type":"code","source":"train_df = train_df.drop([\"Parch\", \"SibSp\", \"FamilySize\"], axis=1)\ntest_df = test_df.drop([\"Parch\", \"SibSp\", \"FamilySize\"], axis=1)\ncombined = [train_df, test_df]\n\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:48.434310Z","iopub.execute_input":"2022-07-18T23:13:48.434677Z","iopub.status.idle":"2022-07-18T23:13:48.452807Z","shell.execute_reply.started":"2022-07-18T23:13:48.434641Z","shell.execute_reply":"2022-07-18T23:13:48.452125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally, let's sort embarkment ports by survival rate, and refactor Embarked as ordinal data. In this step, we will also complete our dataset by replacing the two missing embarkment entries with the most common embarkment point.","metadata":{}},{"cell_type":"code","source":"for i in combined:\n    i[\"Embarked\"] = i[\"Embarked\"].fillna(train_df.Embarked.dropna().mode()[0])\ntrain_df[[\"Embarked\", \"Survived\"]].groupby([\"Embarked\"], as_index=False).mean()","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:48.453875Z","iopub.execute_input":"2022-07-18T23:13:48.454295Z","iopub.status.idle":"2022-07-18T23:13:48.471849Z","shell.execute_reply.started":"2022-07-18T23:13:48.454259Z","shell.execute_reply":"2022-07-18T23:13:48.471151Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in combined:\n    i[\"Embarked\"] = i[\"Embarked\"].map({\"S\": 0, \"Q\": 1, \"C\": 2}).astype('int')\ntrain_df.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:48.473023Z","iopub.execute_input":"2022-07-18T23:13:48.473538Z","iopub.status.idle":"2022-07-18T23:13:48.491181Z","shell.execute_reply.started":"2022-07-18T23:13:48.473497Z","shell.execute_reply":"2022-07-18T23:13:48.490375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally, this is how our cleaned data set looks like.\n\n# L-Layer Neural Network\nNow that we're done with feature engineering, we can construct our neural network. We're going to do this from scratch using Numpy.\n\n### ReLU, Sigmoid Functions\nBecause we are doing a simple binary classification, we will be using the ReLU and Sigmoid functions as our activation functions. ","metadata":{}},{"cell_type":"code","source":"def sigmoid(Z: np.ndarray):\n    return 1/(1+np.exp(-Z)), Z\ndef relu(Z: np.ndarray):\n    return np.maximum(0, Z), Z\n\n# Our activation helper functions also return a cache value we will use for backpropagation","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:48.492387Z","iopub.execute_input":"2022-07-18T23:13:48.492748Z","iopub.status.idle":"2022-07-18T23:13:48.498558Z","shell.execute_reply.started":"2022-07-18T23:13:48.492713Z","shell.execute_reply":"2022-07-18T23:13:48.497739Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Forward Propagation\nLet's write a few helper function to facilitate forward propagation.","metadata":{}},{"cell_type":"code","source":"def linear_forward(A, W, b):\n    \"\"\"\n    Implements linear part of a layer's forward propagation\n    Args:\n    A: Activations from previous layer\n    W: Weight matrices\n    b: bias vector\n    Returns:\n    Z: input of the activation function -- pre-activation parameter\n    cache -- used for backward propagation\n    \"\"\"\n\n    Z = np.dot(W,A)+b\n    cache = (A,W,b)\n    return Z, cache\n\ndef linear_activation_forward(A_prev, W, b, activation: str):\n    \"\"\"\n    Implements forward propagation for the linear->activation layer\n    Args:\n    A_prev: Activations from previous layer\n    W: Weight matrices\n    b: bias vector\n    activation: the activation function to be used in this layer, represented by a string\n    Returns:\n    A: output of the activation function -- post-activation value\n    cache -- used for backward propagation\n    \"\"\" \n    if activation == \"sigmoid\":\n        Z, linear_cache = linear_forward(A_prev, W, b)\n        A, activation_cache = sigmoid(Z)\n    else:\n        Z, linear_cache = linear_forward(A_prev, W, b)\n        A, activation_cache = relu(Z)\n    \n    cache = (linear_cache, activation_cache)\n    \n    return A, cache","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:48.500783Z","iopub.execute_input":"2022-07-18T23:13:48.501252Z","iopub.status.idle":"2022-07-18T23:13:48.510099Z","shell.execute_reply.started":"2022-07-18T23:13:48.501215Z","shell.execute_reply":"2022-07-18T23:13:48.509253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"With these functions written, we can write a function which implements forward propagation.","metadata":{}},{"cell_type":"code","source":"def nn_forward(X, parameters):\n    \"\"\"\n    Implements forward propagation for Linear->ReLU *(L-1) -> Linear-> Sigmoid computation\n    Args:\n    X: Input data\n    parameters: dictionary containing all the components necessary for forward propagation\n    Returns:\n    AL: Activation value from the output layer\n    caches: list of caches containing every cache of linear_activation_forward()\n    \"\"\"\n    caches = []\n    A = X\n    L = len(parameters)//2\n    \n    for i in range(1,L):\n        A_prev = A\n        A, cache = linear_activation_forward(A_prev, parameters[\"W\"+str(i)], parameters[\"b\"+str(i)], \"relu\")\n        caches.append(cache)\n    AL, cache = linear_activation_forward(A, parameters[\"W\"+str(L)], parameters[\"b\"+str(L)], \"sigmoid\")\n    caches.append(cache)\n    return AL, caches","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:48.526887Z","iopub.execute_input":"2022-07-18T23:13:48.527174Z","iopub.status.idle":"2022-07-18T23:13:48.535582Z","shell.execute_reply.started":"2022-07-18T23:13:48.527147Z","shell.execute_reply":"2022-07-18T23:13:48.534855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Cost Function\nLet's write a cost function: we will use this formula:\n$$-\\frac{1}{m} \\sum\\limits_{i = 1}^{m} (y^{(i)}\\log\\left(a^{[L] (i)}\\right) + (1-y^{(i)})\\log\\left(1- a^{[L](i)}\\right))$$","metadata":{}},{"cell_type":"code","source":"def compute_cost(AL, Y):\n    \"\"\"\n    Computes cost of forward propagation.\n    Args:\n    AL: probability vector -- corresponds to label predictions\n    Y: label vector\n    Returns:\n    cost: vector containing computed cost values\n    \"\"\"\n    m = Y.shape[1]\n    cost = np.dot(-1/m, np.sum(np.multiply(np.log(AL+1e-7),Y) + np.multiply(np.log(1-AL+1e-7), 1-Y)))\n    cost = np.squeeze(cost)\n    return cost","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:48.588987Z","iopub.execute_input":"2022-07-18T23:13:48.589259Z","iopub.status.idle":"2022-07-18T23:13:48.595274Z","shell.execute_reply.started":"2022-07-18T23:13:48.589235Z","shell.execute_reply":"2022-07-18T23:13:48.594536Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Backward Propagation\nWe're almost done with creating our neural network. Let's write some helper functions to tackle backward propagation. \n\nThe partial derivatives with respect to Z (weighted sum of the activations shifted by a bias) of the activation functions are as follows:\n\n\n### ReLU\n$$\n\\begin{equation}\n\\partial{J}{\\vec{Z}^{[l]}} = \\partial{J}{\\vec{A}^{[l]}} \\odot I(\\vec{Z}^{[l]} > 0),\n\\end{equation}$$\n\nwhere $\\odot$ denotes element-wise mulitplication.\n\n### Sigmoid\n$$\n\\begin{equation}\n\\partial{J}{\\vec{Z}^{[l]}} = \\partial{J}{\\vec{A}^{[l]}} \\odot \\vec{A}^{[l]} \\odot (1 - \\vec{A}^{[l]}).\n\\end{equation}\n$$\n\nWe will also be using the following formulas to write our backward propagation functions. \n$$ dW^{[l]} = \\frac{\\partial \\mathcal{J} }{\\partial W^{[l]}} = \\frac{1}{m} dZ^{[l]} A^{[l-1] T}$$\n$$ db^{[l]} = \\frac{\\partial \\mathcal{J} }{\\partial b^{[l]}} = \\frac{1}{m} \\sum_{i = 1}^{m} dZ^{[l](i)}$$\n$$ dA^{[l-1]} = \\frac{\\partial \\mathcal{L} }{\\partial A^{[l-1]}} = W^{[l] T} dZ^{[l]}$$","metadata":{}},{"cell_type":"code","source":"def d_relu(dA, activation_cache):\n    Z = activation_cache\n    dZ = np.array(dA, copy=True) #converting dZ to the correct object\n    dZ[Z <= 0] = 0 #fixing negative values to 0\n    return dZ\n    \n\ndef d_sigmoid(dA, activation_cache):\n    Z = activation_cache\n    s, temp  = sigmoid(Z)\n    dZ = dA * s * (1-s)\n    return dZ\n\ndef linear_backward(dZ, cache):\n    \"\"\"\n    Linear portion of backward propagation for a single layer\n    Args:\n    dZ: Gradient of the cost with respect to the linear output\n    cache: tuple of values(A_prex, W, b) coming from the forward propagation in the current layer\n    Returns:\n    dA_prev: Gradient of the cost with respect to the activation of the previous layer (l-1)\n    dW: Gradient of the cost with respect to W in the current layer\n    db: Gradient of the cost with respect to b in the current layer\n    \"\"\"\n    A_prev, W, b = cache\n    m = A_prev.shape[1]\n    \n    dW = np.dot(1/m, np.dot(dZ, A_prev.T))\n    db = np.dot(1/m, np.sum(dZ, axis=1, keepdims=True))\n    dA_prev = np.dot(W.T, dZ)\n    \n    return dA_prev, dW, db\n\ndef linear_activation_backward(dA, cache, activation):\n    \"\"\"\n    Implements backward propagation for the Linear -> Activation layer\n    Args:\n    dA: post-activation gradient for current layer\n    cache: tuple of values(linear_cache, activation_cache)\n    activation: activation function to be used in this layer -- represented by a string\n    \n    Returns:\n    dA_prev: Gradient of the cost with respect to the activation of the previous layer\n    dW: Gradient of the cost with respect to W\n    db: Gradient of the cost with respect to b\n    \"\"\"\n    linear_cache, activation_cache = cache\n    \n    if activation == \"relu\":\n        dZ = d_relu(dA, activation_cache)\n        dA_prev, dW, db = linear_backward(dZ, linear_cache)\n    else:\n        dZ = d_sigmoid(dA, activation_cache)\n        dA_prev, dW, db = linear_backward(dZ, linear_cache)\n        \n    return dA_prev, dW, db","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:48.656552Z","iopub.execute_input":"2022-07-18T23:13:48.657158Z","iopub.status.idle":"2022-07-18T23:13:48.668582Z","shell.execute_reply.started":"2022-07-18T23:13:48.657124Z","shell.execute_reply":"2022-07-18T23:13:48.667705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"With all the helper functions written, let's finally write our backpropagation function. We know our output equals $A^{[L]} = \\sigma(Z^{[L]})$. Through some calculus, we derive the partial derivative of the cost with respect to $A^{[L]}$, which is equal to\n\n$$\\begin{equation}\n\\frac{\\partial{J}}{\\partial\\vec{A}^{[L]}}= \\frac{1}{m} \\Bigl(\\frac{1}{1 - \\vec{A}^{[L]}} \\odot (1 - \\vec{Y}) - \\frac{1}{\\vec{A}^{[L]}} \\odot \\vec{Y}\\Bigr).\n\\end{equation}$$\n\n\nWe will use a for loop to iterate through all the layers of our neural network, storing dA, dW, and db in a dictionary.","metadata":{}},{"cell_type":"code","source":"def nn_backprop(AL, Y, caches):\n    \"\"\"\n    Implements backward propagation for Linear->ReLU * (L-1) -> Linear -> Sigmoid computation\n    Args:\n    AL: output of forward propagation\n    Y: label vectors\n    caches: list of caches containing every cache of linear_activation_forward() with \"relu\" and \"sigmoid\"\n    Returns:\n    grads: A dictionary with our computed gradients\n    \"\"\"\n    grads = {}\n    L = len(caches)\n    m = AL.shape[1]\n    Y = Y.reshape(AL.shape)\n    \n    dAL = -(np.divide(Y,AL+1e-7) - np.divide(1-Y, 1-AL+1e-7))\n    \n    current_cache = caches[L-1]\n    dA_prev_temp, dW_temp, db_temp = linear_activation_backward(dAL, current_cache, \"sigmoid\")\n    grads[\"dA\"+str(L-1)] = dA_prev_temp\n    grads[\"dW\"+str(L)] = dW_temp\n    grads[\"db\"+str(L)] = db_temp\n    \n    for i in reversed(range(L-1)):\n        current_cache = caches[i]\n        dA_prev_temp, dW_temp, db_temp = linear_activation_backward(dA_prev_temp, current_cache, \"relu\")\n        grads[\"dA\"+str(i)] = dA_prev_temp\n        grads[\"dW\"+str(i+1)] = dW_temp\n        grads[\"db\"+str(i+1)] = db_temp\n        \n    return grads","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:48.717046Z","iopub.execute_input":"2022-07-18T23:13:48.717399Z","iopub.status.idle":"2022-07-18T23:13:48.729181Z","shell.execute_reply.started":"2022-07-18T23:13:48.717366Z","shell.execute_reply":"2022-07-18T23:13:48.728390Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Gradient Descent\nLet's write a function to update the parameters of our model using gradient descent. We will use these two formulas: \n\n$$ W^{[l]} = W^{[l]} - \\alpha \\text{ } dW^{[l]}$$\n$$ b^{[l]} = b^{[l]} - \\alpha \\text{ } db^{[l]}$$\n\nwhere $\\alpha$ is the learning rate of our model. \n","metadata":{}},{"cell_type":"code","source":"def update_parameters(params, grads, learning_rate):\n    \"\"\"\n    Implement gradient descent to update our parameters\n    Args:\n    params: dictionary containing parameters of our model\n    grads: dictionary containing our gradient vectors\n    Returns:\n    parameters: updated dictionary\n    \"\"\"\n    parameters = params.copy()\n    L = len(parameters)//2\n    \n    for i in range(L):\n        parameters[\"W\"+str(i+1)] = params[\"W\"+str(i+1)] - learning_rate*grads[\"dW\"+str(i+1)]\n        parameters[\"b\"+str(i+1)] = params[\"b\"+str(i+1)] - learning_rate*grads[\"db\"+str(i+1)]\n        \n    return parameters","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:48.778103Z","iopub.execute_input":"2022-07-18T23:13:48.778394Z","iopub.status.idle":"2022-07-18T23:13:48.784602Z","shell.execute_reply.started":"2022-07-18T23:13:48.778368Z","shell.execute_reply":"2022-07-18T23:13:48.783620Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We finally have all the functions necessary to implement a deep neural network.\n### Putting everything together\nLet's go ahead and make our label sets, and run a check on our matrix dimensions.","metadata":{}},{"cell_type":"code","source":"train_labels = np.array(train_df[\"Survived\"])\ntrain_labels = train_labels[np.newaxis]\ntrain_df = train_df.drop([\"Survived\"], axis=1)\ntrain_df.shape\n","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:48.841826Z","iopub.execute_input":"2022-07-18T23:13:48.842595Z","iopub.status.idle":"2022-07-18T23:13:48.851630Z","shell.execute_reply.started":"2022-07-18T23:13:48.842555Z","shell.execute_reply":"2022-07-18T23:13:48.850705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have 7 input features, and 891 training sets.\nLet's initialize the parameters for our deep neural network. Let's choose to use 3 hidden layers, with 20, 7, and 5 nodes respectively in each layer. Of course, being a binary classifier, our neural network's output layer will be one node. \n\n\nWe will choose our learning rate to be .00075.\n\nFinally, for iterations, we will choose 500000 iterations of training.","metadata":{}},{"cell_type":"markdown","source":"To begin, let's convert our input data from a pandas dataframe to a Numpy array, and initialize our layer dimensions list","metadata":{}},{"cell_type":"code","source":"train_df = train_df.to_numpy().T\nlayer_dims = [7,20,7,5,1]","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:48.938399Z","iopub.execute_input":"2022-07-18T23:13:48.939024Z","iopub.status.idle":"2022-07-18T23:13:48.943449Z","shell.execute_reply.started":"2022-07-18T23:13:48.938987Z","shell.execute_reply":"2022-07-18T23:13:48.942425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's create a function to facilitate our neural network.","metadata":{}},{"cell_type":"code","source":"def L_layer_nn(X, Y, layer_dims, learning_rate, num_iterations):\n    \"\"\"\n    Implements an L-Layer neural network with Linear->ReLU*(L-1) -> Linear->Sigmoid layers.\n    Args:\n    X: Input data -- np array\n    Y: Label vector\n    layer_dims: list containing our layer dimensions\n    num_iterations: number of iterations of the optimization loop\n    Returns:\n    parameters: the parameters learned by our model, which can be used to predict\n    \"\"\"\n    costs = []\n    parameters = {}\n    for i in range(1, len(layer_dims)):\n        parameters[\"W\"+str(i)] = np.random.randn(layer_dims[i], layer_dims[i-1])*np.sqrt(2/layer_dims[i-1])\n        parameters[\"b\"+str(i)] = np.zeros((layer_dims[i],1))\n    for i in range(0, num_iterations):\n        AL, caches = nn_forward(X, parameters)\n        cost = compute_cost(AL, Y)\n        grads = nn_backprop(AL, Y, caches)\n        parameters = update_parameters(parameters, grads, learning_rate)\n        if i % 100000 == 0 or i == num_iterations - 1:\n            print(\"Cost after iteration {}: {}\".format(i, np.squeeze(cost)))\n        if (i % 100 == 0 or i == num_iterations) and cost < 1.0:\n            costs.append(cost)\n    return parameters, costs    ","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:49.038401Z","iopub.execute_input":"2022-07-18T23:13:49.039033Z","iopub.status.idle":"2022-07-18T23:13:49.049921Z","shell.execute_reply.started":"2022-07-18T23:13:49.038999Z","shell.execute_reply":"2022-07-18T23:13:49.048918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that we have our neural network function written, let's train our model.","metadata":{}},{"cell_type":"code","source":"def plot_costs(costs, learning_rate):\n    plt.plot(np.squeeze(costs))\n    plt.ylabel('cost')\n    plt.xlabel('iterations (per hundreds)')\n    plt.title(\"Learning rate =\" + str(learning_rate))\n    plt.ylim(0,1)\n    plt.show()\n#winner - .00075, 86% accuracy\nlearning_rate = .00075\nparameters, costs = L_layer_nn(train_df, train_labels, layer_dims, learning_rate, num_iterations=500000)\nplot_costs(costs, learning_rate)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:13:49.136006Z","iopub.execute_input":"2022-07-18T23:13:49.136590Z","iopub.status.idle":"2022-07-18T23:19:44.610065Z","shell.execute_reply.started":"2022-07-18T23:13:49.136556Z","shell.execute_reply":"2022-07-18T23:19:44.609185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's write a quick prediction function to see how our model is performing.","metadata":{}},{"cell_type":"code","source":"def predict(X, Y, parameters):\n    \"\"\"\n    Predicts results of a L-layer neural network.\n    Args:\n    X: data set of examples being predicted\n    Y: true labels\n    parameters: parameters of the trained model\n    Returns:\n    predictions: predictions of the given dataset X\n    \"\"\"\n    m = X.shape[1]\n    n = len(parameters) // 2\n    p = np.zeros((1,m))\n    \n    probas, caches = nn_forward(X, parameters)\n    for i in range(0, probas.shape[1]):\n        if probas[0, i] > 0.5:\n            p[0, i] = 1\n        else:\n            p[0, i] = 0\n    print(\"Accuracy: \"+str(np.sum((p == Y) / m)))\n    \n    return p","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:19:44.611576Z","iopub.execute_input":"2022-07-18T23:19:44.612005Z","iopub.status.idle":"2022-07-18T23:19:44.618958Z","shell.execute_reply.started":"2022-07-18T23:19:44.611960Z","shell.execute_reply":"2022-07-18T23:19:44.617821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_train = predict(train_df, train_labels, parameters)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:19:44.620072Z","iopub.execute_input":"2022-07-18T23:19:44.620953Z","iopub.status.idle":"2022-07-18T23:19:44.635204Z","shell.execute_reply.started":"2022-07-18T23:19:44.620916Z","shell.execute_reply":"2022-07-18T23:19:44.634409Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Nice. Let's wrap things up by predicting the survivorship of our test set.","metadata":{}},{"cell_type":"code","source":"pass_ids = test_df[\"PassengerId\"]\ntest_df = np.array(test_df.drop([\"PassengerId\"], axis=1)).T\nprobas, caches = nn_forward(test_df, parameters)\np = np.zeros((1, test_df.shape[1]))\nfor i in range(0, probas.shape[1]):\n    if probas[0, i] > 0.5:\n        p[0, i] = 1\n    else:\n        p[0, i] = 0\np = np.resize(p, (p.shape[1],))","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:19:44.636908Z","iopub.execute_input":"2022-07-18T23:19:44.637673Z","iopub.status.idle":"2022-07-18T23:19:44.648552Z","shell.execute_reply.started":"2022-07-18T23:19:44.637623Z","shell.execute_reply":"2022-07-18T23:19:44.647639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And let's finally submit our findings.","metadata":{}},{"cell_type":"code","source":"submission = pd.DataFrame({\n        \"PassengerId\": pass_ids,\n        \"Survived\": p.astype('int')})\nsubmission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T23:43:25.657076Z","iopub.execute_input":"2022-07-18T23:43:25.657428Z","iopub.status.idle":"2022-07-18T23:43:25.666784Z","shell.execute_reply.started":"2022-07-18T23:43:25.657399Z","shell.execute_reply":"2022-07-18T23:43:25.665811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Results\nThis notebook was my personal endeavour for understanding more comprehensively the workings of a neural network. I played around a lot with the hyperparameters, to no avail. I fiddled around with my math, and had to debug many errors such as division by zero due to floating point accuracy, or my code generally not behaving as expected. But I had a lot of fun with this notebook. \n\n#### Inspirations and sources\nOf course, this neural network model is inspired by [Andrew Ng's Coursera Course on Deep Learning and Neural Networks](https://www.coursera.org/learn/neural-networks-deep-learning).\n\nAs for feature engineering, I deferred to [Mr. Manav Sehgal's notebook on this data set](https://www.kaggle.com/code/startupsci/titanic-data-science-solutions/notebook), which provided much needed guidance throughout this journey.\n\n**-- Thank you for reading my notebook --**","metadata":{}}]}