{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1><center>Titanic Survivors</center></h1>\n\nThe following notebook is a simple solution to solve the Titanic survivor problem. Simply put we use data provided from the ill-fated Titanic voyage and attempt to predict who would survive and who would die.\n","metadata":{}},{"cell_type":"markdown","source":"\n## Data Dictionary\n\n__Survival:__ (Did they survive?) 0 = No, 1 = Yes\n\n__Pclass:__ (Ticket class) = 1st, 2 = 2nd, 3 = 3rd\n\n__Sex:__ (Gender)\n\n__Age:__ Age in years\n\n__SibSp:__ Number of siblings / spouses aboard the Titanic\n\n__Parch:__ Number of parents / children aboard the Titanic\n\n__Ticket:__ The ticket number\n\n__Fare:__ Passenger price fare\n\n__Cabin:__ Cabin number\n\n__Embarked:__ Port of Embarkation C = Cherbourg, Q = Queenstown, S = Southampton\n\n\n## Solution Overview\n\nKeras / Tensorflow will be used to create a simple neural network that can be trained to understand the data and perform predictions. This notebook will also be a line-by-line explanation of the code so that the code is easier to follow.\n\n## Presumed Knowledge\nThere is some presumed knowledge of:\n<ul>\n    <li>Python 3+</li>\n    <li>Neural Networks</li>\n    <li>Keras / Tensorflow</li>\n</ul>\n","metadata":{"ExecuteTime":{"end_time":"2022-06-25T13:15:35.07852Z","start_time":"2022-06-25T13:15:35.068804Z"}}},{"cell_type":"markdown","source":"<h1><center>Introduction and Data Analysis</center></h1>\n\n","metadata":{}},{"cell_type":"code","source":"#Importing usual suspects\nimport pandas as pd # LIbrary to help load and explore data\nimport numpy as np # Library for mathematical functions and support for arrays and matrices\n\n#The following 2 lines are basically a config option to allow the cells in the notebook \n#to print all interactive input and not just the last one.\nfrom IPython.core.interactiveshell import InteractiveShell \nInteractiveShell.ast_node_interactivity=\"all\"","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:50.513457Z","start_time":"2022-06-30T13:41:50.511114Z"},"cell_style":"center","code_folding":[],"execution":{"iopub.status.busy":"2022-07-09T01:04:31.667020Z","iopub.execute_input":"2022-07-09T01:04:31.667681Z","iopub.status.idle":"2022-07-09T01:04:31.703619Z","shell.execute_reply.started":"2022-07-09T01:04:31.667584Z","shell.execute_reply":"2022-07-09T01:04:31.702363Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"training_data = pd.read_csv(\"../input/c/titanic/train.csv\") # Load the training data into a dataframe\ntraining_data.head(10) # show 10 line preview if brackets are falso just shows 5 as default","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:50.528415Z","start_time":"2022-06-30T13:41:50.514664Z"},"cell_style":"center","execution":{"iopub.status.busy":"2022-07-09T01:04:31.705504Z","iopub.execute_input":"2022-07-09T01:04:31.706109Z","iopub.status.idle":"2022-07-09T01:04:31.751839Z","shell.execute_reply.started":"2022-07-09T01:04:31.706074Z","shell.execute_reply":"2022-07-09T01:04:31.750954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"training_data.describe() # Show stats info about data - note only works for numerical value fields","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:50.548667Z","start_time":"2022-06-30T13:41:50.529937Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:31.753105Z","iopub.execute_input":"2022-07-09T01:04:31.753592Z","iopub.status.idle":"2022-07-09T01:04:31.801736Z","shell.execute_reply.started":"2022-07-09T01:04:31.753562Z","shell.execute_reply":"2022-07-09T01:04:31.800955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"training_data.info() #Information about the datatypes in the dataframe","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:50.555553Z","start_time":"2022-06-30T13:41:50.549869Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:31.803703Z","iopub.execute_input":"2022-07-09T01:04:31.804263Z","iopub.status.idle":"2022-07-09T01:04:31.823068Z","shell.execute_reply.started":"2022-07-09T01:04:31.804229Z","shell.execute_reply":"2022-07-09T01:04:31.821926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Cleaning\n\nProcessing data through a neural network involves a fair bit of work to get the input in the right was so that the neural network can _understand_ it. This next section looks at the data and some ways of making the data more neural network friendly.","metadata":{}},{"cell_type":"code","source":"training_data.isnull().sum() #Show how many null values there are","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:50.560478Z","start_time":"2022-06-30T13:41:50.556847Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:31.824520Z","iopub.execute_input":"2022-07-09T01:04:31.824966Z","iopub.status.idle":"2022-07-09T01:04:31.836133Z","shell.execute_reply.started":"2022-07-09T01:04:31.824932Z","shell.execute_reply":"2022-07-09T01:04:31.834823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now lets look at each column individually and make some decisions as to whether or not we keep and fix the data or remove thius for the purposes of training our neural network. \n\n_Please note that this are decisions for this particular solution / notebook and other solutions may keep more or remove more data points/columns - which can also be a valid solution_\n1. PassengerId: This column will be dropped from training data as it is not relevant to the output decision of survival probablity\n2. Survived: This will be the column used as the output to train the neural network for\n3. Pclass: The class of passenger would be a relevant data point so this will be kept.\n4. Name: Name is probably not a relevant factor so this can be dropped from the training data \n5. Sex: Gender will be relevant to surcvalval so lets keep\n6. Age: Age might be a relvant factor to determine survival possibility so this should be kept and fixed\n7. SibSp: The count of siblings/spouses will be kept as it would be relevant to the survival probablity\n8. Parch: The count of parents/children will be kept as it would be relevant to the survival probablity\n9. Ticket: The ticket number is probably not relevant and should be ignored\n10. Fare: The price of the ticket could be relevant to survival probablity and will be kept\n11. Cabin: As there is no further information regarding the cabin besides the number then this not relevant and will be ignored.\n\n**_Note that we need to do this to the test/unseen set too when we load that - we need to ensure that the this set resembles the training set_**\n\n","metadata":{}},{"cell_type":"code","source":"median_age_val = training_data[\"Age\"].median() # show the median age\nf\"Median Age: {median_age_val}\" \n","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:50.587463Z","start_time":"2022-06-30T13:41:50.58459Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:31.837680Z","iopub.execute_input":"2022-07-09T01:04:31.838453Z","iopub.status.idle":"2022-07-09T01:04:31.849642Z","shell.execute_reply.started":"2022-07-09T01:04:31.838405Z","shell.execute_reply":"2022-07-09T01:04:31.848771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mean_age_val = training_data[\"Age\"].mean() # show the average age\nf\"Mean Age: {mean_age_val}\" ","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:50.591193Z","start_time":"2022-06-30T13:41:50.588635Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:31.851179Z","iopub.execute_input":"2022-07-09T01:04:31.852118Z","iopub.status.idle":"2022-07-09T01:04:31.863537Z","shell.execute_reply.started":"2022-07-09T01:04:31.852070Z","shell.execute_reply":"2022-07-09T01:04:31.862173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Probably ok to pick either - lets go with mean.\n\nBefore this is done let's build a helper function that preprocesses (ie. cleans and engineers the features as needed). This is so we don't repeat the code over over again for the test set and any potential new tests we may want to do...\n\n","metadata":{}},{"cell_type":"code","source":"#Function to clean the age column\ndef clean_age_column(age_col) -> pd.DataFrame:\n    age_cleaned_df = age_col.fillna(training_data[\"Age\"].mean()) # Habit to create intermediate variables - you could use inplace=True as part of the inplace method params but I like to have original and intermediate variables just in case\n    # this is now the age column we will use later\n    return age_cleaned_df\n        ","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:50.614324Z","start_time":"2022-06-30T13:41:50.612222Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:31.865294Z","iopub.execute_input":"2022-07-09T01:04:31.865970Z","iopub.status.idle":"2022-07-09T01:04:31.875181Z","shell.execute_reply.started":"2022-07-09T01:04:31.865931Z","shell.execute_reply":"2022-07-09T01:04:31.874260Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Function to clean the embarked column\n\ndef clean_embarked_column(embarked_col) -> pd.DataFrame:\n    #a simple way is to just take the most embarked port and fill the nulls. \n    embarked_cleaned =embarked_col.fillna(training_data[\"Embarked\"].mode()[0])\n    return embarked_cleaned","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:50.617808Z","start_time":"2022-06-30T13:41:50.615705Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:31.878198Z","iopub.execute_input":"2022-07-09T01:04:31.878851Z","iopub.status.idle":"2022-07-09T01:04:31.889211Z","shell.execute_reply.started":"2022-07-09T01:04:31.878813Z","shell.execute_reply":"2022-07-09T01:04:31.887990Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#As stated previously we will only work with the columns we need \ndef filter_relevant_data(df) -> pd.DataFrame:\n    clean_training_dataset = df.filter(['Survived', 'Pclass', 'Sex', 'SibSp', 'Parch', 'Fare'])\n    return clean_training_dataset","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:50.621211Z","start_time":"2022-06-30T13:41:50.6192Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:31.890763Z","iopub.execute_input":"2022-07-09T01:04:31.891221Z","iopub.status.idle":"2022-07-09T01:04:31.902401Z","shell.execute_reply.started":"2022-07-09T01:04:31.891185Z","shell.execute_reply":"2022-07-09T01:04:31.900655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def perform_preprocessing(df) -> pd.DataFrame:\n    cleaned_age_column = clean_age_column(df[\"Age\"])\n    cleaned_embarked_column = clean_embarked_column(df[\"Embarked\"])\n    cleaned_df = filter_relevant_data(df)\n    cleaned_df[\"Age\"] = cleaned_age_column\n    cleaned_df[\"Embarked\"] = cleaned_embarked_column\n    return cleaned_df\n        ","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:50.624745Z","start_time":"2022-06-30T13:41:50.622499Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:31.904122Z","iopub.execute_input":"2022-07-09T01:04:31.904472Z","iopub.status.idle":"2022-07-09T01:04:31.920373Z","shell.execute_reply.started":"2022-07-09T01:04:31.904431Z","shell.execute_reply":"2022-07-09T01:04:31.919021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"clean_training_dataset = perform_preprocessing(training_data)\nprint(\"--- STEP 1 COMPLETE: Remove nulls and get relevant columns for training data ---\")\nprint(clean_training_dataset.isna().sum())\n    ","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:50.631817Z","start_time":"2022-06-30T13:41:50.625952Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:31.921672Z","iopub.execute_input":"2022-07-09T01:04:31.922615Z","iopub.status.idle":"2022-07-09T01:04:31.938685Z","shell.execute_reply.started":"2022-07-09T01:04:31.922578Z","shell.execute_reply":"2022-07-09T01:04:31.937834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see from the above we now have no more nulls so for the moment we can consider this data as _clean_ or probably more appropriately _complete_.\n\nThe next section we do some visulisations to get some intuition from our data.","metadata":{}},{"cell_type":"markdown","source":"## Visualising and Understanding\n\nThis next section we will use a matplolib and seaborn (graphing and plotting libraries) to visualise some data and get more intuition and understanding of the data.\n\nLet's start with some simple single variable graphs","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\nfig, ax1 = plt.subplots(2, 2, figsize=(15,15)) #Create a figure of nrows and ncols and add the following plots to that figure\n\n\nsurvived_df = clean_training_dataset[\"Survived\"] # First graph takes the survived column\n# We plot a pie chart of that column based on the counts and add it to the 0,0 position of the figure created before\nsurvived_df.value_counts().plot(kind=\"pie\", ax=ax1[0,0]) \n\n\n#The rest of the graphs follow suit.\nclass_df = clean_training_dataset[\"Pclass\"]\nclass_df.value_counts().plot(kind=\"pie\", ax=ax1[0,1])\n\ngender_df = clean_training_dataset[\"Sex\"]\ngender_df.value_counts().plot(kind=\"pie\", ax=ax1[1,0])\n\n\nsibsp_df = clean_training_dataset[\"SibSp\"]\nsibsp_df.value_counts().plot(kind=\"pie\", ax=ax1[1,1])\n\n","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:50.871879Z","start_time":"2022-06-30T13:41:50.668665Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:31.939975Z","iopub.execute_input":"2022-07-09T01:04:31.940936Z","iopub.status.idle":"2022-07-09T01:04:32.462083Z","shell.execute_reply.started":"2022-07-09T01:04:31.940901Z","shell.execute_reply":"2022-07-09T01:04:32.460870Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Graph 1 - Shows that not a lot survived (remember 0 = Did not survive)\n\nGraph 2 - Shows that there were alot of people on 3rd class tickets\n\nGraph 3 - Shows that there were a lot more men than women\n\nGraph 4 - Shows that there were a lot of people without Siblings or Spouses (single possibly)\n\nNow this is all well and good but lets do some deeper exploration with some examples of multiple variable graphs","metadata":{}},{"cell_type":"code","source":"import seaborn as sns # Seaborn is built on matplotlib and gives some richer and more convenient features\n\nsns.countplot(x=\"Survived\", data=clean_training_dataset) #as an example we can see a bar chart for survived vs not\n","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:50.931309Z","start_time":"2022-06-30T13:41:50.873063Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:32.463888Z","iopub.execute_input":"2022-07-09T01:04:32.464576Z","iopub.status.idle":"2022-07-09T01:04:33.880901Z","shell.execute_reply.started":"2022-07-09T01:04:32.464532Z","shell.execute_reply":"2022-07-09T01:04:33.880114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.catplot(x=\"Sex\", col=\"Survived\", kind=\"count\", data=clean_training_dataset)","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:51.083475Z","start_time":"2022-06-30T13:41:50.932688Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:33.881917Z","iopub.execute_input":"2022-07-09T01:04:33.882588Z","iopub.status.idle":"2022-07-09T01:04:34.243392Z","shell.execute_reply.started":"2022-07-09T01:04:33.882558Z","shell.execute_reply":"2022-07-09T01:04:34.242484Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Looks like more men did not survive than compared to women. Could be because, as the story goes, women and children were saved first.**","metadata":{}},{"cell_type":"code","source":"sns.catplot(x=\"Pclass\", col=\"Survived\", kind=\"count\", data=clean_training_dataset)","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:51.252342Z","start_time":"2022-06-30T13:41:51.084813Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:34.247256Z","iopub.execute_input":"2022-07-09T01:04:34.247944Z","iopub.status.idle":"2022-07-09T01:04:34.635475Z","shell.execute_reply.started":"2022-07-09T01:04:34.247889Z","shell.execute_reply":"2022-07-09T01:04:34.634370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Seems reasonable to see those with worse living arrangements had a higher mortality rate. Intuitively shows that the upper class were probably better off**","metadata":{}},{"cell_type":"code","source":"sns.catplot(x=\"Embarked\", col=\"Survived\", kind=\"count\", data=clean_training_dataset)","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:51.430473Z","start_time":"2022-06-30T13:41:51.253623Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:34.637024Z","iopub.execute_input":"2022-07-09T01:04:34.637475Z","iopub.status.idle":"2022-07-09T01:04:35.016103Z","shell.execute_reply.started":"2022-07-09T01:04:34.637440Z","shell.execute_reply":"2022-07-09T01:04:35.015178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**The above graph shows the mortality for the people who got on at different ports**","metadata":{}},{"cell_type":"code","source":"sns.catplot(x=\"Parch\", col=\"Survived\", kind=\"count\", data=clean_training_dataset)","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:51.617876Z","start_time":"2022-06-30T13:41:51.431521Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:35.017529Z","iopub.execute_input":"2022-07-09T01:04:35.018301Z","iopub.status.idle":"2022-07-09T01:04:35.471769Z","shell.execute_reply.started":"2022-07-09T01:04:35.018253Z","shell.execute_reply":"2022-07-09T01:04:35.470942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**The above graph is interesting and shows that proportinately, those with children or have parents seem to survive more than the ones with none. This supports the common \"women and children first\" adage.**\n\nFor brevity's sake not all the exploration and graphs have been included here and the above are just examples.For the next step we are going to be engineering the data to fit more in line with what a KERAS NN expects. ","metadata":{}},{"cell_type":"markdown","source":"<h1><center>Feature Engineering</center></h1>\n\nThis section creates categories where possible and engineers the features required for NN training to more NN friendly version.","metadata":{}},{"cell_type":"markdown","source":"## Family Size\n\nAs we saw earlier those with parents/children seem to have a better rate of surival. Lets extrapolate this with the size of the family. Lets add people with their parents and children aadn well as siblinsg and spuses and put them into the following categories:\n\n<ul>\n    <li>Singles</li>\n    <li>Small families</li>\n    <li>Standard families</li>\n    <li>Large families</li>\n</ul>\n\n","metadata":{}},{"cell_type":"code","source":"def create_family_size_bucket(family_size)->str:\n\n    if family_size < 2:\n        return \"Single\"\n    if 2<=family_size <=3:\n        return \"Small Family\"\n    if 4 <= family_size <=6:\n        return \"Standard Family\"\n    if family_size > 6:\n        return \"Big Family\"\n\n\ndef create_family_size_category(df)->pd.DataFrame:\n    family_size_df = df[\"SibSp\"] + df[\"Parch\"] + 1 #Parents, children, siblings, spouses + 1 for themselves\n    family_size_bucket = family_size_df.apply(lambda row: create_family_size_bucket(row)) #apply is more effecient way to perform an operation on each row of a dataset\n    return family_size_bucket\n\nclean_training_dataset[\"Family_Size\"] = create_family_size_category(clean_training_dataset)\nclean_training_dataset.head()","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:51.628523Z","start_time":"2022-06-30T13:41:51.618887Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:35.473207Z","iopub.execute_input":"2022-07-09T01:04:35.474301Z","iopub.status.idle":"2022-07-09T01:04:35.497765Z","shell.execute_reply.started":"2022-07-09T01:04:35.474252Z","shell.execute_reply":"2022-07-09T01:04:35.496199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.catplot(x=\"Family_Size\", col=\"Survived\", kind=\"count\", data=clean_training_dataset)","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:51.808101Z","start_time":"2022-06-30T13:41:51.630754Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:35.499397Z","iopub.execute_input":"2022-07-09T01:04:35.500103Z","iopub.status.idle":"2022-07-09T01:04:35.958977Z","shell.execute_reply.started":"2022-07-09T01:04:35.500068Z","shell.execute_reply":"2022-07-09T01:04:35.957929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Seems like the singles were worst off**","metadata":{}},{"cell_type":"markdown","source":"## Age Bracket\n\nLets put the this data point (which is a continuous variable and a little harder to handle) into a category of age brackets.\n\n<ul>\n    <li>Child</li>\n    <li>Teen</li>\n    <li>Young Adult</li>\n    <li>Mature</li>\n    <li>Senior</li>\n    \n</ul>\n\n","metadata":{}},{"cell_type":"code","source":"clean_training_dataset[\"Age\"].max()\nclean_training_dataset[\"Age\"].min()\nclean_training_dataset[\"Age\"].mean()","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:51.814456Z","start_time":"2022-06-30T13:41:51.809618Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:35.960765Z","iopub.execute_input":"2022-07-09T01:04:35.961182Z","iopub.status.idle":"2022-07-09T01:04:35.975816Z","shell.execute_reply.started":"2022-07-09T01:04:35.961148Z","shell.execute_reply":"2022-07-09T01:04:35.974983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def put_in_age_bucket(age):\n    \n    if 0.0 <= age <= 12.99:\n        return \"Child\"\n    if 13.0 <= age <= 19.99:\n        return \"Teen\"\n    if 20.0 <= age <= 39.99:\n        return \"Young Adult\"\n    if 40.0 <= age <= 59.99:\n        return \"Mature Age\"\n    if age >= 60:\n        return \"Senior\"\n\n\ndef create_age_buckets(age_df):\n    age_bucket_df = age_df.apply(lambda row: put_in_age_bucket(row))\n    return age_bucket_df\n\n\nclean_training_dataset[\"Age_Bucket\"] = create_age_buckets(clean_training_dataset[\"Age\"]) \nclean_training_dataset.head()","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:51.825367Z","start_time":"2022-06-30T13:41:51.815542Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:35.977380Z","iopub.execute_input":"2022-07-09T01:04:35.978034Z","iopub.status.idle":"2022-07-09T01:04:36.007167Z","shell.execute_reply.started":"2022-07-09T01:04:35.977991Z","shell.execute_reply":"2022-07-09T01:04:36.006036Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.catplot(x=\"Age_Bucket\", col=\"Survived\", kind=\"count\", data=clean_training_dataset)","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:52.024333Z","start_time":"2022-06-30T13:41:51.826557Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:36.008736Z","iopub.execute_input":"2022-07-09T01:04:36.010261Z","iopub.status.idle":"2022-07-09T01:04:36.425946Z","shell.execute_reply.started":"2022-07-09T01:04:36.010212Z","shell.execute_reply":"2022-07-09T01:04:36.424844Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Again the data visualisation shows that children, young people and the elderly had a beter probability of survival as they were probably saved first which seems intuitively correct**","metadata":{}},{"cell_type":"markdown","source":"## Fare Buckets\n\nNow lets look at another continuous variable fare and see how much the price of admission is related to survival probablity. Again as a continuous variable lets make it easier and break into bucket. \n\nBut let's first look at the fare data and see what we can learn","metadata":{}},{"cell_type":"code","source":"clean_training_dataset[\"Fare\"].max()\nclean_training_dataset[\"Fare\"].min()\nclean_training_dataset[\"Fare\"].mean()\nclean_training_dataset[\"Fare\"].median()\n\nclean_training_dataset.head()\nplt.figure(figsize=(20,5))\nsns.histplot(clean_training_dataset[clean_training_dataset.Survived == 0][\"Fare\"], \n             bins=50, color='r')\n","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:52.223927Z","start_time":"2022-06-30T13:41:52.025576Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:36.427769Z","iopub.execute_input":"2022-07-09T01:04:36.428249Z","iopub.status.idle":"2022-07-09T01:04:36.769136Z","shell.execute_reply.started":"2022-07-09T01:04:36.428205Z","shell.execute_reply":"2022-07-09T01:04:36.767946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(20,5))\nsns.histplot(clean_training_dataset[clean_training_dataset.Survived == 1][\"Fare\"], \n             bins=60, color='g')","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:52.386446Z","start_time":"2022-06-30T13:41:52.225328Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:36.770577Z","iopub.execute_input":"2022-07-09T01:04:36.771018Z","iopub.status.idle":"2022-07-09T01:04:37.249998Z","shell.execute_reply.started":"2022-07-09T01:04:36.770963Z","shell.execute_reply":"2022-07-09T01:04:37.248719Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**From the above data we can definitely see that those who paid more did better. A damning evidence of how the rich were probably looked after better**\n\nUsing the above data the thresholds in the code below seem reasonable to deliniate the following fare buckets:\n\n<ul>\n    <li>Low Fare</li>\n    <li>Cheap Fare</li>\n    <li>Standard Fare</li>\n    <li>High Fare</li>\n    <li>Highest Fare</li>\n</ul>\n\n","metadata":{}},{"cell_type":"code","source":"def put_fare_into_bucket(fare):\n    if 0.0 <= fare <= 10.0:\n        return \"Low Fare\"\n    if 10.0 < fare <= 20.0:\n        return \"Cheap Fare\"\n    if 20.0 < fare <= 50.0:\n        return \"Standard Fare\"\n    if 50.0 < fare <= 100.0:\n        return \"High Fare\"\n    if fare > 100:\n        return \"Highest Fare\"\n\ndef create_fare_buckets(fare_df):\n    fare_bucket_df = fare_df.apply(lambda row: put_fare_into_bucket(row))\n    return fare_bucket_df\n\n\nclean_training_dataset[\"Fare_Bucket\"] = create_fare_buckets(clean_training_dataset[\"Fare\"]) \n","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:52.391436Z","start_time":"2022-06-30T13:41:52.387457Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:37.251507Z","iopub.execute_input":"2022-07-09T01:04:37.252760Z","iopub.status.idle":"2022-07-09T01:04:37.264105Z","shell.execute_reply.started":"2022-07-09T01:04:37.252711Z","shell.execute_reply":"2022-07-09T01:04:37.262507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.catplot(x=\"Fare_Bucket\", col=\"Survived\", kind=\"count\", data=clean_training_dataset)","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:52.574659Z","start_time":"2022-06-30T13:41:52.392668Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:37.265612Z","iopub.execute_input":"2022-07-09T01:04:37.266760Z","iopub.status.idle":"2022-07-09T01:04:37.688075Z","shell.execute_reply.started":"2022-07-09T01:04:37.266696Z","shell.execute_reply":"2022-07-09T01:04:37.685556Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**So again we see that the people with the lowest fare are the unluckiest ones proportionately**","metadata":{}},{"cell_type":"code","source":"del clean_training_dataset[\"Fare\"]\ndel clean_training_dataset[\"Age\"]\ndel clean_training_dataset[\"SibSp\"]\ndel clean_training_dataset[\"Parch\"]\n\nclean_training_dataset.head()\n","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:52.583445Z","start_time":"2022-06-30T13:41:52.575821Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:37.689875Z","iopub.execute_input":"2022-07-09T01:04:37.690517Z","iopub.status.idle":"2022-07-09T01:04:37.711620Z","shell.execute_reply.started":"2022-07-09T01:04:37.690469Z","shell.execute_reply":"2022-07-09T01:04:37.710152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Convert and Scale","metadata":{}},{"cell_type":"markdown","source":"So now we have the data we need and it looks prety good. Next step is to change the data into numerical form for the NN to understand it. \n\nTo do this we have to represent everything as some numerical form and luckily we hanve Pandas to help with this.\n\n__Pandas.getDummies()__ function takes all the categorical features as specified specify and breaks them into category values. \n\nEg. the Sex colmn will be converted to 2 columns where a male is represented by Sex_male= 1 and Sex_female=0 and the Pclass column will be broken into 3 columns\n\nWe can see this when we run the next cell.","metadata":{}},{"cell_type":"code","source":"train_X = pd.get_dummies(clean_training_dataset, columns=[\"Pclass\", \"Sex\", \"Embarked\", \"Family_Size\", \"Age_Bucket\", \"Fare_Bucket\"],\n                                    prefix_sep=[\"_\",\"_\",\"_\",\"_\",\"_\",\"_\"], drop_first=True)\n\n\n\ntrain_X.head()\n","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:52.59638Z","start_time":"2022-06-30T13:41:52.584566Z"},"cell_style":"center","execution":{"iopub.status.busy":"2022-07-09T01:04:37.713910Z","iopub.execute_input":"2022-07-09T01:04:37.714755Z","iopub.status.idle":"2022-07-09T01:04:37.750994Z","shell.execute_reply.started":"2022-07-09T01:04:37.714707Z","shell.execute_reply":"2022-07-09T01:04:37.749736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Almost makes sense. \n\nHowever there seems to be one column missing from each data point mentioned and the reason for this is the _drop_first=True_ attribute that we added to the getDummies function.\n\nThis attribute when set to true drops the first of each clumn that has been categorised and split and this is done so that redundant columns / data is removed.\n\nFor example, taking the case of the Sex column, we don't have to specify Sex_male=1 **and** Sex_female=0 for a male. We can simply say that Sex_male=1 implies male and Sex_male=0 implies female. This follows true for all the other category fields as well.\n\nThis reduces the data data that is needed to be fed into the network making it more effecient. ","metadata":{}},{"cell_type":"markdown","source":"Now create the the y value or the output value we need to train the data for:","metadata":{}},{"cell_type":"code","source":"\ntrain_y = clean_training_dataset[\"Survived\"]#Survived is what we need\n\ndel train_X[\"Survived\"]#This can now be removed from the training data.\ntrain_y.head()\n","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:52.600543Z","start_time":"2022-06-30T13:41:52.597373Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:37.752908Z","iopub.execute_input":"2022-07-09T01:04:37.753684Z","iopub.status.idle":"2022-07-09T01:04:37.776825Z","shell.execute_reply.started":"2022-07-09T01:04:37.753548Z","shell.execute_reply":"2022-07-09T01:04:37.774226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's normalise all the training data for the neural network. \n\nNormalizing the data generally speeds up learning and leads to faster convergence.","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import StandardScaler # \n\n#Create a numpy array from the training dataset\nX = train_X.values\ny = train_y.values\nsc = StandardScaler()\n\nX = sc.fit_transform(X) #Standardize fit and transform the data before input into the machine learning model.\nprint(X)","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:52.733419Z","start_time":"2022-06-30T13:41:52.601531Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:37.778000Z","iopub.execute_input":"2022-07-09T01:04:37.778994Z","iopub.status.idle":"2022-07-09T01:04:37.870048Z","shell.execute_reply.started":"2022-07-09T01:04:37.778947Z","shell.execute_reply":"2022-07-09T01:04:37.868655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Unseen Data\n\nWe now need to load the unseen data or more specifically the data we are going to check our model with.\n\n","metadata":{}},{"cell_type":"markdown","source":"Now we load he data we want to predict (unseen data). After this we i should be checked / ensured that all preprocessing and data cleaning stages are performed.\n\nThis done exactly the same as the training data.","metadata":{}},{"cell_type":"code","source":"#Load the unseen data\nunseen_data_df = pd.read_csv(\"../input/c/titanic/test.csv\")\nunseen_data_df.head()","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:52.753875Z","start_time":"2022-06-30T13:41:52.735571Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:37.871633Z","iopub.execute_input":"2022-07-09T01:04:37.872110Z","iopub.status.idle":"2022-07-09T01:04:37.910335Z","shell.execute_reply.started":"2022-07-09T01:04:37.872074Z","shell.execute_reply":"2022-07-09T01:04:37.908770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"unseen_data_df.isna().sum()","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:52.759151Z","start_time":"2022-06-30T13:41:52.755143Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:37.912059Z","iopub.execute_input":"2022-07-09T01:04:37.913169Z","iopub.status.idle":"2022-07-09T01:04:37.925288Z","shell.execute_reply.started":"2022-07-09T01:04:37.913122Z","shell.execute_reply":"2022-07-09T01:04:37.924048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So now doing similar to the unseen (or test data) to what was done for the training data.","metadata":{}},{"cell_type":"code","source":"clean_unseen_df = perform_preprocessing(unseen_data_df)\n\n\nclean_unseen_df[\"Age_Bucket\"] = create_age_buckets(clean_unseen_df[\"Age\"]) \nclean_unseen_df[\"Fare_Bucket\"] = create_fare_buckets(clean_unseen_df[\"Fare\"])\nclean_unseen_df[\"Family_Size\"] = create_family_size_category(clean_unseen_df)\n\nclean_unseen_df.head()","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:56.511326Z","start_time":"2022-06-30T13:41:56.480879Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:37.927091Z","iopub.execute_input":"2022-07-09T01:04:37.928358Z","iopub.status.idle":"2022-07-09T01:04:37.958435Z","shell.execute_reply.started":"2022-07-09T01:04:37.928311Z","shell.execute_reply":"2022-07-09T01:04:37.957114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Convert and Scale\n\nSo now we have the data we need and it looks prety good. Next step is to change the data into numerical form for the NN to understand it. \n\nTo do this we have to represent everything as some numerical form and luckily we hanve Pandas to help with this.\n\n__Pandas.getDummies()__ function takes all the categorical features as specified specify and breaks them into category values. \n\nEg. the Sex colmn will be converted to 2 columns where a male is represented by Sex_male= 1 and Sex_female=0 and the Pclass column will be broken into 3 columns\n\nWe can see this when we run the next cell.","metadata":{}},{"cell_type":"code","source":"del clean_unseen_df[\"Fare\"]\ndel clean_unseen_df[\"Age\"]\ndel clean_unseen_df[\"SibSp\"]\ndel clean_unseen_df[\"Parch\"]\n\ntest_X = pd.get_dummies(clean_unseen_df, columns=[\"Pclass\", \"Sex\", \"Family_Size\", \"Embarked\",\"Age_Bucket\", \"Fare_Bucket\"],\n                                    prefix_sep=[\"_\",\"_\",\"_\",\"_\",\"_\",\"_\"], drop_first=True)\n\n\n\ntest_X.head()\ntest_X.info()","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:57.479441Z","start_time":"2022-06-30T13:41:57.456327Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:37.960132Z","iopub.execute_input":"2022-07-09T01:04:37.960833Z","iopub.status.idle":"2022-07-09T01:04:38.012870Z","shell.execute_reply.started":"2022-07-09T01:04:37.960769Z","shell.execute_reply":"2022-07-09T01:04:38.009881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"testX = test_X.values # create numpy array of  values\ntestX = testX.astype(np.float64, copy=False)","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:58.42782Z","start_time":"2022-06-30T13:41:58.42271Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:38.019808Z","iopub.execute_input":"2022-07-09T01:04:38.020614Z","iopub.status.idle":"2022-07-09T01:04:38.027252Z","shell.execute_reply.started":"2022-07-09T01:04:38.020563Z","shell.execute_reply":"2022-07-09T01:04:38.026384Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"testX = sc.fit_transform(testX)\n\nX.shape\ntestX.shape","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:41:59.248308Z","start_time":"2022-06-30T13:41:59.239542Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:38.028626Z","iopub.execute_input":"2022-07-09T01:04:38.029305Z","iopub.status.idle":"2022-07-09T01:04:38.048182Z","shell.execute_reply.started":"2022-07-09T01:04:38.029272Z","shell.execute_reply":"2022-07-09T01:04:38.046695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So now everything looks good and the unseen data is in the same shape and format as the training data. The next step is to create the model.","metadata":{}},{"cell_type":"markdown","source":"<h1><center>Neural Network Development</center></h1>","metadata":{}},{"cell_type":"markdown","source":"## Overview\n\nWe now need to develop an _artificial neural network_ which can learn with the training data and predict with a the probablity of survival with a reasonable degree of accuracy.\n\nWithout getting into the nitty gritty of neral networks we need to create:\n\n<ul>\n    <li>The Input layer - This will be the layer that takes the input </li>\n    <li>Output Layer - The final layer that gives us the prbablity of survival </li>\n    <li>Hidden Layers - An arbitrary number of layers which have nodes and wieights that can be altered/trained</li>\n</ul>\n\nThese will be combined together to create a neural network\n\nWe also need to make choices for:\n\n<ul>\n    <li>Activation Function</li>\n    <li>Loss Function</li>\n    <li>Optimizer</li>\n\n</ul>\n\n","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nimport keras\n# Debug checking\nprint(\"Tensorflow version: \" + tf.__version__)\nprint(\"Keras version: \"  + keras.__version__)\n\n\nfrom keras.models import Sequential\nfrom keras.layers import Dense, Activation, Dropout, Input\nfrom tensorflow.keras.optimizers import SGD\nfrom keras.callbacks import EarlyStopping","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:42:38.37274Z","start_time":"2022-06-30T13:42:37.011078Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:38.050482Z","iopub.execute_input":"2022-07-09T01:04:38.051338Z","iopub.status.idle":"2022-07-09T01:04:48.554202Z","shell.execute_reply.started":"2022-07-09T01:04:38.051292Z","shell.execute_reply":"2022-07-09T01:04:48.552972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Whilst a full theory on Keras and Tensorflow is out of scope here, the following section/code is commented to give as much detail as possible.\n\nThe architecture chosen was that of:\n\n<ul>\n    <li>2 Hidden Layers</li>\n    <li>RELU activation function</li>\n    <li>2 Dropout layers</li>\n    <li>2 Hidden Layers</li>\n    <li>Output layer with Sigmoid activation</li>\n    <li>Stochastic Gradient Descent for Optimizer</li>\n    <li>Loss Function - Binary Crossentropy</li>\n</ul>\n\nThis is just one example of an architecure that worked. There are many ways to develop different NN architectures (and/or different hyperparameters that can be tuned).\n","metadata":{}},{"cell_type":"markdown","source":"## Model Creation\n\nNow we implement the architecture.","metadata":{}},{"cell_type":"code","source":"predictive_model = Sequential() # a sequential neural network\n\ninput_layer = Input(shape=(X.shape[1],)) #the input layer must match the shape of the input\n\nhidden_layer_1 = Dense(32, activation=\"relu\", kernel_initializer=\"uniform\") # hidden layer with 32 nodes, relu activation and uniform initialized weights\ndropout_1 = Dropout(0.5) #droput layer - drop half the nodes - prevents overfitting\n\nhidden_layer_2 = Dense(64, activation=\"relu\", kernel_initializer=\"uniform\") \ndropout_2 = Dropout(0.5)\n\noutput_layer = Dense(1, kernel_initializer=\"uniform\", activation=\"sigmoid\") #the output layer is a one node layer that gives a probablity\n\n#put them all together\npredictive_model.add(input_layer)\npredictive_model.add(hidden_layer_1)\npredictive_model.add(dropout_1)\n\npredictive_model.add(hidden_layer_2)\npredictive_model.add(dropout_2)\n\npredictive_model.add(output_layer)\n\npredictive_model.summary() #print what it looks like","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:46:43.923202Z","start_time":"2022-06-30T13:42:39.221806Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:48.555713Z","iopub.execute_input":"2022-07-09T01:04:48.556372Z","iopub.status.idle":"2022-07-09T01:04:48.909315Z","shell.execute_reply.started":"2022-07-09T01:04:48.556340Z","shell.execute_reply":"2022-07-09T01:04:48.908022Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Training\n\nThe following section shows the training of the simple neural network developed using:\n\n**100 Epochs**: An epoch is basicaly when all data has passed therough the network (forward and backward propagation) once\n\n**Batch Size of 50**: The training data is broken up into batches of 50 \n\n**Binary Cross Entropy** loss function: As this is a binary classification problem (0 death or 1 for survival) this is a valid choice. \n\nWe also use the **stochastic gradient descent** as the optimizer (with learning rate 0.1 and momentum of 0.6) which is the chosen algorithm to reduce the loss.","metadata":{}},{"cell_type":"code","source":"callback = EarlyStopping(monitor=\"accuracy\", patience=10) # Stop training early if accuracy does not really change after 10 epochs\n\nsgd = SGD(learning_rate=0.1, momentum=0.6) \npredictive_model.compile(optimizer=sgd, loss=\"binary_crossentropy\", metrics=[\"accuracy\"])\n\n#Start training - this may take a minute based on uyour hardware\nhistory = predictive_model.fit(X, y, validation_split=0.2, batch_size=50, epochs=100, \n                               verbose=2, callbacks=[callback])","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:48:19.586565Z","start_time":"2022-06-30T13:47:00.599253Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:48.913357Z","iopub.execute_input":"2022-07-09T01:04:48.913707Z","iopub.status.idle":"2022-07-09T01:04:54.266233Z","shell.execute_reply.started":"2022-07-09T01:04:48.913676Z","shell.execute_reply":"2022-07-09T01:04:54.264998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if len(history.history[\"accuracy\"]) < 100:\n    print(\"Stopped early after \" + str(len(history.history['accuracy'])) + \" epochs\")\n\nprint(history.history.keys())","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:48:29.042979Z","start_time":"2022-06-30T13:48:29.032652Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:54.267715Z","iopub.execute_input":"2022-07-09T01:04:54.268090Z","iopub.status.idle":"2022-07-09T01:04:54.274204Z","shell.execute_reply.started":"2022-07-09T01:04:54.268060Z","shell.execute_reply":"2022-07-09T01:04:54.273160Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(history.history['accuracy'])\nplt.plot(history.history['val_accuracy'])\nplt.title('Model Accuracy')\nplt.ylabel('Accuracy')\nplt.xlabel('Epoch')\nplt.legend(['train', 'test'], loc='upper left')\n","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:49:06.649884Z","start_time":"2022-06-30T13:49:06.530129Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:54.275951Z","iopub.execute_input":"2022-07-09T01:04:54.276642Z","iopub.status.idle":"2022-07-09T01:04:54.510877Z","shell.execute_reply.started":"2022-07-09T01:04:54.276610Z","shell.execute_reply":"2022-07-09T01:04:54.509836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create Sumbmission (Optional)","metadata":{}},{"cell_type":"markdown","source":"This section just outlines how the submission for the kaggle competition was created.","metadata":{}},{"cell_type":"code","source":"#optional\npredictions = predictive_model.predict(testX)\nsubmissions = np.rint(predictions)\n\nsubmission_list = []\nfor s in submissions:\n    submission_list.append(int(s[0]))\n\nout = pd.DataFrame({\"PassengerId\": unseen_data_df.PassengerId, \"Survived\":submission_list})\n\nout.head(20)\n","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:49:12.29166Z","start_time":"2022-06-30T13:49:12.182382Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:54.512251Z","iopub.execute_input":"2022-07-09T01:04:54.512688Z","iopub.status.idle":"2022-07-09T01:04:54.682842Z","shell.execute_reply.started":"2022-07-09T01:04:54.512654Z","shell.execute_reply":"2022-07-09T01:04:54.681841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"out.to_csv(\"submission.csv\", index=False)","metadata":{"ExecuteTime":{"end_time":"2022-06-30T13:49:18.084975Z","start_time":"2022-06-30T13:49:18.074171Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:54.684105Z","iopub.execute_input":"2022-07-09T01:04:54.684445Z","iopub.status.idle":"2022-07-09T01:04:54.693094Z","shell.execute_reply.started":"2022-07-09T01:04:54.684414Z","shell.execute_reply":"2022-07-09T01:04:54.692111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1><center>Intuition</center></h1>","metadata":{"ExecuteTime":{"end_time":"2022-07-07T01:31:16.614289Z","start_time":"2022-07-07T01:31:16.600326Z"}}},{"cell_type":"markdown","source":"Intuitively we should be able to sanity check our model by taking a random set of the training rows (for which we know the survival prediction) and pass it through our model.\n\nThe model should then be able to predict(verify) correctly what we already know.\n\n## Setup\n\nLets take 2 of the rows, randomly from the training data which has one record showing the passenger survived and the other that shows the passenger wasn't as lucky.\n\n","metadata":{}},{"cell_type":"code","source":"survived_row = training_data.iloc[3] #This passenger survived\nperished_row = training_data.iloc[4] #This passenger did not survive \n\n\nintuition_df = pd.DataFrame() # createa new dataframe and add it to the rows\n\nintuition_df = intuition_df.append(survived_row)\nintuition_df = intuition_df.append(perished_row)\n\nintuition_df.head()","metadata":{"ExecuteTime":{"end_time":"2022-07-07T01:48:16.303995Z","start_time":"2022-07-07T01:48:16.273421Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:54.694405Z","iopub.execute_input":"2022-07-09T01:04:54.694720Z","iopub.status.idle":"2022-07-09T01:04:54.728163Z","shell.execute_reply.started":"2022-07-09T01:04:54.694691Z","shell.execute_reply":"2022-07-09T01:04:54.726727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Clean and Feature Engineer\n\nWe have to take the same steps that we did for the unseen data to pass it through our model.\n","metadata":{"ExecuteTime":{"end_time":"2022-06-25T14:42:35.588332Z","start_time":"2022-06-25T14:42:35.585636Z"}}},{"cell_type":"code","source":"\n\n\n#1. Clean data \nclean_intuition_df = perform_preprocessing(intuition_df)\n\n#2. Categorise the data\nclean_intuition_df[\"Age_Bucket\"] = create_age_buckets(clean_intuition_df[\"Age\"]) \nclean_intuition_df[\"Fare_Bucket\"] = create_fare_buckets(clean_intuition_df[\"Fare\"])\nclean_intuition_df[\"Family_Size\"] = create_family_size_category(clean_intuition_df)\n\n#3. Remove the unwanted / not needed columns\ndel clean_intuition_df[\"Fare\"]\ndel clean_intuition_df[\"Age\"]\ndel clean_intuition_df[\"SibSp\"]\ndel clean_intuition_df[\"Parch\"]\ndel clean_intuition_df[\"Survived\"]\n\n\nclean_intuition_df.head()","metadata":{"ExecuteTime":{"end_time":"2022-07-07T02:03:34.789322Z","start_time":"2022-07-07T02:03:34.754543Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:54.729964Z","iopub.execute_input":"2022-07-09T01:04:54.730301Z","iopub.status.idle":"2022-07-09T01:04:54.755944Z","shell.execute_reply.started":"2022-07-09T01:04:54.730269Z","shell.execute_reply":"2022-07-09T01:04:54.754714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now the above needs to be changed to the same shape as the training set and the unseen data:\n\n(891, 16)\n\n(418, 16)\n\nso basically for our inution data it would need to be \n(2, 16)\n\nTherefore the inputs need to have the following columns:","metadata":{}},{"cell_type":"code","source":"test_X.info()","metadata":{"ExecuteTime":{"end_time":"2022-07-07T06:04:04.146196Z","start_time":"2022-07-07T06:04:04.133741Z"},"execution":{"iopub.status.busy":"2022-07-09T01:04:54.758014Z","iopub.execute_input":"2022-07-09T01:04:54.758835Z","iopub.status.idle":"2022-07-09T01:04:54.772692Z","shell.execute_reply.started":"2022-07-09T01:04:54.758768Z","shell.execute_reply":"2022-07-09T01:04:54.771503Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We use the data that we have in the _intuition_df_ to then create rows similar","metadata":{"cell_style":"center"}},{"cell_type":"code","source":"row1 = {\n    \"Pclass_2\": 0,\n    \"Pclass_3\": 0,\n    \"Sex_male\": 0,\n    \"Family_Size_Single\": 0,\n    \"Family_Size_Small Family\": 1,\n    \"Family_Size_Standard Family\": 0,\n    \"Embarked_Q\": 0,\n    \"Embarked_S\": 1,\n    \"Age_Bucket_Mature Age\": 0,\n    \"Age_Bucket_Senior\": 0,\n    \"Age_Bucket_Teen\": 0,\n    \"Age_Bucket_Young Adult\": 1,\n    \"Fare_Bucket_High Fare\": 1,\n    \"Fare_Bucket_Highest Fare\": 0,\n    \"Fare_Bucket_Low Fare\": 0,\n    \"Fare_Bucket_Standard Fare\": 0\n}\n\nrow2 = {\n    \"Pclass_2\": 0,\n    \"Pclass_3\": 1,\n    \"Sex_male\": 1,\n    \"Family_Size_Single\": 1,\n    \"Family_Size_Small Family\": 0,\n    \"Family_Size_Standard Family\": 0,\n    \"Embarked_Q\": 0,\n    \"Embarked_S\": 1,\n    \"Age_Bucket_Mature Age\": 0,\n    \"Age_Bucket_Senior\": 0,\n    \"Age_Bucket_Teen\": 0,\n    \"Age_Bucket_Young Adult\": 1,\n    \"Fare_Bucket_High Fare\": 0,\n    \"Fare_Bucket_Highest Fare\": 0,\n    \"Fare_Bucket_Low Fare\": 1,\n    \"Fare_Bucket_Standard Fare\": 0\n}\n\nintuition_input = pd.DataFrame()\nintuition_input = intuition_input.append(row1, ignore_index=True)\nintuition_input = intuition_input.append(row2, ignore_index=True)\n\nintuition_input.head()\nintuition_input.info()\n\ninput_x = intuition_input.values\ninput_x = input_x.astype(np.float64, copy=False)\ninput_x = sc.fit_transform(input_x)\n\ninput_x.shape\n","metadata":{"execution":{"iopub.status.busy":"2022-07-09T01:04:54.774941Z","iopub.execute_input":"2022-07-09T01:04:54.775748Z","iopub.status.idle":"2022-07-09T01:04:54.823180Z","shell.execute_reply.started":"2022-07-09T01:04:54.775701Z","shell.execute_reply":"2022-07-09T01:04:54.822003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Run Model","metadata":{}},{"cell_type":"markdown","source":"Looking Good\n\nNow let's push this through the model and look at our predictions... same we we did for unseen","metadata":{}},{"cell_type":"code","source":"intuition_preds = predictive_model.predict(input_x)\np_asint = np.rint(intuition_preds)\n\nfor p in p_asint:\n    print(\"Survived: \" + str(p))\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-09T01:04:54.824840Z","iopub.execute_input":"2022-07-09T01:04:54.825467Z","iopub.status.idle":"2022-07-09T01:04:54.885528Z","shell.execute_reply.started":"2022-07-09T01:04:54.825430Z","shell.execute_reply":"2022-07-09T01:04:54.884315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Conclusion\n\nViola!!! The model looks like it has learnt the correct things!!.. Sure more could be tuned and better models could be found and different features could be engineered - but we are definitely on the right track.","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}