{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### Objectives: \n\nIn this case study my objectives are as follows: \n\n1) Find the people that were most likely to survive the disaster. \n\n2) Find out through the data why those that survived did. \n","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nfrom matplotlib.ticker import PercentFormatter\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-21T07:53:25.662130Z","iopub.execute_input":"2022-07-21T07:53:25.662793Z","iopub.status.idle":"2022-07-21T07:53:26.415278Z","shell.execute_reply.started":"2022-07-21T07:53:25.662653Z","shell.execute_reply":"2022-07-21T07:53:26.413892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"First, I load the datasets into a pandas dataframe.","metadata":{}},{"cell_type":"code","source":"training = pd.read_csv('/kaggle/input/titanic/train.csv')\ntest = pd.read_csv('/kaggle/input/titanic/test.csv')\n\ntraining['train_test'] = 1\ntest['train_test'] = 0\ntest['Survived'] = np.NaN\nall_data = pd.concat([training,test])\n\n%matplotlib inline\nall_data.columns","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:26.417671Z","iopub.execute_input":"2022-07-21T07:53:26.418100Z","iopub.status.idle":"2022-07-21T07:53:26.476436Z","shell.execute_reply.started":"2022-07-21T07:53:26.418065Z","shell.execute_reply":"2022-07-21T07:53:26.475157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(training)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:26.478167Z","iopub.execute_input":"2022-07-21T07:53:26.478878Z","iopub.status.idle":"2022-07-21T07:53:26.512775Z","shell.execute_reply.started":"2022-07-21T07:53:26.478753Z","shell.execute_reply":"2022-07-21T07:53:26.511741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Then I try to get a better understanding of the types of data that I will be working with. ","metadata":{}},{"cell_type":"code","source":"training.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:26.514269Z","iopub.execute_input":"2022-07-21T07:53:26.514643Z","iopub.status.idle":"2022-07-21T07:53:26.536957Z","shell.execute_reply.started":"2022-07-21T07:53:26.514580Z","shell.execute_reply":"2022-07-21T07:53:26.535540Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Cleaning \n\nFirst, I'm going to use a heat map to try and understand how many null values exist in the data. \n","metadata":{}},{"cell_type":"code","source":"sns.heatmap(training.isnull(),yticklabels=False,cbar=False,cmap=\"magma\")","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:26.540257Z","iopub.execute_input":"2022-07-21T07:53:26.540917Z","iopub.status.idle":"2022-07-21T07:53:26.808906Z","shell.execute_reply.started":"2022-07-21T07:53:26.540879Z","shell.execute_reply":"2022-07-21T07:53:26.807293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The Cabin variable has almost all missing values so I won't be able to extract any useful informationn out of them. So, I will remove it from the training datafram all together. ","metadata":{}},{"cell_type":"markdown","source":"There are some passengers that have their age missing as well. I will fill them in using the mean age of the place that they embarked from. ","metadata":{}},{"cell_type":"code","source":"sns.boxplot(x='Embarked',y='Age',data=training,palette='rocket')\n","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:26.810780Z","iopub.execute_input":"2022-07-21T07:53:26.811167Z","iopub.status.idle":"2022-07-21T07:53:26.999254Z","shell.execute_reply.started":"2022-07-21T07:53:26.811133Z","shell.execute_reply":"2022-07-21T07:53:26.998361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The box plot tells me the median age values for the passengers grouped by embarkment point. I will now write a function that can add these median values to their respective places.","metadata":{}},{"cell_type":"code","source":"def impute(cols):\n    Age = cols[0]\n    Embarked = cols[1]\n    \n    if pd.isnull(Age):\n        if Embarked == 'S':\n            return 28\n        elif Embarked == 'C':\n            return 29\n        else:\n            return 27\n        \n    else:\n        return Age","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:27.000891Z","iopub.execute_input":"2022-07-21T07:53:27.002030Z","iopub.status.idle":"2022-07-21T07:53:27.009154Z","shell.execute_reply.started":"2022-07-21T07:53:27.001978Z","shell.execute_reply":"2022-07-21T07:53:27.008029Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, I will apply the code from above to the age column to fill in the missing values.","metadata":{}},{"cell_type":"code","source":"training['Age']=training[['Age','Pclass']].apply(impute,axis=1)\n\ntraining.isnull().sum(axis = 0)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:27.010719Z","iopub.execute_input":"2022-07-21T07:53:27.012059Z","iopub.status.idle":"2022-07-21T07:53:27.048185Z","shell.execute_reply.started":"2022-07-21T07:53:27.012018Z","shell.execute_reply":"2022-07-21T07:53:27.046393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Cabin has 687 null, so analysis of it will likely not reveal too much.","metadata":{}},{"cell_type":"markdown","source":"## Data Exploration and Analysis \n\nNow, I will get a better understanding of the data. I will do this by seperating the categorical and numeric variables and exploring them seperately. ","metadata":{}},{"cell_type":"code","source":"numeric = training[['Age','SibSp','Fare','Parch']]\ncategorical = training[['Survived','Pclass','Sex','Ticket','Embarked']]","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:27.049738Z","iopub.execute_input":"2022-07-21T07:53:27.050096Z","iopub.status.idle":"2022-07-21T07:53:27.058445Z","shell.execute_reply.started":"2022-07-21T07:53:27.050066Z","shell.execute_reply":"2022-07-21T07:53:27.057552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I will explore all of the numeric data first. To do this, I will create histograms of each to understand how the data is distributed. ","metadata":{}},{"cell_type":"code","source":"numeric.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:27.059801Z","iopub.execute_input":"2022-07-21T07:53:27.060640Z","iopub.status.idle":"2022-07-21T07:53:27.095129Z","shell.execute_reply.started":"2022-07-21T07:53:27.060581Z","shell.execute_reply":"2022-07-21T07:53:27.093971Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in numeric.columns:\n    plt.hist(numeric[i],color = 'black')\n    plt.title(i)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:27.096618Z","iopub.execute_input":"2022-07-21T07:53:27.097383Z","iopub.status.idle":"2022-07-21T07:53:27.932297Z","shell.execute_reply.started":"2022-07-21T07:53:27.097336Z","shell.execute_reply":"2022-07-21T07:53:27.931107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, I will create bar graphs to understand the categorical variables better.","metadata":{}},{"cell_type":"code","source":"for i in categorical.columns:\n    sns.barplot(categorical[i].value_counts().index,categorical[i].value_counts()).set_title(i)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:27.934106Z","iopub.execute_input":"2022-07-21T07:53:27.934574Z","iopub.status.idle":"2022-07-21T07:53:37.400549Z","shell.execute_reply.started":"2022-07-21T07:53:27.934526Z","shell.execute_reply":"2022-07-21T07:53:37.399180Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, I'd like to explore how some of these variables are related to the survived category. In particular, I am interested in how Age, Sex, Pclass, Fare, embarkment location, and ticket number are related to whether or not you survived. ","metadata":{}},{"cell_type":"markdown","source":"#### Pclass:","metadata":{}},{"cell_type":"code","source":"survivedByClass = pd.pivot_table(training,index='Survived', columns='Pclass', values = 'PassengerId', aggfunc='count')\nsurvivedByClass.plot(kind='bar')","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:37.402073Z","iopub.execute_input":"2022-07-21T07:53:37.402439Z","iopub.status.idle":"2022-07-21T07:53:37.652810Z","shell.execute_reply.started":"2022-07-21T07:53:37.402406Z","shell.execute_reply":"2022-07-21T07:53:37.651481Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(survivedByClass)\nprint(\" \")\nprint(\"Survival percentage in Class 1:\",(136/(80+136))*100)\nprint(\"Survival percentage in Class 2:\",(87/(97+87))*100)\nprint(\"Survival percentage in Class 3:\",(119/(372+119))*100)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:37.657608Z","iopub.execute_input":"2022-07-21T07:53:37.658536Z","iopub.status.idle":"2022-07-21T07:53:37.670199Z","shell.execute_reply.started":"2022-07-21T07:53:37.658494Z","shell.execute_reply":"2022-07-21T07:53:37.668800Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"When comparing people in first class (1) to the lower classes, it is clear that if you were in a higher class, your odds of survival are greater. In the lowest class, only about 25% of the passengers survived, while in the middle class, there was about a 47% chance of survival, and if you were in the highest class, you had a 63% chance of survival.","metadata":{}},{"cell_type":"markdown","source":"#### Embarked:","metadata":{}},{"cell_type":"code","source":"survivedByEmbarkment = pd.pivot_table(training,index='Survived', columns='Embarked', values = 'PassengerId', aggfunc='count')\nsurvivedByEmbarkment.plot(kind='bar')","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:37.671923Z","iopub.execute_input":"2022-07-21T07:53:37.672807Z","iopub.status.idle":"2022-07-21T07:53:37.864376Z","shell.execute_reply.started":"2022-07-21T07:53:37.672756Z","shell.execute_reply":"2022-07-21T07:53:37.863436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(survivedByEmbarkment)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:37.865689Z","iopub.execute_input":"2022-07-21T07:53:37.866451Z","iopub.status.idle":"2022-07-21T07:53:37.873479Z","shell.execute_reply.started":"2022-07-21T07:53:37.866417Z","shell.execute_reply":"2022-07-21T07:53:37.872298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Survival percentage in Cherbourg Embarkment:\",(93/(93+75))*100)\nprint(\"Survival percentage in Queenstown Embarkmet:\",(30/(30+47))*100)\nprint(\"Survival percentage in Southampton Embarkment:\",(217/(217+427))*100)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:37.875515Z","iopub.execute_input":"2022-07-21T07:53:37.876563Z","iopub.status.idle":"2022-07-21T07:53:37.887245Z","shell.execute_reply.started":"2022-07-21T07:53:37.876511Z","shell.execute_reply":"2022-07-21T07:53:37.886145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The location of where the passengers embarked from doesn't seem to be related to survival rate. The ratios are relatively the same.","metadata":{}},{"cell_type":"markdown","source":"#### Age: ","metadata":{}},{"cell_type":"markdown","source":"to analyze how age is related to survival, I will first divide the passengers into four age groups, kids, young adults, adults, and senior citizens. The I will look into the survival rate of each of the cohorts. ","metadata":{}},{"cell_type":"code","source":"kids = training.loc[(training.Age <=14)]\nkids.info()\n\nyoung = training.loc[(training.Age <= 25)&(training.Age > 14)]\nyoung.info()\n\nadults = training.loc[(training.Age > 25) & (training.Age < 65)]\nadults.info()\n\nseniors = training.loc[(training.Age >=60)]\nseniors.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:37.888690Z","iopub.execute_input":"2022-07-21T07:53:37.889045Z","iopub.status.idle":"2022-07-21T07:53:37.941162Z","shell.execute_reply.started":"2022-07-21T07:53:37.889012Z","shell.execute_reply":"2022-07-21T07:53:37.940001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"survivedKids = kids.loc[(kids.Survived==1)]\nprint(\"Kids Survival Rate:\",(len(survivedKids)/len(kids))*100, \"%\")\n\nsurvivedYoung = young.loc[(young.Survived==1)]\nprint(\"Young Adults Survival Rate:\",(len(survivedYoung)/len(young))*100, \"%\")\n\nsurvivedAdults = adults.loc[(adults.Survived==1)]\nprint(\"Adults Survival Rate:\",(len(survivedAdults)/len(adults))*100, \"%\")\n\nsurvivedSeniors = seniors.loc[(seniors.Survived==1)]\nprint(\"Seniors Survival Rate:\",(len(survivedSeniors)/len(seniors))*100, \"%\")\n","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:37.942882Z","iopub.execute_input":"2022-07-21T07:53:37.943455Z","iopub.status.idle":"2022-07-21T07:53:37.955776Z","shell.execute_reply.started":"2022-07-21T07:53:37.943423Z","shell.execute_reply":"2022-07-21T07:53:37.954852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The survival rate of kids was much higher than those of adults, indicating that their lives were put in priority when the ship went down. Now, I'd like to compare children in the different classes and see if whether they were in class 3,2, or 1 determined how likely they were to survive. ","metadata":{}},{"cell_type":"code","source":"p3Kids = kids.loc[(kids.Pclass == 3)]\np3KidsSurvived = p3Kids.loc[(p3Kids.Survived == 1)]\nprint(\"Number of kids in class 3:\",len(p3Kids))\nprint(\"Percent of kids in class 3 that survived: \",len(p3KidsSurvived)/len(p3Kids)*100)\n\n\np2Kids = kids.loc[(kids.Pclass == 2)]\np2KidsSurvived = p2Kids.loc[(p2Kids.Survived == 1)]\nprint(\"Number of kids in class 2:\",len(p2Kids))\nprint(\"Percent of kids in class 2 that survived: \",len(p2KidsSurvived)/len(p2Kids)*100)\n\np1Kids = kids.loc[(kids.Pclass == 1)]\np1KidsSurvived = p1Kids.loc[(p1Kids.Survived == 1)]\nprint(\"Number of kids in class 1:\",len(p1Kids))\nprint(\"Percent of kids in class 1 that survived: \",len(p1KidsSurvived)/len(p1Kids)*100)\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:37.957089Z","iopub.execute_input":"2022-07-21T07:53:37.957581Z","iopub.status.idle":"2022-07-21T07:53:37.974960Z","shell.execute_reply.started":"2022-07-21T07:53:37.957550Z","shell.execute_reply":"2022-07-21T07:53:37.973665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Although the sample size of kids in their respective classes is different, the number of kids that survived in the lowest class, class 3, is significantly lower than those in class 2 or in class 1. The reason for this could be two. Either the class three residents resided in a place on the ship that resulted in more difficulty to escape, and/or the members in class 3 got the least amount of support by the crew in their attempts to escape. ","metadata":{}},{"cell_type":"markdown","source":"#### Fares:","metadata":{}},{"cell_type":"code","source":"sns.scatterplot(data=training, x=\"Pclass\", y=\"Fare\")","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:37.976342Z","iopub.execute_input":"2022-07-21T07:53:37.976881Z","iopub.status.idle":"2022-07-21T07:53:38.177092Z","shell.execute_reply.started":"2022-07-21T07:53:37.976840Z","shell.execute_reply":"2022-07-21T07:53:38.175844Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So, first of all, through the scatterplot above, it is clear that those that paid more for their ticket, ended up in a better class. Now, I'd like to see how the amount paid, is related to the passengers' survival. ","metadata":{}},{"cell_type":"code","source":"sns.boxplot(data=training, x=\"Survived\", y=\"Fare\")","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:38.178637Z","iopub.execute_input":"2022-07-21T07:53:38.179505Z","iopub.status.idle":"2022-07-21T07:53:38.360529Z","shell.execute_reply.started":"2022-07-21T07:53:38.179465Z","shell.execute_reply":"2022-07-21T07:53:38.358512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"survived = training.loc[(training.Survived == 1)]\ndied = training.loc[(training.Survived == 0)]\n\n\nprint(\"Mean fare of those that died: \",died.Fare.mean())\nprint(\"Mean fare of those that survived: \", survived.Fare.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:38.362893Z","iopub.execute_input":"2022-07-21T07:53:38.364198Z","iopub.status.idle":"2022-07-21T07:53:38.375668Z","shell.execute_reply.started":"2022-07-21T07:53:38.364140Z","shell.execute_reply":"2022-07-21T07:53:38.374388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Those that survived paid an average of almost double of those that died. My next step will be to try to see if there is a correlation between the fare, the cabin number, and whether one survived or not. This could potentially give an indication as to whether those located in a certain cabin had a better chance to survive. ","metadata":{}},{"cell_type":"code","source":"nullCabins = training.loc[(training.Cabin.isnull())]\ndisplay(nullCabins)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:38.377415Z","iopub.execute_input":"2022-07-21T07:53:38.377993Z","iopub.status.idle":"2022-07-21T07:53:38.406541Z","shell.execute_reply.started":"2022-07-21T07:53:38.377955Z","shell.execute_reply":"2022-07-21T07:53:38.405512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"nonNullCabins = training.loc[(training.Cabin.notnull())]\nprint(display(nonNullCabins))\n","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:38.407968Z","iopub.execute_input":"2022-07-21T07:53:38.408479Z","iopub.status.idle":"2022-07-21T07:53:38.438271Z","shell.execute_reply.started":"2022-07-21T07:53:38.408447Z","shell.execute_reply":"2022-07-21T07:53:38.436877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 687 cabins that are filled with a null value and only 204 with non-null values. So, analysis will have to be done using only the cabins that have values in it. \n\nThe cabins have a letter before them, so I will use that letter to group them and see how they relate to fare, pclass, and survival. But, first I need to seperate out the letter from the numbers. ","metadata":{}},{"cell_type":"code","source":"nonNullCabins['cabinLetter'] = training.Cabin.apply(lambda x: str(x)[0])\nprint(nonNullCabins.cabinLetter.value_counts())","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:38.440057Z","iopub.execute_input":"2022-07-21T07:53:38.440428Z","iopub.status.idle":"2022-07-21T07:53:38.451653Z","shell.execute_reply.started":"2022-07-21T07:53:38.440394Z","shell.execute_reply":"2022-07-21T07:53:38.450367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The cabin T only has one non null value in it, so I will be dropping it for my analysis.","metadata":{}},{"cell_type":"code","source":"nonNullCabins = nonNullCabins.drop(nonNullCabins[nonNullCabins[\"cabinLetter\"]=='T'].index)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:38.453537Z","iopub.execute_input":"2022-07-21T07:53:38.454043Z","iopub.status.idle":"2022-07-21T07:53:38.467212Z","shell.execute_reply.started":"2022-07-21T07:53:38.453997Z","shell.execute_reply":"2022-07-21T07:53:38.465908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, I will create a dataframe in which the cabin letters are grouped by the percent survived of each cabin and by the mean fare paid at each cabin.","metadata":{}},{"cell_type":"code","source":"grouping = nonNullCabins.groupby('cabinLetter', as_index=False)['Fare'].mean().sort_values('Fare')\ngrouping2=nonNullCabins.groupby('cabinLetter')['Survived'].apply(lambda x: (x==1).sum()).reset_index(name='count')\ngrouping3=nonNullCabins.groupby('cabinLetter')['Survived'].apply(lambda x: (x==0).sum()).reset_index(name='count')","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:38.468886Z","iopub.execute_input":"2022-07-21T07:53:38.469358Z","iopub.status.idle":"2022-07-21T07:53:38.494675Z","shell.execute_reply.started":"2022-07-21T07:53:38.469313Z","shell.execute_reply":"2022-07-21T07:53:38.493688Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grouping['percentSurvived'] = (grouping2['count'] / (grouping2['count']+grouping3['count'])*100)\nprint(grouping)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:38.496161Z","iopub.execute_input":"2022-07-21T07:53:38.497265Z","iopub.status.idle":"2022-07-21T07:53:38.510830Z","shell.execute_reply.started":"2022-07-21T07:53:38.497214Z","shell.execute_reply":"2022-07-21T07:53:38.509867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ax = sns.barplot(x = 'cabinLetter', y= 'Fare',data = grouping)\nax2 = ax.twinx()\nsns.lineplot(x='cabinLetter', y= 'percentSurvived',data = grouping, lw=3, ax=ax2,color='black')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:38.512237Z","iopub.execute_input":"2022-07-21T07:53:38.513467Z","iopub.status.idle":"2022-07-21T07:53:38.825750Z","shell.execute_reply.started":"2022-07-21T07:53:38.513419Z","shell.execute_reply":"2022-07-21T07:53:38.824691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the dual axis chart above, the line represents the percentage of people that survived in their respective cabins, while the bar represents the fare paid for the cabin. There is a general upward trend, in that the more you paid, the higher chance you survived. Cabins G,F,A have only 4,13,and 15 data points respectively, which could explain the volatility in the line graph. If there were less null values, a more accurate upward trend would be able to be detected.","metadata":{}},{"cell_type":"code","source":"pd.pivot_table(nonNullCabins,index='Survived',columns='cabinLetter', values = 'Ticket', aggfunc='count')","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:38.827208Z","iopub.execute_input":"2022-07-21T07:53:38.827545Z","iopub.status.idle":"2022-07-21T07:53:38.855751Z","shell.execute_reply.started":"2022-07-21T07:53:38.827517Z","shell.execute_reply":"2022-07-21T07:53:38.854524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In all of the cabins, other than the null values, the number of survivers is higher or equal to the number of people that died. The reason for this could be that the ones who perished didn't have their cabin number to provide. Because there were so many null cabin values, it isn't possible to find a relationship between survival and the cabin in which people resided in.","metadata":{}},{"cell_type":"markdown","source":"#### Sex:\n","metadata":{}},{"cell_type":"code","source":"survivedBySex = pd.pivot_table(training,index='Survived', columns='Sex', values = 'PassengerId', aggfunc='count')\nsurvivedBySex.plot(kind='bar')\nprint(survivedBySex)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:38.857348Z","iopub.execute_input":"2022-07-21T07:53:38.858283Z","iopub.status.idle":"2022-07-21T07:53:39.059686Z","shell.execute_reply.started":"2022-07-21T07:53:38.858239Z","shell.execute_reply":"2022-07-21T07:53:39.058112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print((233/(81+233))*100,'% of women survived.')\nprint((109/(109+468))*100,'% of women survived.')","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:39.061153Z","iopub.execute_input":"2022-07-21T07:53:39.063065Z","iopub.status.idle":"2022-07-21T07:53:39.070475Z","shell.execute_reply.started":"2022-07-21T07:53:39.063020Z","shell.execute_reply":"2022-07-21T07:53:39.068618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Women clearly survived at a much greater rate than the me did on the Titanic. Now I will further filter women by class and see if how their survival rate defers among their class. ","metadata":{}},{"cell_type":"code","source":"Female = training.loc[(training.Sex=='female')]\n\np1Female = Female.loc[(Female.Pclass == 1)]\np1FemaleSurvived = p1Female.loc[(p1Female.Survived == 1)]\n\np2Female = Female.loc[(Female.Pclass == 2)]\np2FemaleSurvived = p2Female.loc[(p2Female.Survived == 1)]\n\np3Female = Female.loc[(Female.Pclass == 3)]\np3FemaleSurvived = p3Female.loc[(p3Female.Survived == 1)]\n\n\nprint(\"Number of females in class 1:\", len(p1Female))\nprint(\"Percent of females in class 1 that survived: \",len(p1FemaleSurvived)/len(p1Female)*100)\nprint(\" \")\nprint(\"Number of females in class 2:\", len(p2Female))\nprint(\"Percent of females in class 2 that survived: \",len(p2FemaleSurvived)/len(p2Female)*100)\nprint(\" \")\nprint(\"Number of females in class 3:\", len(p3Female))\nprint(\"Percent of females in class 3 that survived: \",len(p3FemaleSurvived)/len(p3Female)*100)\n\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:39.072084Z","iopub.execute_input":"2022-07-21T07:53:39.072709Z","iopub.status.idle":"2022-07-21T07:53:39.090612Z","shell.execute_reply.started":"2022-07-21T07:53:39.072659Z","shell.execute_reply":"2022-07-21T07:53:39.089646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"74% of total women survived the crash. But, when diving into the women seperated by their classes, class 3 women, drops to 50%. This finding further explains that being in class 1 or 2 significantly increases chance of survival with a 96% and 92% survival rate among women. Now, I will do the same analysis with the men aboard the ship. ","metadata":{}},{"cell_type":"code","source":"Male = training.loc[(training.Sex=='male')]\n\np1Male = Male.loc[(Male.Pclass == 1)]\np1MaleSurvived = p1Male.loc[(p1Male.Survived == 1)]\n\np2Male = Male.loc[(Male.Pclass == 2)]\np2MaleSurvived = p2Male.loc[(p2Male.Survived == 1)]\n\np3Male = Male.loc[(Male.Pclass == 3)]\np3MaleSurvived = p3Male.loc[(p3Male.Survived == 1)]\n\n\nprint(\"Number of males in class 1:\", len(p1Male))\nprint(\"Percent of males in class 1 that survived: \",len(p3MaleSurvived)/len(p1Male)*100)\nprint(\" \")\nprint(\"Number of males in class 2:\", len(p2Male))\nprint(\"Percent of males in class 2 that survived: \",len(p2MaleSurvived)/len(p2Male)*100)\nprint(\" \")\nprint(\"Number of males in class 3:\", len(p3Male))\nprint(\"Percent of males in class 3 that survived: \",len(p1MaleSurvived)/len(p3Male)*100)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:53:39.091824Z","iopub.execute_input":"2022-07-21T07:53:39.092917Z","iopub.status.idle":"2022-07-21T07:53:39.113715Z","shell.execute_reply.started":"2022-07-21T07:53:39.092879Z","shell.execute_reply":"2022-07-21T07:53:39.112340Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Although a lot lower, similar analysis of the males on the ship indicate similar results. With passengers in class 1 having a 39% survival rate, class 2 having 16% and class 3 having 13%.","metadata":{}},{"cell_type":"markdown","source":"#### Data Exploration Conclusion\n\n\nThe following variables are key in determinig whether a passenger survived or not: Age, Sex, Pclass, and Fare. \nThe younger you are the more likely you are to survive. If you were a women, you were more likely to survive. If you paid more and or were in a higher class, you were more likely to survive. If you combine the variables that make it more likely to survive, your chances would be higher. For example, if you were a female child, in a high class, your likelihood of survival is quite high in comparisonn to an older male in the lowest class. \n","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}