{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<a id=\"p1\"></a>\n# **1. Importing Libraries and Packages**\nWe will use these packages to help us manipulate the data and visualize the features/labels as well as measure how well our model performed. Numpy and Pandas are helpful for manipulating the dataframe and its columns and cells. We will use matplotlib along with Seaborn to visualize our data.","metadata":{"_uuid":"b5a9e650ed4ee4bcb942881e39476de033684e8f"}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \n\nimport seaborn as sns\nfrom matplotlib import pyplot as plt\nsns.set_style(\"whitegrid\")\n%matplotlib inline\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\nimport os \nprint(os.listdir(\"../input\"))","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"p2\"></a>\n# **2. Loading and Viewing Data Set**\nWith Pandas, we can load both the training and testing set that we wil later use to train and test our model. Before we begin, we should take a look at our data table to see the values that we'll be working with. We can use the head and describe function to look at some sample data and statistics. We can also look at its keys and column names.","metadata":{"_uuid":"9b574000bd7d01fc11bd1a2a129512c420ae9d60"}},{"cell_type":"code","source":"training = pd.read_csv(\"../input/train.csv\")\ntesting = pd.read_csv(\"../input/test.csv\")","metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","scrolled":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"training.head()","metadata":{"_uuid":"17d595b96b41eecfaeb9185bca85c7b5abd8d2a1","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"testing.head()","metadata":{"_uuid":"ab600060187dd559c1b05508a8a525fab6f56c50","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This data looks very messy! We're going to have to preprocess it before it's ready to be used in Machine Learning models.","metadata":{"_uuid":"a94a05068ed0bb0872ab21cf0c5f30a5d4251eda"}},{"cell_type":"code","source":"print(training.keys())\nprint(testing.keys())","metadata":{"_uuid":"3d25ce4c089559faf5ab3b78667e6201a62e57da","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"types_train = training.dtypes\nnum_values = types_train[(types_train == float)]\n\nprint(\"These are the numerical features:\")\nprint(num_values)","metadata":{"_uuid":"2460f0f9071f3a06263ca657d14656642a2d348c","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"training.describe()","metadata":{"_uuid":"a8398875d9d3d71fc7b20c3f6bf102f301c4fbe2","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":">**Note:** Only Age and Fare are the actual numerical values above and that the other features are just represented with numbers.","metadata":{"_uuid":"a94f4700f53a7850007b1c9eea2c3a1e14d4af9d"}},{"cell_type":"code","source":"training.info()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"testing.info()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"p3\"></a>\n# **3. Dealing with NaN Values (Imputation)**\nThere are NaN values in our data set in the age column. Furthermore, the Cabin column has a lot of missing values as well. These NaN values will get in the way of training our model. We need to fill in the NaN values with replacement values in order for the model to have a complete prediction for every row in the data set. This process is known as **imputation** and we will show how to replace the missing data.","metadata":{"_uuid":"e6d7efeb33bee9b004b0a309c21feb9c90bde389"}},{"cell_type":"code","source":"def null_table(training, testing):\n    print(\"Training Data Frame\")\n    print(pd.isnull(training).sum()) \n    print(\" \")\n    print(\"Testing Data Frame\")\n    print(pd.isnull(testing).sum())\n\nnull_table(training, testing)","metadata":{"_uuid":"fc61dad51cca67aafca7447fb02b58d8bbb2a226","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Wow! Cabin has a lot of missing values. Also, it seems as though the Ticket feature is too noisy to be useful. We can probably drop both features without it impacting the performance of our model.","metadata":{"_uuid":"3b6171e3a905719a4bfca530faa349695506b42d"}},{"cell_type":"code","source":"training.drop(labels = [\"Cabin\", \"Ticket\"], axis = 1, inplace = True)\ntesting.drop(labels = [\"Cabin\", \"Ticket\"], axis = 1, inplace = True)\n\nnull_table(training, testing)","metadata":{"_uuid":"4dfb986a769ce38cb258a3e8f0fc2b4c9490612f","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We take a look at the distribution of the Age column to see if it's skewed or symmetrical. This will help us determine what value to replace the NaN values.","metadata":{"_uuid":"b8cd1cb47f574b46b2058f2e3635d034ce9a57e5"}},{"cell_type":"code","source":"copy = training.copy()\ncopy.dropna(inplace = True)\nsns.distplot(copy[\"Age\"])","metadata":{"_uuid":"b791d174a23371a6a9748906771189bf45051ff2","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looks like the distribution of ages is slightly skewed right. Because of this, we can fill in the null values with the median for the most accuracy. \n> **Note:** We do not want to fill with the mean because the skewed distribution means that very large values on one end will greatly impact the mean, as opposed to the median, which will only be slightly impacted.","metadata":{"_uuid":"c04adcb431b42cb574ee0741fdb86925b83863ff"}},{"cell_type":"code","source":"#the median will be an acceptable value to place in the NaN cells\ntraining[\"Age\"].fillna(training[\"Age\"].median(), inplace = True)\ntesting[\"Age\"].fillna(testing[\"Age\"].median(), inplace = True) \ntraining[\"Embarked\"].fillna(\"S\", inplace = True)\ntesting[\"Fare\"].fillna(testing[\"Fare\"].median(), inplace = True)\n\nnull_table(training, testing)","metadata":{"_uuid":"dfa607948743e80892fdf82a000c7a96873d38ce","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Nice! No more missing values. Now let's take a look at our imputed data.","metadata":{"_uuid":"e0d9d4968c5e23024be4a9abf1979f083b057f84"}},{"cell_type":"code","source":"training.head()","metadata":{"_uuid":"2db013984d3de21c8f73d6086db83bd6e99deaa2","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"testing.head()","metadata":{"_uuid":"e72f7e7c35bc63cf6c4fe1104bc2b9bae0b79985","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"p4\"></a>\n# **4. Plotting and Visualizing Data**\nIt is very important to understand and visualize any data we are going to use in a machine learning model. By visualizing, we can see the trends and general associations of variables like Sex and Age with survival rate. We can make several different graphs for each feature we want to work with to see the entropy and information gain of the feature. ","metadata":{"_uuid":"8ae1ecf061805ccaf6663b50bf4c21252f4195a4"}},{"cell_type":"markdown","source":"**Gender**","metadata":{"_uuid":"14ac153fd1d8992bed821c28532da3dd99a3ce2d"}},{"cell_type":"code","source":"#can ignore the testing set for now\nsns.barplot(x=\"Sex\", y=\"Survived\", data=training)\nplt.title(\"Distribution of Survival based on Gender\")\nplt.show()\n\ntotal_survived_females = training[training.Sex == \"female\"][\"Survived\"].sum()\ntotal_survived_males = training[training.Sex == \"male\"][\"Survived\"].sum()\n\nprint(\"Total people survived is: \" + str((total_survived_females + total_survived_males)))\nprint(\"Proportion of Females who survived:\") \nprint(total_survived_females/(total_survived_females + total_survived_males))\nprint(\"Proportion of Males who survived:\")\nprint(total_survived_males/(total_survived_females + total_survived_males))","metadata":{"_uuid":"a313f660ca8133f5c625814a6c599f5c527a9b22","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> **Note:** The numbers printed above are the proportion of male/female survivors of all the surviviors ONLY. The graph shows the propotion of male/females out of ALL the passengers including those that didn't survive.","metadata":{"_uuid":"e924d85d07f0da76ad085c5ca8f26172a80ad987"}},{"cell_type":"markdown","source":"Gender appears to be a very good feature to use to predict survival, as shown by the large difference in propotion survived. Let's take a look at how class plays a role in survival as well.","metadata":{"_uuid":"69f48dadbaa468414c542c74af8bfe98a68ec20d"}},{"cell_type":"markdown","source":"**Class**","metadata":{"_uuid":"8479cb62de593c2945b112b645a527be7aa776f3"}},{"cell_type":"code","source":"sns.barplot(x=\"Pclass\", y=\"Survived\", data=training)\nplt.ylabel(\"Survival Rate\")\nplt.title(\"Distribution of Survival Based on Class\")\nplt.show()\n\ntotal_survived_one = training[training.Pclass == 1][\"Survived\"].sum()\ntotal_survived_two = training[training.Pclass == 2][\"Survived\"].sum()\ntotal_survived_three = training[training.Pclass == 3][\"Survived\"].sum()\ntotal_survived_class = total_survived_one + total_survived_two + total_survived_three\n\nprint(\"Total people survived is: \" + str(total_survived_class))\nprint(\"Proportion of Class 1 Passengers who survived:\") \nprint(total_survived_one/total_survived_class)\nprint(\"Proportion of Class 2 Passengers who survived:\")\nprint(total_survived_two/total_survived_class)\nprint(\"Proportion of Class 3 Passengers who survived:\")\nprint(total_survived_three/total_survived_class)","metadata":{"_uuid":"d55d78e57c418f2ea9f57c6ae649dc9f044ef553","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.barplot(x=\"Pclass\", y=\"Survived\", hue=\"Sex\", data=training)\nplt.ylabel(\"Survival Rate\")\nplt.title(\"Survival Rates Based on Gender and Class\")","metadata":{"_uuid":"205515b33d6dd0962cab4fb825caa45c77e5ec12","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.barplot(x=\"Sex\", y=\"Survived\", hue=\"Pclass\", data=training)\nplt.ylabel(\"Survival Rate\")\nplt.title(\"Survival Rates Based on Gender and Class\")","metadata":{"_uuid":"ac54f705c3fc462445cad2de978d7077646f933a","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It appears that class also plays a role in survival, as shown by the bar graph. People in Pclass 1 were more likely to survive than people in the other 2 Pclasses.","metadata":{"_uuid":"13dde8426cb8a654e6f50cb1855e4369c362ea68"}},{"cell_type":"markdown","source":"**Age**","metadata":{"_uuid":"ad10453d5d42100bfb65be57900ff1554508a95c"}},{"cell_type":"code","source":"survived_ages = training[training.Survived == 1][\"Age\"]\nnot_survived_ages = training[training.Survived == 0][\"Age\"]\nplt.subplot(1, 2, 1)\nsns.distplot(survived_ages, kde=False)\nplt.axis([0, 100, 0, 100])\nplt.title(\"Survived\")\nplt.ylabel(\"Proportion\")\nplt.subplot(1, 2, 2)\nsns.distplot(not_survived_ages, kde=False)\nplt.axis([0, 100, 0, 100])\nplt.title(\"Didn't Survive\")\nplt.subplots_adjust(right=1.7)\nplt.show()","metadata":{"_uuid":"323af5a70d2033f9fa1b5cc8286a675869ba302e","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.stripplot(x=\"Survived\", y=\"Age\", data=training, jitter=True)","metadata":{"_uuid":"46374be19c92dbeb146493226b47701820a1bc7f","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It appears as though passengers in the younger range of ages were more likely to survive than those in the older range of ages, as seen by the clustering in the strip plot, as well as the survival distributions of the histogram.","metadata":{"_uuid":"56bcea2157a75ed314388f1804aa2271cc631f25"}},{"cell_type":"markdown","source":"Here is one final cumulative graph of a pair plot that shows the relations between all of the different features","metadata":{"_uuid":"60d068f1155860f57f3395bd8bc6b752c050c8e5"}},{"cell_type":"code","source":"sns.pairplot(training)","metadata":{"_uuid":"7d98aac15ffac0c0a3282961850dabc42ae79b1c","trusted":true},"execution_count":null,"outputs":[]}]}