{"cells":[{"metadata":{"_cell_guid":"5767a33c-8f18-4034-e52d-bf7a8f7d8ab8","_uuid":"f740edfefd28b012fd6e603f22ac741561664dfb","trusted":true},"cell_type":"code","source":"# data analysis and wrangling\nimport pandas as pd\nimport numpy as np\nimport random as rnd\n\n# visualization\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\n# machine learning\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.svm import SVC, LinearSVC\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.naive_bayes import GaussianNB\nfrom sklearn.linear_model import Perceptron\nfrom sklearn.linear_model import SGDClassifier\nfrom sklearn.tree import DecisionTreeClassifier","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"6b5dc743-15b1-aac6-405e-081def6ecca1","_uuid":"18881c681759caf91f8928da1246102d9cc0facf"},"cell_type":"markdown","source":"## Acquire data\n\nThe Python Pandas packages helps us work with our datasets. We start by acquiring the training and testing datasets into Pandas DataFrames. We also combine these datasets to run certain operations on both datasets together."},{"metadata":{"_cell_guid":"e7319668-86fe-8adc-438d-0eef3fd0a982","_uuid":"91877fb86c005ce0b59cb65132e770dbe3edc6dc","trusted":true},"cell_type":"code","source":"train_df = pd.read_csv('../input/train.csv')\ntest_df = pd.read_csv('../input/test.csv')\ncombine = [train_df, test_df]","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"3d6188f3-dc82-8ae6-dabd-83e28fcbf10d","_uuid":"383413307f109de2f941e57635ae4df4f7b12e10"},"cell_type":"markdown","source":"## Analyze by describing data\n\nPandas also helps describe the datasets answering following questions early in our project.\n\n**Which features are available in the dataset?**\n\nNoting the feature names for directly manipulating or analyzing these. These feature names are described on the [Kaggle data page here](https://www.kaggle.com/c/titanic/data)."},{"metadata":{"_cell_guid":"ce473d29-8d19-76b8-24a4-48c217286e42","_uuid":"6f52ab20c6d4d6ce4f2c3394cf05a615519c4168","trusted":true},"cell_type":"code","source":"print(train_df.columns.values)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"cd19a6f6-347f-be19-607b-dca950590b37","_uuid":"e7a20450d7415e6aaed2cef3061d9cce7776e6e5"},"cell_type":"markdown","source":"**Which features are categorical?**\n\nThese values classify the samples into sets of similar samples. Within categorical features are the values nominal, ordinal, ratio, or interval based? Among other things this helps us select the appropriate plots for visualization.\n\n- Categorical: Survived, Sex, and Embarked. Ordinal: Pclass.\n\n**Which features are numerical?**\n\nWhich features are numerical? These values change from sample to sample. Within numerical features are the values discrete, continuous, or timeseries based? Among other things this helps us select the appropriate plots for visualization.\n\n- Continous: Age, Fare. Discrete: SibSp, Parch."},{"metadata":{"_cell_guid":"8d7ac195-ac1a-30a4-3f3f-80b8cf2c1c0f","_uuid":"0d09e73f695ba4e9b46b37c172323ddef22dc000","trusted":true},"cell_type":"code","source":"# preview the data\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"97f4e6f8-2fea-46c4-e4e8-b69062ee3d46","_uuid":"6d40ca779917cc92b0b3b666abb277f73d791b1a"},"cell_type":"markdown","source":"**Which features are mixed data types?**\n\nNumerical, alphanumeric data within same feature. These are candidates for correcting goal.\n\n- Ticket is a mix of numeric and alphanumeric data types. Cabin is alphanumeric.\n\n**Which features may contain errors or typos?**\n\nThis is harder to review for a large dataset, however reviewing a few samples from a smaller dataset may just tell us outright, which features may require correcting.\n\n- Name feature may contain errors or typos as there are several ways used to describe a name including titles, round brackets, and quotes used for alternative or short names."},{"metadata":{"_cell_guid":"f6e761c2-e2ff-d300-164c-af257083bb46","_uuid":"854cdd6b938f35f1db4057f733ce04e10e66b453","trusted":true},"cell_type":"code","source":"train_df.tail()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"8bfe9610-689a-29b2-26ee-f67cd4719079","_uuid":"1f6cfa4266456708b205981a863705c7fcc9bc86"},"cell_type":"markdown","source":"**Which features contain blank, null or empty values?**\n\nThese will require correcting.\n\n- Cabin > Age > Embarked features contain a number of null values in that order for the training dataset.\n- Cabin > Age are incomplete in case of test dataset.\n\n**What are the data types for various features?**\n\nHelping us during converting goal.\n\n- Seven features are integer or floats. Six in case of test dataset.\n- Five features are strings (object)."},{"metadata":{"_cell_guid":"9b805f69-665a-2b2e-f31d-50d87d52865d","_uuid":"16dba7c150bd03e6f2fbdcc464633976b16c190f","trusted":true},"cell_type":"code","source":"train_df.info()\nprint('_'*40)\ntest_df.info()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"859102e1-10df-d451-2649-2d4571e5f082","_uuid":"7851ba31ea9848b932b544e63eee0bf786bbcb8b"},"cell_type":"markdown","source":"**What is the distribution of numerical feature values across the samples?**\n\nThis helps us determine, among other early insights, how representative is the training dataset of the actual problem domain.\n\n- Total samples are 891 or 40% of the actual number of passengers on board the Titanic (2,224).\n- Survived is a categorical feature with 0 or 1 values.\n- Around 38% samples survived representative of the actual survival rate at 32%.\n- Most passengers (> 75%) did not travel with parents or children.\n- Nearly 30% of the passengers had siblings and/or spouse aboard.\n- Fares varied significantly with few passengers (<1%) paying as high as $512.\n- Few elderly passengers (<1%) within age range 65-80."},{"metadata":{"_cell_guid":"58e387fe-86e4-e068-8307-70e37fe3f37b","_uuid":"904fc518e4c13983721e2e0fef766c4a9f092152","trusted":true},"cell_type":"code","source":"train_df.describe()\n# Review survived rate using `percentiles=[.61, .62]` knowing our problem description mentions 38% survival rate.\n# Review Parch distribution using `percentiles=[.75, .8]`\n# SibSp distribution `[.68, .69]`\n# Age and Fare `[.1, .2, .3, .4, .5, .6, .7, .8, .9, .99]`","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"5462bc60-258c-76bf-0a73-9adc00a2f493","_uuid":"40cef093ec980cd84abccad7dbe90cec5c93c620"},"cell_type":"markdown","source":"**What is the distribution of categorical features?**\n\n- Names are unique across the dataset (count=unique=891)\n- Sex variable as two possible values with 65% male (top=male, freq=577/count=891).\n- Cabin values have several dupicates across samples. Alternatively several passengers shared a cabin.\n- Embarked takes three possible values. S port used by most passengers (top=S)\n- Ticket feature has high ratio (22%) of duplicate values (unique=681)."},{"metadata":{"_cell_guid":"8066b378-1964-92e8-1352-dcac934c6af3","_uuid":"fa3f147f565b94b70d544f8f4280710baa01b9d0","trusted":true},"cell_type":"code","source":"train_df.describe(include=['O'])","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"2cb22b88-937d-6f14-8b06-ea3361357889","_uuid":"effb69910d0c4908f8c7cda1af23070c20787b91"},"cell_type":"markdown","source":"### Assumtions based on data analysis\n\nWe arrive at following assumptions based on data analysis done so far. We may validate these assumptions further before taking appropriate actions.\n\n**Correlating.**\n\nWe want to know how well does each feature correlate with Survival. We want to do this early in our project and match these quick correlations with modelled correlations later in the project.\n\n**Completing.**\n\n1. We may want to complete Age feature as it is definitely correlated to survival.\n2. We may want to complete the Embarked feature as it may also correlate with survival or another important feature.\n\n**Correcting.**\n\n1. Ticket feature may be dropped from our analysis as it contains high ratio of duplicates (22%) and there may not be a correlation between Ticket and survival.\n2. Cabin feature may be dropped as it is highly incomplete or contains many null values both in training and test dataset.\n3. PassengerId may be dropped from training dataset as it does not contribute to survival.\n4. Name feature is relatively non-standard, may not contribute directly to survival, so maybe dropped.\n\n**Creating.**\n\n1. We may want to create a new feature called Family based on Parch and SibSp to get total count of family members on board.\n2. We may want to engineer the Name feature to extract Title as a new feature.\n3. We may want to create new feature for Age bands. This turns a continous numerical feature into an ordinal categorical feature.\n4. We may also want to create a Fare range feature if it helps our analysis.\n\n**Classifying.**\n\nWe may also add to our assumptions based on the problem description noted earlier.\n\n1. Women (Sex=female) were more likely to have survived.\n2. Children (Age<?) were more likely to have survived. \n3. The upper-class passengers (Pclass=1) were more likely to have survived."},{"metadata":{"_cell_guid":"6db63a30-1d86-266e-2799-dded03c45816","_uuid":"4b9a8cd75d1048fc22566f902ed6337a28a23049"},"cell_type":"markdown","source":"## Analyze by pivoting features\n\nTo confirm some of our observations and assumptions, we can quickly analyze our feature correlations by pivoting features against each other. We can only do so at this stage for features which do not have any empty values. It also makes sense doing so only for features which are categorical (Sex), ordinal (Pclass) or discrete (SibSp, Parch) type.\n\n- **Pclass** We observe significant correlation (>0.5) among Pclass=1 and Survived (classifying #3). We decide to include this feature in our model.\n- **Sex** We confirm the observation during problem definition that Sex=female had very high survival rate at 74% (classifying #1).\n- **SibSp and Parch** These features have zero correlation for certain values. It may be best to derive a feature or a set of features from these individual features (creating #1)."},{"metadata":{"_cell_guid":"0964832a-a4be-2d6f-a89e-63526389cee9","_uuid":"c79ee4464a3bda5842ed7763e03da587a747b3d6","trusted":true},"cell_type":"code","source":"train_df[['Pclass', 'Survived']].groupby(['Pclass'], as_index=False).mean().sort_values(by='Survived', ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"68908ba6-bfe9-5b31-cfde-6987fc0fbe9a","_uuid":"519d35fa3ea920b27ac87c06bffa1258ac96fe3a","trusted":true},"cell_type":"code","source":"train_df[[\"Sex\", \"Survived\"]].groupby(['Sex'], as_index=False).mean().sort_values(by='Survived', ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"01c06927-c5a6-342a-5aa8-2e486ec3fd7c","_uuid":"6926b5f26f11aedcd2c351428e3614588f3dcab6","trusted":true},"cell_type":"code","source":"train_df[[\"SibSp\", \"Survived\"]].groupby(['SibSp'], as_index=False).mean().sort_values(by='Survived', ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"e686f98b-a8c9-68f8-36a4-d4598638bbd5","_uuid":"3d11ff5bb02569da009468be9e437f9fe7528eac","trusted":true},"cell_type":"code","source":"train_df[[\"Parch\", \"Survived\"]].groupby(['Parch'], as_index=False).mean().sort_values(by='Survived', ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"0d43550e-9eff-3859-3568-8856570eff76","_uuid":"ae8d91dc4ae624f82a3b96eb7dbd44910d33afb5"},"cell_type":"markdown","source":"## Analyze by visualizing data\n\nNow we can continue confirming some of our assumptions using visualizations for analyzing the data.\n\n### Correlating numerical features\n\nLet us start by understanding correlations between numerical features and our solution goal (Survived).\n\nA histogram chart is useful for analyzing continous numerical variables like Age where banding or ranges will help identify useful patterns. The histogram can indicate distribution of samples using automatically defined bins or equally ranged bands. This helps us answer questions relating to specific bands (Did infants have better survival rate?)\n\nNote that x-axis in historgram visualizations represents the count of samples or passengers.\n\n**Observations.**\n\n- Infants (Age <=4) had high survival rate.\n- Oldest passengers (Age = 80) survived.\n- Large number of 15-25 year olds did not survive.\n- Most passengers are in 15-35 age range.\n\n**Decisions.**\n\nThis simple analysis confirms our assumptions as decisions for subsequent workflow stages.\n\n- We should consider Age (our assumption classifying #2) in our model training.\n- Complete the Age feature for null values (completing #1).\n- We should band age groups (creating #3)."},{"metadata":{"_cell_guid":"50294eac-263a-af78-cb7e-3778eb9ad41f","_uuid":"f5560a1dc8224b597c75b3359929ae537e91115d","trusted":true},"cell_type":"code","source":"g = sns.FacetGrid(train_df, col='Survived')\ng.map(plt.hist, 'Age', bins=20)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"87096158-4017-9213-7225-a19aea67a800","_uuid":"cf8c6a97856154c3452f746eeaa401ea0474603f"},"cell_type":"markdown","source":"### Correlating numerical and ordinal features\n\nWe can combine multiple features for identifying correlations using a single plot. This can be done with numerical and categorical features which have numeric values.\n\n**Observations.**\n\n- Pclass=3 had most passengers, however most did not survive. Confirms our classifying assumption #2.\n- Infant passengers in Pclass=2 and Pclass=3 mostly survived. Further qualifies our classifying assumption #2.\n- Most passengers in Pclass=1 survived. Confirms our classifying assumption #3.\n- Pclass varies in terms of Age distribution of passengers.\n\n**Decisions.**\n\n- Consider Pclass for model training."},{"metadata":{"_cell_guid":"916fdc6b-0190-9267-1ea9-907a3d87330d","_uuid":"3056eb9308329c83b1bafc178ea0c4ac77b65d2e","trusted":true},"cell_type":"code","source":"# grid = sns.FacetGrid(train_df, col='Pclass', hue='Survived')\ngrid = sns.FacetGrid(train_df, col='Survived', row='Pclass', size=2.2, aspect=1.6)\ngrid.map(plt.hist, 'Age', alpha=.5, bins=20)\ngrid.add_legend();","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"36f5a7c0-c55c-f76f-fdf8-945a32a68cb0","_uuid":"3fa1033dd0ae08675123077319bd7c01f09babc6"},"cell_type":"markdown","source":"### Correlating categorical features\n\nNow we can correlate categorical features with our solution goal.\n\n**Observations.**\n\n- Female passengers had much better survival rate than males. Confirms classifying (#1).\n- Exception in Embarked=C where males had higher survival rate. This could be a correlation between Pclass and Embarked and in turn Pclass and Survived, not necessarily direct correlation between Embarked and Survived.\n- Males had better survival rate in Pclass=3 when compared with Pclass=2 for C and Q ports. Completing (#2).\n- Ports of embarkation have varying survival rates for Pclass=3 and among male passengers. Correlating (#1).\n\n**Decisions.**\n\n- Add Sex feature to model training.\n- Complete and add Embarked feature to model training."},{"metadata":{"_cell_guid":"db57aabd-0e26-9ff9-9ebd-56d401cdf6e8","_uuid":"0b037679a1a6194331ce4ff80cbc2774a5ae21c5","trusted":true},"cell_type":"code","source":"# grid = sns.FacetGrid(train_df, col='Embarked')\ngrid = sns.FacetGrid(train_df, row='Embarked', size=2.2, aspect=1.6)\ngrid.map(sns.pointplot, 'Pclass', 'Survived', 'Sex', palette='deep')\ngrid.add_legend()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"6b3f73f4-4600-c1ce-34e0-bd7d9eeb074a","_uuid":"d0752c49d60c97ec22c6f0563b690c84e9c85947"},"cell_type":"markdown","source":"### Correlating categorical and numerical features\n\nWe may also want to correlate categorical features (with non-numeric values) and numeric features. We can consider correlating Embarked (Categorical non-numeric), Sex (Categorical non-numeric), Fare (Numeric continuous), with Survived (Categorical numeric).\n\n**Observations.**\n\n- Higher fare paying passengers had better survival. Confirms our assumption for creating (#4) fare ranges.\n- Port of embarkation correlates with survival rates. Confirms correlating (#1) and completing (#2).\n\n**Decisions.**\n\n- Consider banding Fare feature."},{"metadata":{"_cell_guid":"a21f66ac-c30d-f429-cc64-1da5460d16a9","_uuid":"77fc02a9b90b6ace4a6fb9ce1cf0978fd8171b7e","trusted":true},"cell_type":"code","source":"# grid = sns.FacetGrid(train_df, col='Embarked', hue='Survived', palette={0: 'k', 1: 'w'})\ngrid = sns.FacetGrid(train_df, row='Embarked', col='Survived', size=2.2, aspect=1.6)\ngrid.map(sns.barplot, 'Sex', 'Fare', alpha=.5, ci=None)\ngrid.add_legend()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"cfac6291-33cc-506e-e548-6cad9408623d","_uuid":"f6e710a6e3b2e94cee474d5ecf7b4c2a2a8a785f"},"cell_type":"markdown","source":"## Wrangle data\n\nWe have collected several assumptions and decisions regarding our datasets and solution requirements. So far we did not have to change a single feature or value to arrive at these. Let us now execute our decisions and assumptions for correcting, creating, and completing goals.\n\n### Correcting by dropping features\n\nThis is a good starting goal to execute. By dropping features we are dealing with fewer data points. Speeds up our notebook and eases the analysis.\n\nBased on our assumptions and decisions we want to drop the Cabin (correcting #2) and Ticket (correcting #1) features.\n\nNote that where applicable we perform operations on both training and testing datasets together to stay consistent."},{"metadata":{"_cell_guid":"da057efe-88f0-bf49-917b-bb2fec418ed9","_uuid":"bd3b0e2b0f565a2bbfad72950c506bddb4f9a57e","trusted":true},"cell_type":"code","source":"print(\"Before\", train_df.shape, test_df.shape, combine[0].shape, combine[1].shape)\n\ntrain_df = train_df.drop(['Ticket', 'Cabin'], axis=1)\ntest_df = test_df.drop(['Ticket', 'Cabin'], axis=1)\ncombine = [train_df, test_df]\n\n\"After\", train_df.shape, test_df.shape, combine[0].shape, combine[1].shape","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"6b3a1216-64b6-7fe2-50bc-e89cc964a41c","_uuid":"ef94d905ec97575a6d0293eb16190f1901a15598"},"cell_type":"markdown","source":"### Creating new feature extracting from existing\n\nWe want to analyze if Name feature can be engineered to extract titles and test correlation between titles and survival, before dropping Name and PassengerId features.\n\nIn the following code we extract Title feature using regular expressions. The RegEx pattern `(\\w+\\.)` matches the first word which ends with a dot character within Name feature. The `expand=False` flag returns a DataFrame.\n\n**Observations.**\n\nWhen we plot Title, Age, and Survived, we note the following observations.\n\n- Most titles band Age groups accurately. For example: Master title has Age mean of 5 years.\n- Survival among Title Age bands varies slightly.\n- Certain titles mostly survived (Mme, Lady, Sir) or did not (Don, Rev, Jonkheer).\n\n**Decision.**\n\n- We decide to retain the new Title feature for model training."},{"metadata":{"_cell_guid":"df7f0cd4-992c-4a79-fb19-bf6f0c024d4b","_uuid":"ab7ac0b0c936cef53d7d94a1f2bc96fce84d1d8a","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['Title'] = dataset.Name.str.extract(' ([A-Za-z]+)\\.', expand=False)\n\npd.crosstab(train_df['Title'], train_df['Sex'])","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"908c08a6-3395-19a5-0cd7-13341054012a","_uuid":"07c1f61babcfc28fc43ab0cb57bebb66e8418d02"},"cell_type":"markdown","source":"We can replace many titles with a more common name or classify them as `Rare`."},{"metadata":{"_cell_guid":"553f56d7-002a-ee63-21a4-c0efad10cfe9","_uuid":"3320045352e88e063906fce5e12631e4ea671dde","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['Title'] = dataset['Title'].replace(['Lady', 'Countess','Capt', 'Col',\\\n \t'Don', 'Dr', 'Major', 'Rev', 'Sir', 'Jonkheer', 'Dona'], 'Rare')\n\n    dataset['Title'] = dataset['Title'].replace('Mlle', 'Miss')\n    dataset['Title'] = dataset['Title'].replace('Ms', 'Miss')\n    dataset['Title'] = dataset['Title'].replace('Mme', 'Mrs')\n    \ntrain_df[['Title', 'Survived']].groupby(['Title'], as_index=False).mean()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"6d46be9a-812a-f334-73b9-56ed912c9eca","_uuid":"cf17269ba24e49a1f8e02fcb84af8e32bf0216de"},"cell_type":"markdown","source":"We can convert the categorical titles to ordinal."},{"metadata":{"_cell_guid":"67444ebc-4d11-bac1-74a6-059133b6e2e8","_uuid":"7a0c787c1ffe23c11d95d22f9c1ef711287e873b","trusted":true},"cell_type":"code","source":"title_mapping = {\"Mr\": 1, \"Miss\": 2, \"Mrs\": 3, \"Master\": 4, \"Rare\": 5}\nfor dataset in combine:\n    dataset['Title'] = dataset['Title'].map(title_mapping)\n    dataset['Title'] = dataset['Title'].fillna(0)\n\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"f27bb974-a3d7-07a1-f7e4-876f6da87e62","_uuid":"ca159a0180491fe3b17886a93effb0b8bd09b812"},"cell_type":"markdown","source":"Now we can safely drop the Name feature from training and testing datasets. We also do not need the PassengerId feature in the training dataset."},{"metadata":{"_cell_guid":"9d61dded-5ff0-5018-7580-aecb4ea17506","_uuid":"019a0e24beb9bea988c7278507ffefdc49b77eeb","trusted":true},"cell_type":"code","source":"train_df = train_df.drop(['Name', 'PassengerId'], axis=1)\ntest_df = test_df.drop(['Name'], axis=1)\ncombine = [train_df, test_df]\ntrain_df.shape, test_df.shape","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"2c8e84bb-196d-bd4a-4df9-f5213561b5d3","_uuid":"0e2af6d0a5943b9c4473978ecb7cf80cd0a3cf47"},"cell_type":"markdown","source":"### Converting a categorical feature\n\nNow we can convert features which contain strings to numerical values. This is required by most model algorithms. Doing so will also help us in achieving the feature completing goal.\n\nLet us start by converting Sex feature to a new feature called Gender where female=1 and male=0."},{"metadata":{"_cell_guid":"c20c1df2-157c-e5a0-3e24-15a828095c96","_uuid":"0ecd25dd2d9b0eabf97f1665469b565ddb575bb3","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['Sex'] = dataset['Sex'].map( {'female': 1, 'male': 0} ).astype(int)\n\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"d72cb29e-5034-1597-b459-83a9640d3d3a","_uuid":"ab8112001c3dbf313277ad0e17ae9166cab2de1b"},"cell_type":"markdown","source":"### Completing a numerical continuous feature\n\nNow we should start estimating and completing features with missing or null values. We will first do this for the Age feature.\n\nWe can consider three methods to complete a numerical continuous feature.\n\n1. A simple way is to generate random numbers between mean and [standard deviation](https://en.wikipedia.org/wiki/Standard_deviation).\n\n2. More accurate way of guessing missing values is to use other correlated features. In our case we note correlation among Age, Gender, and Pclass. Guess Age values using [median](https://en.wikipedia.org/wiki/Median) values for Age across sets of Pclass and Gender feature combinations. So, median Age for Pclass=1 and Gender=0, Pclass=1 and Gender=1, and so on...\n\n3. Combine methods 1 and 2. So instead of guessing age values based on median, use random numbers between mean and standard deviation, based on sets of Pclass and Gender combinations.\n\nMethod 1 and 3 will introduce random noise into our models. The results from multiple executions might vary. We will prefer method 2."},{"metadata":{"_cell_guid":"c311c43d-6554-3b52-8ef8-533ca08b2f68","_uuid":"2182a561929886d93d718e5d3f5ddb9ee0366259","trusted":true},"cell_type":"code","source":"# grid = sns.FacetGrid(train_df, col='Pclass', hue='Gender')\ngrid = sns.FacetGrid(train_df, row='Pclass', col='Sex', size=2.2, aspect=1.6)\ngrid.map(plt.hist, 'Age', alpha=.5, bins=20)\ngrid.add_legend()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"a4f166f9-f5f9-1819-66c3-d89dd5b0d8ff","_uuid":"0fffe1bd419524cad994fe205c5bea5c05c15400"},"cell_type":"markdown","source":"Let us start by preparing an empty array to contain guessed Age values based on Pclass x Gender combinations."},{"metadata":{"_cell_guid":"9299523c-dcf1-fb00-e52f-e2fb860a3920","_uuid":"cb73bec742e006be8ea45c8c717b13a754418199","trusted":true},"cell_type":"code","source":"guess_ages = np.zeros((2,3))\nguess_ages","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"ec9fed37-16b1-5518-4fa8-0a7f579dbc82","_uuid":"0d1f60d4fea5d1d0cb9cbd06776f5433a0378be2"},"cell_type":"markdown","source":"Now we iterate over Sex (0 or 1) and Pclass (1, 2, 3) to calculate guessed values of Age for the six combinations."},{"metadata":{"_cell_guid":"a4015dfa-a0ab-65bc-0cbe-efecf1eb2569","_uuid":"64cde468817d63eccdea0803c76741d7d370bb7b","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    for i in range(0, 2):\n        for j in range(0, 3):\n            guess_df = dataset[(dataset['Sex'] == i) & \\\n                                  (dataset['Pclass'] == j+1)]['Age'].dropna()\n\n            # age_mean = guess_df.mean()\n            # age_std = guess_df.std()\n            # age_guess = rnd.uniform(age_mean - age_std, age_mean + age_std)\n\n            age_guess = guess_df.median()\n\n            # Convert random age float to nearest .5 age\n            guess_ages[i,j] = int( age_guess/0.5 + 0.5 ) * 0.5\n            \n    for i in range(0, 2):\n        for j in range(0, 3):\n            dataset.loc[ (dataset.Age.isnull()) & (dataset.Sex == i) & (dataset.Pclass == j+1),\\\n                    'Age'] = guess_ages[i,j]\n\n    dataset['Age'] = dataset['Age'].astype(int)\n\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"dbe0a8bf-40bc-c581-e10e-76f07b3b71d4","_uuid":"fc22e5bdfe3cdbbf136823036dfa7dde8529ced2"},"cell_type":"markdown","source":"Let us create Age bands and determine correlations with Survived."},{"metadata":{"_cell_guid":"725d1c84-6323-9d70-5812-baf9994d3aa1","_uuid":"75c192a1dccd6a785cdc723df70ba2d7638495ad","trusted":true},"cell_type":"code","source":"train_df['AgeBand'] = pd.cut(train_df['Age'], 5)\ntrain_df[['AgeBand', 'Survived']].groupby(['AgeBand'], as_index=False).mean().sort_values(by='AgeBand', ascending=True)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"ba4be3a0-e524-9c57-fbec-c8ecc5cde5c6","_uuid":"39d25d400f99813c55f6c70c32bf1120db155535"},"cell_type":"markdown","source":"Let us replace Age with ordinals based on these bands."},{"metadata":{"_cell_guid":"797b986d-2c45-a9ee-e5b5-088de817c8b2","_uuid":"4f2f177f954d6977e0062857da1af1a50d5da64f","trusted":true},"cell_type":"code","source":"for dataset in combine:    \n    dataset.loc[ dataset['Age'] <= 16, 'Age'] = 0\n    dataset.loc[(dataset['Age'] > 16) & (dataset['Age'] <= 32), 'Age'] = 1\n    dataset.loc[(dataset['Age'] > 32) & (dataset['Age'] <= 48), 'Age'] = 2\n    dataset.loc[(dataset['Age'] > 48) & (dataset['Age'] <= 64), 'Age'] = 3\n    dataset.loc[ dataset['Age'] > 64, 'Age']\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"004568b6-dd9a-ff89-43d5-13d4e9370b1d","_uuid":"e9d4cf5d4371a9033574b8d079d70b52211122fb"},"cell_type":"markdown","source":"We can not remove the AgeBand feature."},{"metadata":{"_cell_guid":"875e55d4-51b0-5061-b72c-8a23946133a3","_uuid":"fff0e6c13536c2373326d3d103de6a78183b1bb0","trusted":true},"cell_type":"code","source":"train_df = train_df.drop(['AgeBand'], axis=1)\ncombine = [train_df, test_df]\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"1c237b76-d7ac-098f-0156-480a838a64a9","_uuid":"b785d5c5c389ad3afc23cbcf1457c980ac4e74ef"},"cell_type":"markdown","source":"### Create new feature combining existing features\n\nWe can create a new feature for FamilySize which combines Parch and SibSp. This will enable us to drop Parch and SibSp from our datasets."},{"metadata":{"_cell_guid":"7e6c04ed-cfaa-3139-4378-574fd095d6ba","_uuid":"bd34f2179ccef65fbbe216e0ba3c1a4f221df678","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['FamilySize'] = dataset['SibSp'] + dataset['Parch'] + 1\n\ntrain_df[['FamilySize', 'Survived']].groupby(['FamilySize'], as_index=False).mean().sort_values(by='Survived', ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"842188e6-acf8-2476-ccec-9e3451e4fa86","_uuid":"ee3cc2b3a2cfef186e6895a5519a61f98fd0ceb0"},"cell_type":"markdown","source":"We can create another feature called IsAlone."},{"metadata":{"_cell_guid":"5c778c69-a9ae-1b6b-44fe-a0898d07be7a","_uuid":"571f5bedf380d5f68e11377345dc4b28d4eab916","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['IsAlone'] = 0\n    dataset.loc[dataset['FamilySize'] == 1, 'IsAlone'] = 1\n\ntrain_df[['IsAlone', 'Survived']].groupby(['IsAlone'], as_index=False).mean()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"e6b87c09-e7b2-f098-5b04-4360080d26bc","_uuid":"cd284a6a1f6f233b430c5d40e5faea4f91a092bb"},"cell_type":"markdown","source":"Let us drop Parch, SibSp, and FamilySize features in favor of IsAlone."},{"metadata":{"_cell_guid":"74ee56a6-7357-f3bc-b605-6c41f8aa6566","_uuid":"3060b88fca180c581581f3cde6412725e3136ac4","trusted":true},"cell_type":"code","source":"train_df = train_df.drop(['Parch', 'SibSp', 'FamilySize'], axis=1)\ntest_df = test_df.drop(['Parch', 'SibSp', 'FamilySize'], axis=1)\ncombine = [train_df, test_df]\n\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"f890b730-b1fe-919e-fb07-352fbd7edd44","_uuid":"92a531f666e97294655228ade67c956d058370bd"},"cell_type":"markdown","source":"We can also create an artificial feature combining Pclass and Age."},{"metadata":{"_cell_guid":"305402aa-1ea1-c245-c367-056eef8fe453","_uuid":"178e142bc97dca6fb4ab9b50ea3070c572608f77","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['Age*Class'] = dataset.Age * dataset.Pclass\n\ntrain_df.loc[:, ['Age*Class', 'Age', 'Pclass']].head(10)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"13292c1b-020d-d9aa-525c-941331bb996a","_uuid":"d9912e0c6a1563c0c564c0e82e670afc21b7e074"},"cell_type":"markdown","source":"### Completing a categorical feature\n\nEmbarked feature takes S, Q, C values based on port of embarkation. Our training dataset has two missing values. We simply fill these with the most common occurance."},{"metadata":{"_cell_guid":"bf351113-9b7f-ef56-7211-e8dd00665b18","_uuid":"f13407dcbd5cf5263a6a7f5730b0bae1fdf30df0","trusted":true},"cell_type":"code","source":"freq_port = train_df.Embarked.dropna().mode()[0]\nfreq_port","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"51c21fcc-f066-cd80-18c8-3d140be6cbae","_uuid":"bf4a55028148881f244b247c5942bd8a00e860e1","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['Embarked'] = dataset['Embarked'].fillna(freq_port)\n    \ntrain_df[['Embarked', 'Survived']].groupby(['Embarked'], as_index=False).mean().sort_values(by='Survived', ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"f6acf7b2-0db3-e583-de50-7e14b495de34","_uuid":"d0f36769354c2941d6f4c5cb8b87d406033d23a3"},"cell_type":"markdown","source":"### Converting categorical feature to numeric\n\nWe can now convert the EmbarkedFill feature by creating a new numeric Port feature."},{"metadata":{"_cell_guid":"89a91d76-2cc0-9bbb-c5c5-3c9ecae33c66","_uuid":"170fe060610dc3d02fdbf795e6a0f73d0f35f9f2","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['Embarked'] = dataset['Embarked'].map( {'S': 0, 'C': 1, 'Q': 2} ).astype(int)\n\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"e3dfc817-e1c1-a274-a111-62c1c814cecf","_uuid":"41f4b8eb901765461e88a87a1ff2b645b573134c"},"cell_type":"markdown","source":"### Quick completing and converting a numeric feature\n\nWe can now complete the Fare feature for single missing value in test dataset using mode to get the value that occurs most frequently for this feature. We do this in a single line of code.\n\nNote that we are not creating an intermediate new feature or doing any further analysis for correlation to guess missing feature as we are replacing only a single value. The completion goal achieves desired requirement for model algorithm to operate on non-null values.\n\nWe may also want round off the fare to two decimals as it represents currency."},{"metadata":{"_cell_guid":"3600cb86-cf5f-d87b-1b33-638dc8db1564","_uuid":"cff24e42e82d634a1a5cfa4645b4f4265d9562ea","trusted":true},"cell_type":"code","source":"test_df['Fare'].fillna(test_df['Fare'].dropna().median(), inplace=True)\ntest_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"4b816bc7-d1fb-c02b-ed1d-ee34b819497d","_uuid":"361308a98bf2902d9f7c2796ab34525cf65fea41"},"cell_type":"markdown","source":"We can not create FareBand."},{"metadata":{"_cell_guid":"0e9018b1-ced5-9999-8ce1-258a0952cbf2","_uuid":"9efe4d060470730e4beea62d9c2d3c3ad1a5cffa","trusted":true},"cell_type":"code","source":"train_df['FareBand'] = pd.qcut(train_df['Fare'], 4)\ntrain_df[['FareBand', 'Survived']].groupby(['FareBand'], as_index=False).mean().sort_values(by='FareBand', ascending=True)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"d65901a5-3684-6869-e904-5f1a7cce8a6d","_uuid":"09845a7d7a142d231f5d38b0e95554b48d26c30b"},"cell_type":"markdown","source":"Convert the Fare feature to ordinal values based on the FareBand."},{"metadata":{"_cell_guid":"385f217a-4e00-76dc-1570-1de4eec0c29c","_uuid":"9f18b3a976bd902e184e60113ff31de9a1671a3b","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset.loc[ dataset['Fare'] <= 7.91, 'Fare'] = 0\n    dataset.loc[(dataset['Fare'] > 7.91) & (dataset['Fare'] <= 14.454), 'Fare'] = 1\n    dataset.loc[(dataset['Fare'] > 14.454) & (dataset['Fare'] <= 31), 'Fare']   = 2\n    dataset.loc[ dataset['Fare'] > 31, 'Fare'] = 3\n    dataset['Fare'] = dataset['Fare'].astype(int)\n\ntrain_df = train_df.drop(['FareBand'], axis=1)\ncombine = [train_df, test_df]\n    \ntrain_df.head(10)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"27272bb9-3c64-4f9a-4a3b-54f02e1c8289","_uuid":"7b26c822e0daa19e14ea14bb6e7ea46b0607c5ef"},"cell_type":"markdown","source":"And the test dataset."},{"metadata":{"_cell_guid":"d2334d33-4fe5-964d-beac-6aa620066e15","_uuid":"55b55c74a16db56a63dfd4a2db5ec4df1923b30f","trusted":true},"cell_type":"code","source":"test_df.head(10)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"752b1791-2bcd-e05e-f1d3-af59bc1e8e92","_uuid":"1bc5d639688e4ff741127a27b8246ef3802bc750"},"cell_type":"markdown","source":"# Write the preprocessed data frames to csv"},{"metadata":{"_cell_guid":"21ccca92-2db1-de4d-5dc0-c84dbb1b7407","_uuid":"965fe2b5b3edf78b4180df8c276d63f1c08f4ca3","trusted":true},"cell_type":"code","source":"train_df.to_csv('train_df.csv',header=True, index=False)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"086c84e3-8aa6-ae6f-4000-af09583345c8","_uuid":"03a13e7cc5eeca820f50d68a102962800b305c0f","trusted":true},"cell_type":"code","source":"test_df.to_csv('test_df.csv',header=True, index=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a2eaf46becc448df327f668ea0c6b822b68200ec"},"cell_type":"markdown","source":"# Feature importance"},{"metadata":{"trusted":true,"_uuid":"a521f00b507eb524399523a4fefac6684bc24ba7"},"cell_type":"code","source":"def cohen_effect_size(X, y):\n    \"\"\"Calculates the Cohen effect size of each feature.\n    \n        Parameters\n        ----------\n        X : {array-like, sparse matrix}, shape = [n_samples, n_features]\n            Training vector, where n_samples in the number of samples and\n            n_features is the number of features.\n        y : array-like, shape = [n_samples]\n            Target vector relative to X\n        Returns\n        -------\n        cohen_effect_size : array, shape = [n_features,]\n            The set of Cohen effect values.\n    \"\"\"\n    group1 = X[y==0]\n    group2 = X[y==1]\n    diff = group1.mean() - group2.mean()\n    var1, var2 = group1.var(), group2.var()\n    n1 = group1.shape[0]\n    n2 = group2.shape[0]\n    pooled_var = (n1 * var1 + n2 * var2) / (n1 + n2)\n    d = diff / np.sqrt(pooled_var)\n    return d","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fae41abdfa44f9b2065ac3d664a0fe743398045c"},"cell_type":"code","source":"X_train = train_df.drop(\"Survived\", axis=1)\nY_train = train_df[\"Survived\"]\nX_test  = test_df.drop(\"PassengerId\", axis=1).copy()\nX_train.shape, Y_train.shape, X_test.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"98ce264c61ef86efe9e14fb34eb09d29c9d2fe6a"},"cell_type":"code","source":"effect_sizes = cohen_effect_size(X_train, Y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"314047291dd1ef8e53ab5eb88759c5e9f6c9ff0e"},"cell_type":"code","source":"effect_sizes_s = effect_sizes.reindex(effect_sizes.abs().sort_values(ascending=False).index)\neffect_sizes_s.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f434800e71e7e1ba6ab7e0d19c75604b22515c88"},"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(6, 8))\neffect_sizes_s[::-1].plot.barh(ax=ax);","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2d89cdba380d446e0056270535f6b779f2964245"},"cell_type":"code","source":"from sklearn.utils import shuffle\n\ndef p_value_effect(X, y, nr_iters=1000):\n    actual = cohen_effect_size(X, y)\n    results = np.zeros(actual.shape[0])\n    for i in range(nr_iters):\n        target_shuffled = shuffle(y.values)\n        results = results + (cohen_effect_size(X, target_shuffled).abs() >= actual.abs())\n    p_values = results/nr_iters\n    return pd.DataFrame({'cohen_effect_size':actual, 'p_value':p_values}, index=actual.index)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e9eba36b57fdc8c6a4f331a9f1f618175d5e2ca8"},"cell_type":"code","source":"train_df['Survived'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"786eafc60467a0eb34a134f10a473adb519f9343"},"cell_type":"code","source":"    X = train_df.drop('Survived', axis=1)\n    y = train_df['Survived']","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"0acf54f9-6cf5-24b5-72d9-29b30052823a","_uuid":"500dd31a8fb1d936ea9214370873befd35e2bb8a","trusted":true},"cell_type":"code","source":"%time df_ces_p = p_value_effect(X, y, 10000)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"dcf67c66a863a4dacb66da2838f9384b7ec4bb11"},"cell_type":"code","source":"df_ces_p","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9b20763f4bb7072be15c0c93a70d80d23ef6db9d"},"cell_type":"markdown","source":"From the results we see that Age is possibly not significant because it's p-value is close to 0.05. On the other hand, Age*Class featyre is definitly significant.  We must not be surprised that the effect size of Age*Class is 0.20 and is rather low. Hence using univariate effect size for feature pruning might drop features that have significant interactions with other features. Since the effect size is small, it might not influence the prediction result a lot."},{"metadata":{"trusted":true,"_uuid":"42a222bae9085c2e9e0d4026302701e32a69cf2b"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"_change_revision":0,"_is_fork":false,"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}