{"cells":[{"metadata":{"_uuid":"2b5ec99544519a8dcbd9b92852c0d6914ca2ee44"},"cell_type":"markdown","source":"### First ML Notebook || TITANIC SURVIVAL"},{"metadata":{"_uuid":"f1ce8a0682a8af6a689740abb2b5ff885de2a7de"},"cell_type":"markdown","source":"**Berroug Med Amine**"},{"metadata":{"trusted":true,"_uuid":"754ddc51043e40af9fdc29f47b9bbba3825a25df"},"cell_type":"code","source":"from IPython.display import Image\nImage(url= \"https://static1.squarespace.com/static/5006453fe4b09ef2252ba068/5095eabce4b06cb305058603/5095eabce4b02d37bef4c24c/1352002236895/100_anniversary_titanic_sinking_by_esai8mellows-d4xbme8.jpg?format=1000w\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2a9791096a022f2e00ce2ae2d644309f41d3b6c3"},"cell_type":"markdown","source":"Import all libraries required\n- pandas is used for data manipulation \n+ numpy is used for mathematacial functions\n+ matplotlib and seaborn are for Data Visualisation"},{"metadata":{"_uuid":"ef913c1c0fd72eb763167bd09f9199b70658aad1"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\">\n<b>1) Importing libraries</b> \n</div>"},{"metadata":{"trusted":true,"_uuid":"9961eb32638c1e395e166e5da42d21930ba29dd2"},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom pandas.plotting  import scatter_matrix\nimport matplotlib.pyplot as plt\nfrom sklearn import model_selection \nfrom sklearn.metrics import classification_report \nfrom sklearn.metrics import confusion_matrix \nfrom sklearn.metrics import accuracy_score \nfrom sklearn.tree import DecisionTreeClassifier \nfrom sklearn.neighbors import KNeighborsClassifier\nfrom collections import Counter\nimport seaborn as sns\nimport warnings#ignore alertes\nwarnings.filterwarnings('ignore')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"aa52fbb91e351f94406cd768a70c7c456fb4ee38"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\">\n<b>2) Importing data Set </b> \n</div>"},{"metadata":{"trusted":true,"_uuid":"446a85f3008ec1ea566c9ad2e359fba55b772fc0"},"cell_type":"code","source":"\ntrain_df = pd.read_csv(\"../input/train.csv\")\n\ntest_df = pd.read_csv(\"../input/test.csv\")\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"66d7b870cc57918c7c7ad99d188d8e3fcdba4554"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\">\n<b>3) Data Understanding</b> \n</div>"},{"metadata":{"trusted":true,"_uuid":"e997e5b670bf227ce5970ba2c9a589e8bbc8ab86"},"cell_type":"code","source":"\ntrain_df.head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"972477026c94d9cfdfc835777298e8844253d07c"},"cell_type":"code","source":"\ntest_df.head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"24f1fec476e541c91961230e4f294a30af15b62e"},"cell_type":"code","source":"\ntrain_df.describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d53c7ccb883239564b336ea070b8dea21a51f516"},"cell_type":"code","source":"train_df.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9b34433d802b6f35f26514ecf40d72d38418eec3"},"cell_type":"code","source":"test_df.info()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"561b0c35deb5cb295c1614dee0843a414b7a1bfc"},"cell_type":"markdown","source":"For the visualization of the data, it is important to note the type of variables:\n\n**Nominal:** Survived, Sex, Embarked, SibSp, Parch<br>\n**Ordinal:** Pclass<br>\n**Numerical:** Age, Fare<br>"},{"metadata":{"_uuid":"c53bdb2392dee2ccfee9d8aabc6113d0b0d75cca"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\">\n<b>4) Data Visualisation</b> \n</div>"},{"metadata":{"_uuid":"35d16f37d10579af2619269ed6f3b1791fde27a9"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-warning\">\n<b></b> \nwe will display some attributes according to the attribute label = 'Survived'</div>"},{"metadata":{"trusted":true,"_uuid":"9d2a3f53815b56d6f3f9afd71a967584c3ae57df"},"cell_type":"code","source":"#Detection of possible correlations between class and other attributes\ng = sns.heatmap(train_df[[\"Survived\",\"SibSp\",\"Parch\",\"Age\",\"Fare\"]].corr(),annot=True, fmt = \".2f\", cmap = \"coolwarm\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c4f9ba48b1f7cdf49ef26f8a3ad05f5dd7ddb8ff"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-warning\">\n <b>Two interesting correlations are:</b>\npositive (0.41) for <b>SibSp and Parch</b>\nnegative (-0.31) pour <b>SibSp and Age</b>\n</div>"},{"metadata":{"trusted":true,"_uuid":"8ddbfd7e1f4a0ce3399c5ec75a523b3bc713200f"},"cell_type":"code","source":"#Defined a function that represents the Sex, Pclass in barplot\ndef bar_chart(feature):\n    survécu = train_df[train_df['Survived']==1][feature].value_counts()\n    mort = train_df[train_df['Survived']==0][feature].value_counts()\n    df = pd.DataFrame([survécu,mort])\n    df.index = ['Survived','dead']\n    df.plot(kind='bar',stacked=True, figsize=(10,5))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0da3c709a03cbb558d757aefcad580df651e4eb5"},"cell_type":"markdown","source":"**Sexe**"},{"metadata":{"trusted":true,"_uuid":"c729399c7519ad1b8a322a7e7111b8306b42564f"},"cell_type":"code","source":"bar_chart('Sex')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1a5c109ffd3a410b4f21a6592ec0ee77f1c76807"},"cell_type":"markdown","source":"\n<div class=\"alert alert-block alert-warning\">\n<b>Note:</b> \n    The chart confirms that <b> women are more likely </b> to survive than <b> men </b> || women and children first</div>"},{"metadata":{"trusted":true,"_uuid":"510e30548f54b600c203f2a48d73c01d40bde38f"},"cell_type":"code","source":"#Explanation\ntrain_df[[\"Sex\",\"Survived\"]].groupby('Sex').mean()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5370795145cfc3bb06935e374415c120000e41bf"},"cell_type":"code","source":"#display data by seaborn barplot\ng = sns.barplot(x=\"Sex\",y=\"Survived\",data=train_df)\ng = g.set_ylabel(\"Survival Probability\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"eb8b71e851843646c6ba8a1de08ff30f9d2e13fe"},"cell_type":"markdown","source":"**Pclass**"},{"metadata":{"trusted":true,"_uuid":"a725c7088620357b6ddcf0e23dadd5f7b68a4bbd"},"cell_type":"code","source":"bar_chart('Pclass')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"63acea188743ab996320253eb0af7744e81b69ff"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-warning\">\n<b>Note:</b> \n  The graph confirms that 1st class people are more likely to survive than other classes. The chart confirms that the 3rd class people are more likely to die than the others || the rich people first</div>\n"},{"metadata":{"trusted":true,"_uuid":"fb3c04a2e9a935b62dd6777fb838b85de25e9892"},"cell_type":"code","source":"\ng = sns.factorplot(x=\"Pclass\",y=\"Survived\",data=train_df,kind=\"bar\", size = 6 , palette = \"muted\")\ng.despine(left=True)\ng = g.set_ylabels(\"Survival Probability\")\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"509ea22ce9bfc2643e2e432afd55841903d00348"},"cell_type":"code","source":"#more details\ng = sns.factorplot(x=\"Pclass\", y=\"Survived\", hue=\"Sex\", data=train_df,size=6, kind=\"bar\", palette=\"muted\")\ng.despine(left=True)\ng = g.set_ylabels(\"probabilité de survie par class ticket et Sex\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5caccb8a656d40495472a60b2842fd621ebc808d"},"cell_type":"markdown","source":"**SibSp**"},{"metadata":{"trusted":true,"_uuid":"ef3bb551f5d8706f385e97b0d68932c3ce772f55"},"cell_type":"code","source":"bar_chart('SibSp')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1ed7f8ebbee52f89e98f5c781174ff97f55780b8"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-warning\">\n<b>Note:</b> \n   the chart confirms that a person having more than two borthers/sisters or spouse is more likely to survive.moreover the chart confirms that a person on board without borthers/sisters/spouse is more likely to die</div>"},{"metadata":{"trusted":true,"_uuid":"fa22e46fd0dc2d0acc0862144d6aec00a415af3e"},"cell_type":"code","source":"g  = sns.factorplot(x=\"SibSp\",y=\"Survived\",data=train_df,kind=\"bar\", size = 6 , palette = \"muted\")\ng.despine(left=True)\ng = g.set_ylabels(\"Survival Probability\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"dc76a3c9308740f967e821aba4cca46e866f1030"},"cell_type":"markdown","source":"**Parch**"},{"metadata":{"trusted":true,"_uuid":"c72144ee0e36db1c35714dfe1abfeb3acdf37113"},"cell_type":"code","source":"bar_chart('Parch')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ea7bf7159b1523d5fe1c7dd61d55e1098c53e2cc"},"cell_type":"code","source":"g  = sns.factorplot(x=\"Parch\",y=\"Survived\",data=train_df,kind=\"bar\", size = 6 , palette = \"muted\")\ng.despine(left=True)\ng = g.set_ylabels(\"Survival Probability\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2376c8f9a5fcca4b89bac5c43197248d36c6b7e7"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-warning\">\n<b>Note:</b> \nThe graph confirms that a person on board with more than 2 parents or children is more likely to survive\n </div>\n"},{"metadata":{"_uuid":"22ba5573cf2fb5610e90460d519a8bd033bc9a7d"},"cell_type":"markdown","source":"**Age**"},{"metadata":{"trusted":true,"_uuid":"1f4ec24b5ca47f5e627d1167175743fee2d4a440"},"cell_type":"code","source":"g = sns.FacetGrid(train_df, col='Survived')\ng = g.map(sns.distplot, \"Age\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"98a1f166cc84e6eefe54d94d70150b3832ee6bea"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-warning\">\n<b>Note:</b> \nWe can see that the age distributions are not the same in the surviving and non-surviving subpopulations. Indeed, there is a peak corresponding to the young passengers who survived. We also find that passengers between 60 and 80 years have survived less.\n\nThus, even if \"Age\" does not correlate with \"Survived\", we can see that there are age categories of passengers that are more or less likely to survive.\n\nIt seems that very young passengers (between 0-5 years old) are more likely to survive.</div>\n"},{"metadata":{"_uuid":"7fd2b4affebb7d84809609b4bcda486c1bf613a3"},"cell_type":"markdown","source":"**Embarked**"},{"metadata":{"trusted":true,"_uuid":"77b7aac2b35f37132c3d5d9c78c01a3a6eb7f938"},"cell_type":"code","source":"# C = Cherbourg(France),Q = Queenstown(New-Zelande),S = Southampton(England)\nbar_chart('Embarked')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"712ca47fee295cebc8511bd8723881f84ead0107"},"cell_type":"code","source":"g = sns.factorplot(x=\"Embarked\", y=\"Survived\",  data=train_df,size=6, kind=\"bar\", palette=\"muted\")\ng.despine(left=True)\ng = g.set_ylabels(\"Survival Probability\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ffa3ae2b8a35bfa4b9b8d764a531c729968f9db5"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-warning\">\n<b>Note:</b> <br>\nThe graph confirms that a person onboard C is slightly more likely to survive<br> \nThe graph confirms that a person embarked from Q,S is more likely to die <br></div>"},{"metadata":{"_uuid":"c328364c34e6a35709907669bbed6105f73d03c6"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-warning\">\n<b>Note 2:</b> <br>\nsince we already analyzed survival probability by Pclass (1st class more likely to survive),the next step is to see the passengers distribution by PClass<br> \n</div>"},{"metadata":{"trusted":true,"_uuid":"63dccbb122e6019607fa35856313d0e441ee004e"},"cell_type":"code","source":"g = sns.factorplot(\"Pclass\", col=\"Embarked\",  data=train_df,size=6, kind=\"count\", palette=\"muted\")\ng.despine(left=True)\ng = g.set_ylabels(\"Count\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"82f8798b02ffbe7deac93a0af7690fb053d5933b"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\">\n<b>Hypothesis result supported by data:</b> <br>\nIndeed, the third class is the most frequent for passengers coming from Southampton (S) and Queenstown (Q), while the passengers of C (Cherbourg-France) are mostly first class and have higher survival rate .\n\nAt this point, we can not explain why the first class has a higher survival rate. our hypothesis is that first-class passengers were given priority during the evacuation due to their influence. the graph below represents this hypothesis.</div>"},{"metadata":{"trusted":true,"_uuid":"2bcb25d213e18ccb6d2534b30c4d66c9a1fc1b44"},"cell_type":"code","source":"from IPython.display import Image\nImage(url= \"https://static1.squarespace.com/static/5006453fe4b09ef2252ba068/t/5090b249e4b047ba54dfd258/1351660113175/TItanic-Survival-Infographic.jpg?format=1500w\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b35f1caccf7a61c82846399a8a84c38e779ccfc4"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\">\n<b>5) Missing values</b></div>"},{"metadata":{"_uuid":"4cdf5fe427cc4d541940c27a088155437c72c365"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\">\n<b>a.Outliers Treatement :</b> \n</div>"},{"metadata":{"_uuid":"4f8416c1e33502687dd6ff73bb0cfa40aac8896c"},"cell_type":"markdown","source":"_there are several methods of detecting outliers:_<br>\n\n- Modified Z-score method\n-  Z-score Method\n+ IQR interquartille"},{"metadata":{"_uuid":"5287bb9a0453a04655cf06037169b103b2501b66"},"cell_type":"markdown","source":"_**IQR interquartille** ._\n"},{"metadata":{"trusted":true,"_uuid":"0d7ccbf09c4072e577e32edd3590f71ff642f60a"},"cell_type":"code","source":"from collections import Counter\ndef detect_outliers(df,n,features):\n\n    outlier_indices = []\n    \n\n    for col in features:\n        # 1st quartile (25%)\n        Q1 = np.percentile(df[col], 25)\n        # 3rd quartile (75%)\n        Q3 = np.percentile(df[col],75)\n        # Interquartile range (IQR)\n        IQR = Q3 - Q1\n        \n\n        outlier_step = 1.5 * IQR\n\n        outlier_list_col = df[(df[col] < Q1 - outlier_step) | (df[col] > Q3 + outlier_step )].index\n\n        outlier_indices.extend(outlier_list_col)\n\n    outlier_indices = Counter(outlier_indices)        \n    multiple_outliers = list( k for k, v in outlier_indices.items() if v > n )\n    \n    return multiple_outliers   \n\nOutliers_to_drop = detect_outliers(train_df,2,[\"Age\",\"SibSp\",\"Parch\",\"Fare\"])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e79296d92c2adb063655da62e97f91b8b307b8c2"},"cell_type":"code","source":"train_df.loc[Outliers_to_drop] ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"350eb4c372ddbd7d03b22b5a3c5064d0ca6a66d4"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\">\n<b>for this we have defined our own method of outliers detection</b> by defining the Max and Min for each attribute\n</div>"},{"metadata":{"trusted":true,"_uuid":"755c8cda7cad6c601e6c65ef2ec0fd6032bfbd52"},"cell_type":"code","source":"# I defined outliers as being above of 99% percentile here\n# get lists of people above 99% percentile for each feature\nimport pprint\nhighest = {}\nfor column in train_df.columns:\n    if train_df[column].dtypes != \"object\": # exclude string data typed columns\n        highest[column]=[]\n        q = train_df[column].quantile(0.99)\n        highest[column] = train_df[train_df[column] > q].index.tolist()\n    \npprint.pprint(highest)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"321a895f59b073a47e436242c1ffea42289454b7"},"cell_type":"code","source":"# delete 'PassengerId' from dictionary highest\nhighest.pop('PassengerId', 0)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5a71b40d6d4ae49421005c1e22a554cc8508aabb"},"cell_type":"markdown","source":"_outliers repeatedly shown among the features_"},{"metadata":{"trusted":true,"_uuid":"8e3f86c6d1b23feee2d754e0547d44afdac37759"},"cell_type":"code","source":"# summarize the previous dictionary, highest\n# create a dictionary of outliers and the frequency of being outlier\nhighest_count = {}\nfor feature in highest:\n    for person in highest[feature]:\n        if person not in highest_count:\n            highest_count[person] = 1\n        else:\n            highest_count[person] += 1\n             \nhighest_count","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2562629d5d863f37b27fb6cc39d9bbcee203c614"},"cell_type":"code","source":"# This time, I defined outliers as being below of 1% percentile here\n# get lists of people below 1% percentile for each feature\nlowest = {}\nfor column in train_df.columns:\n    if train_df[column].dtypes != \"object\": # exception string \n        lowest[column]=[]\n        q = train_df[column].quantile(0.01)\n        lowest[column] = train_df[train_df[column] < q].index.tolist()\n\n# supp 'PassengerId' \nlowest.pop('PassengerId', 0)\n\npprint.pprint(lowest)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2fd66ab144a3960ef8271b5639441a1da4e937c6"},"cell_type":"code","source":"for person in lowest['Age']:\n    if person not in highest_count:\n        highest_count[person] = 1\n    else:\n        highest_count[person] += 1\n \nhighest_count","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"84cf3015acc35f355caf431569fb5aca3dee7f28"},"cell_type":"markdown","source":"_Overall, there is no outlier that are repeatedly shown among the features._\n\n_We can focus on age and Fare for continous values and Parch and SibSp for integer values to further outliers detection._"},{"metadata":{"_uuid":"974e7e5e51741803bc98735c5ef004a45b35b818"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\">\n<b>Outliers Detection:</b> Fare\n</div>"},{"metadata":{"trusted":true,"_uuid":"3714675ef4f39d4e4c8a4a587814e8c1d9b570a3"},"cell_type":"code","source":"# fare >99%\ntrain_df.loc[highest['Fare'],['Fare', 'Survived']]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"23709a102e6a3dfd4514e97f88c5b093bcc948f2"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-warning\">\n<b>Note :</b>\nThe mean fare is 32 units but there are outliers who paid 262, 263, or 512 units, which are 8 to 16 times higher than the mean fare. I am going to keep these outliers because this might help to classify survival as extreme cases in such decision tree algorithm.</div>"},{"metadata":{"_uuid":"c211fb57513696b2c1cb78d4d30bedc386572d6a"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\">\n<b>Outliers detection :</b> Age\n</div>"},{"metadata":{"trusted":true,"_uuid":"59a27e97a5e1bdbc0294caba79a80286dd90bb73"},"cell_type":"code","source":"# age above 99% percentile\ntrain_df.loc[highest['Age'],['Age', 'Survived']]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0b2144e9ddb9cc3c6a06ce8963d9e2e6de43ddc2"},"cell_type":"code","source":"# age below 1% percentile\ntrain_df.loc[lowest['Age'],['Age','Survived']]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7a9fd68e1b9cf65a76b6660fd8bba7d6dbd7bd20"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-warning\">\n<b>Note :</b>The average age is 29, but there are outliers who are 66 to 80 years old, which are 2.3 to 2.8 times higher than the average age, and who are younger than one year old. I am going to keep these outliers for now because the extreme cases of age might be usefull.</div>"},{"metadata":{"_uuid":"bdc3fd3a88c6c7fef4702fce0afa42126fa79d2e"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\">\n<b>b.Missing Values :</b> \n</div>\n"},{"metadata":{"scrolled":true,"trusted":true,"_uuid":"edd07e9f6e023bb59108005fb223a5885c068843"},"cell_type":"code","source":"\ntrain_df.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"scrolled":true,"trusted":true,"_uuid":"4cd5720364024de3b0ecb18a45e8259432626bf5"},"cell_type":"code","source":"\ntest_df.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d35c1d903c885b677f1f07452c6f81f97f022e2e"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\">\n<b>Missing values :</b> Fare\n</div>"},{"metadata":{"trusted":true,"_uuid":"1d1a51fb65529a5a2530e882eba5dc8ec25322cd"},"cell_type":"code","source":"\ntest_df[\"Fare\"].isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3c75c23df890c9257db036ab856f52fb730b4e22"},"cell_type":"code","source":"\ntest_df[\"Fare\"] = test_df[\"Fare\"].fillna(test_df[\"Fare\"].median())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1cdcd3126ca95254bbd2073620b08ae43c220692"},"cell_type":"code","source":"\ntest_df[\"Fare\"].isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6ca3a1f388a23a57f7f880c54cd8e6d189406d4e"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\">\n<b>Missing values:</b> Embarked\n</div>\n"},{"metadata":{"trusted":true,"_uuid":"69ceeace4c254b5c20f714b1958911451b347e82"},"cell_type":"code","source":"\ntrain_df[\"Embarked\"].isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"af185bb88b469a1ea540356cae2b9a7b28560c28"},"cell_type":"code","source":"\nPclass1 = train_df[train_df['Pclass']==1]['Embarked'].value_counts()\nPclass2 = train_df[train_df['Pclass']==2]['Embarked'].value_counts()\nPclass3 = train_df[train_df['Pclass']==3]['Embarked'].value_counts()\ndf = pd.DataFrame([Pclass1, Pclass2, Pclass3])\ndf.index = ['1st class','2nd class', '3rd class']\ndf.plot(kind='bar',stacked=True, figsize=(10,5));","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4572c219b8a37f94a6ac704564a88b9d72488f36"},"cell_type":"code","source":"train_test_data = [train_df, test_df]\nfor dataset in train_test_data:\n    dataset['Embarked'] = dataset['Embarked'].fillna('S')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d902d843f782885a77c6a6f0de03223969513844"},"cell_type":"code","source":"train_df[\"Embarked\"].isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"27baf761bf1cb98b0453f8e52da0c6e983554de8"},"cell_type":"code","source":"train_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fb0e43178961a8eb380642d5757764c5befd5d36"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\">\n<b>Missing values :</b> Cabin\n</div>"},{"metadata":{"_uuid":"e8767b08dbad709fb4b3443f610bb5d8468ed233"},"cell_type":"markdown","source":"_On pense que les personnes ayant un numéro de cabine enregistré appartiennent à une classe socioéconomique supérieure sont donc plus susceptibles de survivre._"},{"metadata":{"trusted":true,"_uuid":"640a21ffa1debf31a2388d0908aabba4636e6e6e"},"cell_type":"code","source":"train_df[\"Cabin\"] = (train_df[\"Cabin\"].notnull().astype('int'))\ntest_df[\"Cabin\"] = (test_df[\"Cabin\"].notnull().astype('int'))\n\n\nprint(\"survival % in cabins  = 1 :\", train_df[\"Survived\"][train_df[\"Cabin\"] == 1].value_counts(normalize = True)[1]*100)\n\nprint(\"survival % in cabins  = 0 :\", train_df[\"Survived\"][train_df[\"Cabin\"] == 0].value_counts(normalize = True)[1]*100)\n\nsns.barplot(x=\"Cabin\", y=\"Survived\", data=train_df)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"65b4a54f03366391fd904fd801fae89745fc85cf"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-warning\"><b>Note :</b><br>\npeople with registred cabin have more chance of survival (66,6% vs 29,9%)<br>\nso it would be legit to delete cabin from dataset.\n</div>"},{"metadata":{"trusted":true,"_uuid":"31c4158166f9e6706ee97667b1e1c8cdc58f811d"},"cell_type":"code","source":"\ntrain_df = train_df.drop(['Cabin'], axis = 1)\ntest_df = test_df.drop(['Cabin'], axis = 1)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"20ac491683bd2a37a1ec927067049f426698318c"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\">\n<b>Missing Values :</b> Age\n</div>\n"},{"metadata":{"trusted":true,"_uuid":"59cbc9f943dbbeb5563e811a4116bad503f9e7e9"},"cell_type":"code","source":"#train\nprint('pourcentage of missing values in \"Age\" is %.2f%%' %((train_df['Age'].isnull().sum()/train_df.shape[0])*100))\ntrain_df[\"Age\"].isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9c4a97ce022714337ad5d31ca3a28d04c156e284"},"cell_type":"code","source":"#test\nprint('pourc of missing values in \"Age\" is %.2f%%' %((test_df['Age'].isnull().sum()/test_df.shape[0])*100))\ntest_df[\"Age\"].isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"67cff29fe79be07d57c90dcac534d03434d0f538"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\">\n<b>Note </b>: \nSince some subpopulations are more likely to survive (eg children), it is best to keep the age function and impute the missing values.\n<b>Solution :</b> To solve this problem, we examined the most correlated characteristics with <b> Age (Sex, Parch, Pclass and SibSP)</b>,i guess this solution is more logical  because it is based on correlated data ,moreover a baby or child he will be in company with their parents or their brothers\n</div>\n"},{"metadata":{"trusted":true,"_uuid":"f9bab67197861e7c46a45902952ddc567f7b2da4"},"cell_type":"code","source":"g = sns.factorplot(y=\"Age\",x=\"Sex\",data=train_df,kind=\"box\")\ng = sns.factorplot(y=\"Age\",x=\"Sex\",hue=\"Pclass\", data=train_df,kind=\"box\")\ng = sns.factorplot(y=\"Age\",x=\"Parch\", data=train_df,kind=\"box\")\ng = sns.factorplot(y=\"Age\",x=\"SibSp\", data=train_df,kind=\"box\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c53fff3779b0fcc06f9dd198fc97f09fc6de8b2a"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-warning\">\n<b>Note :</b><br>The age distribution seems to be the same in the subpopulations of men and women, so <b> sex is not informative for age predicting.</b>\nHowever, <b> 1st class passengers are older than 2nd class passengers </b> and are also older than 3rd class passengers.\nMoreover, the more a passenger has parents/children the older he is and the more a passenger has siblings/spouses the younger he is .</div>"},{"metadata":{"trusted":true,"_uuid":"9c35d1e159f1c91e0fe3b7dfd2f05fef74d6ff5d"},"cell_type":"code","source":"\ng = sns.heatmap(train_df[[\"Age\",\"Sex\",\"SibSp\",\"Parch\",\"Pclass\"]].corr(),cmap=\"BrBG\",annot=True)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8b266a034dc7c711b1f1460e0043ac84716754a1"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-warning\">\n<b>the strategie :</b>is to fill in Age with the median age of similar rows according to Pclass, Parch and SibSp.</div>"},{"metadata":{"trusted":true,"_uuid":"afd9b05a6606f58258637bb2f96da60cb8404f09"},"cell_type":"code","source":"dataset =  train_df \n\n# NaN values\nindex_NaN_age = list(dataset[\"Age\"][dataset[\"Age\"].isnull()].index)\n\nfor i in index_NaN_age :\n    age_med = dataset[\"Age\"].median()#Median of age\n    age_pred = dataset[\"Age\"][((dataset['SibSp'] == dataset.iloc[i][\"SibSp\"]) & (dataset['Parch'] == dataset.iloc[i][\"Parch\"]) & (dataset['Pclass'] == dataset.iloc[i][\"Pclass\"]))].median()\n    if not np.isnan(age_pred) :\n        dataset['Age'].iloc[i] = age_pred\n    else :\n        dataset['Age'].iloc[i] = age_med","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2051cb92e41b480ec40f755d13194523bbae6410"},"cell_type":"code","source":"g = sns.factorplot(x=\"Survived\", y = \"Age\",data = train_df, kind=\"box\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"60fcfbd6258cf4638c82c79485ceac4be2e688fc"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-warning\">\n<b>Note :</b>\nNo difference between median value of age in survived and not survived subpopulation.<br>\n\nBut in the violin plot of survived passengers, we still notice that very young passengers have higher survival rate.</div>\n"},{"metadata":{"trusted":true,"_uuid":"ced9e579f04ce49769021c4870e817de1fd6fdd8"},"cell_type":"code","source":"\ntrain_df[\"Age\"].isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a2ad016653b678c104954d26d06ef3a0a8666d21"},"cell_type":"code","source":"\n\ndataset =  test_df \n\n# Nan values\nindex_NaN_age = list(dataset[\"Age\"][dataset[\"Age\"].isnull()].index)\n\nfor i in index_NaN_age :\n    age_med = dataset[\"Age\"].median()#Median of age\n    age_pred = dataset[\"Age\"][((dataset['SibSp'] == dataset.iloc[i][\"SibSp\"]) & (dataset['Parch'] == dataset.iloc[i][\"Parch\"]) & (dataset['Pclass'] == dataset.iloc[i][\"Pclass\"]))].median()\n    if not np.isnan(age_pred) :\n        dataset['Age'].iloc[i] = age_pred\n    else :\n        dataset['Age'].iloc[i] = age_med","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7f4794ffae2351aecbded2e7e0fdcbe0e4f3b7b4"},"cell_type":"code","source":"test_df[\"Age\"].isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a759cacb725a5153fee3df52741bd206af32e546"},"cell_type":"code","source":"train_df.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1001e8800867215dbdeb45442a1ba3f2b205027e"},"cell_type":"code","source":"\ntest_df.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"111c0022ac90de0779ae64359738b1fbd55aea19"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\">\n<b>6) Feature engeneering</b></div>"},{"metadata":{"_uuid":"d635c1711467ee8c4c60df2eb05db313dcea51ae"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\">\n<b> NAME</b> \n</div>"},{"metadata":{"trusted":true,"_uuid":"4479e114c6a8ac25a2b51210d9df26c00be5f0a7"},"cell_type":"code","source":"train_df[\"Name\"].head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e2d8746dec92a53002fb47ec22a4634f1e4d4bd3"},"cell_type":"markdown","source":"\n_The Name feature contains information on passenger's title.\nSince some passenger with distingused title may be preferred during the evacuation, it is interesting to add them to the model._\n"},{"metadata":{"trusted":true,"_uuid":"948fb11fec3b5b292966a19c6a3a27ecb810a9ba"},"cell_type":"code","source":"train_test_data = [train_df, test_df]\n\nfor dataset in train_test_data:\n    dataset['Titre'] = dataset['Name'].str.extract(' ([A-Za-z]+)\\.', expand=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"66b0591bd0dbf21ce50d5f138ca39d96ee7cf93a"},"cell_type":"code","source":"\ntrain_df['Titre'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7ad89b43c6696d9656f19b02489dec16ac254059"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-warning\">\nThere is 17 titles in the dataset, most of them are very rare and we can group them in 4 categories.</div>"},{"metadata":{"trusted":true,"_uuid":"4064a82dcf209b46746e71bc6c9367c958737a13"},"cell_type":"code","source":"train_len = len(train_df)\ndataset =  pd.concat(objs=[train_df, test_df], axis=0).reset_index(drop=True)\n\ng = sns.countplot(x=\"Titre\",data=dataset)\ng = plt.setp(g.get_xticklabels(), rotation=45) ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d7d037f56bd4f1b03cd836ab1c63838cba2a0b16"},"cell_type":"code","source":"\ndataset[['Titre', 'Survived']].groupby(['Titre'], as_index=False).mean()","execution_count":null,"outputs":[]},{"metadata":{"scrolled":true,"trusted":true,"_uuid":"1e4b7432833ffbd79c200d016c1ca5f17b179ad3"},"cell_type":"code","source":"\nbar_chart('Titre')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4f47f5293d26336b423f3d735b4ae6bb0fc4e494"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-warning\">\n    <b>Note : </b>\"Women and children first\",It is interesting to note that passengers with a rare title are more likely to survive.</div>"},{"metadata":{"trusted":true,"_uuid":"86630f1faffb77ecfc870d731d438ebcb2639fe2"},"cell_type":"code","source":"#mapping list\nmapping_nom = {\"Mr\": 0, \"Miss\": 1, \"Mrs\": 2, \"Master\": 3, \"Dr\": 3, \"Rev\": 3, \"Col\": 3, \"Major\": 3, \"Mlle\": 3,\"Countess\": 3,\n                 \"Ms\": 3, \"Lady\": 3, \"Jonkheer\": 3, \"Don\": 3, \"Dona\" : 3, \"Mme\": 3,\"Capt\": 3,\"Sir\": 3 }\n#application au dataset(train+test)\nfor dataset in train_test_data:\n    dataset['Titre'] = dataset['Titre'].map(mapping_nom)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e58d193ac6ca073dc08a35b8b742624726c63882"},"cell_type":"code","source":"train_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"57e0101bbc2f943b1ccae559ec8ee8ee5d42536b"},"cell_type":"code","source":"\ntest_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"55b0069340be40bd1823e440ae510c439cb8dc77"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\">\n<b>SEXE</b> \n</div>"},{"metadata":{"trusted":true,"_uuid":"ab382e8b7c069572b87a42e2dc8922d32a638cc3"},"cell_type":"code","source":"# M:0 et F:1\nsexe_mapping = {\"male\": 0, \"female\": 1}\nfor dataset in train_test_data:\n    dataset['Sex'] = dataset['Sex'].map(sexe_mapping)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f409833cad3aa0ac36d8b93256f855437e0c17c2"},"cell_type":"code","source":"bar_chart('Sex')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6f0dbf6bb6da5736877fefe28a0479c049091a25"},"cell_type":"code","source":"\ntrain_df.drop('Name', axis=1, inplace=True)\ntest_df.drop('Name', axis=1, inplace=True)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9a87c9da2e860275b14b2f7b9250e483ee8b8c3e"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\">\n<b>EMBARKED </b> \n</div>"},{"metadata":{"trusted":true,"_uuid":"3f77b59dae88ffa770e3f070f581474a4e017f8f"},"cell_type":"code","source":"\nembarked_mapping = {\"S\": 0, \"C\": 1, \"Q\": 2}\nfor dataset in train_test_data:\n    dataset['Embarked'] = dataset['Embarked'].map(embarked_mapping)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c07a43173d24f4582919e27db30e56ac51e64d7e"},"cell_type":"code","source":"train_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9b306e66ab04d814d8ed0e129b38dcaef6a5ef76"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\">\n<b>FAMILY SIZE</b> \n</div>"},{"metadata":{"trusted":true,"_uuid":"bad1cd23708ac0ced9c7f4786658fc313cc06e6e"},"cell_type":"code","source":"train_df[\"taillefamille\"] = train_df[\"SibSp\"] + train_df[\"Parch\"] + 1\ntest_df[\"taillefamille\"] = test_df[\"SibSp\"] + test_df[\"Parch\"] + 1","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"975d307d485b47c62a857f8fbb3c43a2e7674be9"},"cell_type":"code","source":"\ng = sns.factorplot(x=\"taillefamille\",y=\"Survived\",data = train_df)\ng = g.set_ylabels(\"Probabilité de survie\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"64d04dc9f88b3ca74995e838771d6fdeec8bc5ab"},"cell_type":"markdown","source":"_la taille de famille semble avoir un role important dans la survie des passagers pour ceci on a décidé de les classer encore en 4 catégories_"},{"metadata":{"trusted":true,"_uuid":"48e08404286fe4aea1aca2eb982af21f81d6c2e2"},"cell_type":"code","source":"#solo= alone / pfamille=smallfamily / Mfamille=Medium / Gfamille=Bigfamily\ntrain_df['Solo'] = train_df['taillefamille'].map(lambda s: 1 if s == 1 else 0)\ntrain_df['Pfamille'] = train_df['taillefamille'].map(lambda s: 1 if s == 2  else 0)\ntrain_df['Mfamille'] = train_df['taillefamille'].map(lambda s: 1 if 3 <= s <= 4 else 0)\ntrain_df['Gfamille'] = train_df['taillefamille'].map(lambda s: 1 if s >= 5 else 0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1f3f7116e8277918049004191ba80410bbaa9032"},"cell_type":"code","source":"\ng = sns.factorplot(x=\"Solo\",y=\"Survived\",data=train_df,kind=\"bar\")\ng = g.set_ylabels(\"Survival Probability\")\ng = sns.factorplot(x=\"Pfamille\",y=\"Survived\",data=train_df,kind=\"bar\")\ng = g.set_ylabels(\"Survival Probability\")\ng = sns.factorplot(x=\"Mfamille\",y=\"Survived\",data=train_df,kind=\"bar\")\ng = g.set_ylabels(\"Survival Probability\")\ng = sns.factorplot(x=\"Gfamille\",y=\"Survived\",data=train_df,kind=\"bar\")\ng = g.set_ylabels(\"Survival Probability\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1566a6641efc0748557434594682b82dc168daf4"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-warning\">\n    Factorplots of family size categories show that <b>Small and Medium</b> families have more chance to survive than single passenger and large families.</div>"},{"metadata":{"trusted":true,"_uuid":"287ecbe66e569e76e955726c1d09e96a498f39fd"},"cell_type":"code","source":"train_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8418534e9bb2144f25812620d0efc2986ceb1317"},"cell_type":"code","source":"#for test data\ntest_df['Solo'] = test_df['taillefamille'].map(lambda s: 1 if s == 1 else 0)\ntest_df['Pfamille'] = test_df['taillefamille'].map(lambda s: 1 if s == 2  else 0)\ntest_df['Mfamille'] = test_df['taillefamille'].map(lambda s: 1 if 3 <= s <= 4 else 0)\ntest_df['Gfamille'] = test_df['taillefamille'].map(lambda s: 1 if s >= 5 else 0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1364db027178fedb6bb725aa6a6f369ffb480bef"},"cell_type":"code","source":"test_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b72d690875bae73b9dcddc28da4b550be5a2e3ee"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\">\n<b></b> Removing unnecessary attributes\n</div>"},{"metadata":{"trusted":true,"_uuid":"7c80d6198772ff0a07bf1ebcaf210d6ea5d7d10b"},"cell_type":"code","source":"\nfeatures_drop = ['Ticket', 'SibSp', 'Parch']\n\ntrain_df = train_df.drop(features_drop, axis=1)\ntest_df = test_df.drop(features_drop, axis=1)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"33a33af86e47e2ac097980de96b23e8156a76e3a"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\">\n<b>Final Step </b> \n</div>"},{"metadata":{"trusted":true,"_uuid":"77dd3fd37bedde60657830d23ddf8d36c18be687"},"cell_type":"code","source":"\ntrain_df.head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a9083802397003fd3c07a5d0f903701327a37a0a"},"cell_type":"code","source":"\ntest_df.head(5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"02cf8914e2716eb327f06ad5b8268ef016240679"},"cell_type":"code","source":"train_df.info()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"acc64399202739839f240d8134bf75022c74dee4"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\">\n<b>7) Modeling Section</b> </div>"},{"metadata":{"_uuid":"676f33c6351702daeb30c59fba1155bae59fec10"},"cell_type":"markdown","source":"Testing Models:\n- k-Nearest Neighbors KNN\n+ Decision Tree Classifier\n+ Random Forest Classifier\n+ Gradient Boosting Classifier\n"},{"metadata":{"trusted":true,"_uuid":"2b58bde4aac419bef03a801f20710541c6753d04"},"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\nX_all = train_df.drop(['Survived', 'PassengerId'], axis=1)\nY_all = train_df['Survived']\n\nnum_test = 0.20 #20% for test\nX_train, X_test, Y_train, Y_test = train_test_split(X_all, Y_all, test_size=num_test, random_state=23)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6626502a61ab05dc088445bfcaab37d3118235e5"},"cell_type":"code","source":"#KNN\n\nfrom sklearn.metrics import make_scorer, accuracy_score\nfrom sklearn.model_selection import GridSearchCV\n\n# Choose the type of classifier. \nknn = KNeighborsClassifier()\n\n# Choose some parameter combinations to try\nparameters = {'n_neighbors':[3, 5, 7],\n              'weights':['uniform'], \n              'algorithm':['auto'], \n              'leaf_size':[30]\n             }\n\n# Type of scoring used to compare parameter combinations\nacc_scorer = make_scorer(accuracy_score)\n\n# Run the grid search\ngrid_obj = GridSearchCV(knn, parameters, scoring=acc_scorer)\ngrid_obj = grid_obj.fit(X_train, Y_train)\n\n# Set the clf to the best combination of parameters\nclf = grid_obj.best_estimator_\n\n# Fit the best algorithm to the data. \nclf.fit(X_train, Y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"dc7d39325313a404d909bb268786ce158d6f140f"},"cell_type":"code","source":"predictions = clf.predict(X_test)\nprint(accuracy_score(Y_test, predictions))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a7a9a12f962357704d474a876578e4be4d78b3f8"},"cell_type":"code","source":"\n# Gradient Boosting Classifier\nfrom sklearn.ensemble import GradientBoostingClassifier\n# Choose the type of classifier. \nGBC = GradientBoostingClassifier()\n\n# Choose some parameter combinations to try\nparameters = {\n              'max_depth': [1, 2, 3, 4, 5],\n              'max_features': [1, 2, 3, 4]}\n             \n\n# Type of scoring used to compare parameter combinations\nacc_scorer = make_scorer(accuracy_score)\n\n# Run the grid search\ngrid_obj = GridSearchCV(GBC, parameters, scoring=acc_scorer)\ngrid_obj = grid_obj.fit(X_train, Y_train)\n\n# Set the clf to the best combination of parameters\nclf = grid_obj.best_estimator_\n\n# Fit the best algorithm to the data. \nclf.fit(X_train, Y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3fadec4f0f14a18aa11c948bebbad56b398e693f"},"cell_type":"code","source":"predictions = clf.predict(X_test)\nprint(accuracy_score(Y_test, predictions))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"781a67332eac56c19f6e17764ce9dcee76458fcb"},"cell_type":"code","source":"#Random Forest Classifier\n\nfrom sklearn.ensemble import RandomForestClassifier\n# Choose the type of classifier. \nRFC = RandomForestClassifier()\n\n# Choose some parameter combinations to try\nparameters = {\n              'max_depth': [1, 2, 3, 4, 5],\n              'max_features': [1, 2, 3, 4]}\n             \n\n# Type of scoring used to compare parameter combinations\nacc_scorer = make_scorer(accuracy_score)\n\n# Run the grid search\ngrid_obj = GridSearchCV(RFC, parameters, scoring=acc_scorer)\ngrid_obj = grid_obj.fit(X_train, Y_train)\n\n# Set the clf to the best combination of parameters\nclf = grid_obj.best_estimator_\n\n# Fit the best algorithm to the data. \nclf.fit(X_train, Y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"16bc7cf0a0967eff2157c2015599c37e4326fe92"},"cell_type":"code","source":"predictions = clf.predict(X_test)\nprint(accuracy_score(Y_test, predictions))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7c7305d99bb45be6be6c10471fbe406a4a5c2458"},"cell_type":"code","source":"#CART\n\n# Choose the type of classifier. \nCART = DecisionTreeClassifier()\n\n# Choose some parameter combinations to try\nparameters = {'max_depth': [1, 2, 3, 4, 5],\n              'max_features': [1, 2, 3, 4]}\n             \n\n# Type of scoring used to compare parameter combinations\nacc_scorer = make_scorer(accuracy_score)\n\n# Run the grid search\ngrid_obj = GridSearchCV(CART, parameters, scoring=acc_scorer)\ngrid_obj = grid_obj.fit(X_train, Y_train)\n\n# Set the clf to the best combination of parameters\nclf = grid_obj.best_estimator_\n\n# Fit the best algorithm to the data. \nclf.fit(X_train, Y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9524daf82551a83126a69946ded3c9c0cb4f34b0"},"cell_type":"code","source":"predictions = clf.predict(X_test)\nprint(accuracy_score(Y_test, predictions))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"37926699e971f5b9e8e9a76e1d9e02d88ffee26c"},"cell_type":"markdown","source":"so i decided to use the Decision tree classifier model for the testing data. "},{"metadata":{"_uuid":"c8730eaeb813a484e2094aeebe28c24c801ca327"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\">\n<b>5.Test du Modèle:</b> d'après ce qui precède on va continuer par le modèle CART\n</div>"},{"metadata":{"trusted":true,"_uuid":"4cce8ca191a507aff50acb9af2b6ce7eaac8f326"},"cell_type":"code","source":"\n\npredictions = clf.predict(X_test)\nprint(accuracy_score(Y_test, predictions))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"aea955103fc9a8c096fb009d7d7f5847616c84d4"},"cell_type":"code","source":"#Confusion matrix\n\npredictions = clf.predict(X_test)\nprint(confusion_matrix(Y_test, predictions))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"bb0163166b8db30c0549a8252abb90bd3ffbdd96"},"cell_type":"code","source":"predictions = clf.predict(X_test)\nprint(classification_report(Y_test, predictions))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"99dc2664c0dfe3a5aba9a4cb1368898b564d0361"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\">\n<b>6.Extraction of Test_results</b><br>\n</div>"},{"metadata":{"trusted":false,"_uuid":"9d27e85d4bde55ec09a9cdab16405b5b59df188a"},"cell_type":"code","source":"ids = test_df['PassengerId']\npredictions = clf.predict(test_df.drop('PassengerId', axis=1))\n\noutput = pd.DataFrame({ 'PassengerId' : ids, 'Survived': predictions })\noutput.to_csv('test_result.csv', index = False)\noutput.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b838406eb80285ac07d284c1aead804190b1a992"},"cell_type":"code","source":"from IPython.display import Image\nImage(url= \"https://static1.squarespace.com/static/5006453fe4b09ef2252ba068/t/5090b249e4b047ba54dfd258/1351660113175/TItanic-Survival-Infographic.jpg?format=1500w\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0574ca80226cdeda7053d31d1d7fd0df5cd9bf77"},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\">\n<b>End of this nb</b>\n</div>"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.7.1"}},"nbformat":4,"nbformat_minor":1}