{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **TITANIC** database analysis and prediction\n\nOur objective in this study is to predict whether a person survived the Titanic's crash or not\n\nThe minimum metric we want to achieve for our prediction is 90% or even better for our models. ","metadata":{}},{"cell_type":"markdown","source":"## Part I : __Exploratory analyzes of the TITANIC database__\n","metadata":{}},{"cell_type":"markdown","source":"For our work we will need the following packages: ","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nfrom sklearn.preprocessing import OrdinalEncoder\nfrom sklearn.model_selection import train_test_split, StratifiedShuffleSplit \nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.naive_bayes import GaussianNB\nfrom sklearn.neural_network import MLPClassifier\nfrom sklearn.metrics import accuracy_score  \nprint(\"import success\")","metadata":{"execution":{"iopub.status.busy":"2022-08-07T19:54:41.940017Z","iopub.execute_input":"2022-08-07T19:54:41.940500Z","iopub.status.idle":"2022-08-07T19:54:42.928581Z","shell.execute_reply.started":"2022-08-07T19:54:41.940412Z","shell.execute_reply":"2022-08-07T19:54:42.926880Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Database import","metadata":{}},{"cell_type":"code","source":"dataf = pd.read_csv(\"../input/titanic/train.csv\", index_col=\"PassengerId\")","metadata":{"execution":{"iopub.status.busy":"2022-08-07T19:54:42.930820Z","iopub.execute_input":"2022-08-07T19:54:42.931521Z","iopub.status.idle":"2022-08-07T19:54:42.960638Z","shell.execute_reply.started":"2022-08-07T19:54:42.931471Z","shell.execute_reply":"2022-08-07T19:54:42.959420Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's visualize the structure of our Database","metadata":{}},{"cell_type":"code","source":"dataf.info()\ndataf.tail(10)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T19:54:42.962300Z","iopub.execute_input":"2022-08-07T19:54:42.964205Z","iopub.status.idle":"2022-08-07T19:54:43.006445Z","shell.execute_reply.started":"2022-08-07T19:54:42.964155Z","shell.execute_reply":"2022-08-07T19:54:43.005142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Our data set consists of 11 columns and 891 individuals (observations) of which:\n\n* survival: The variable indicates whether the individual survived or not with the modality 0 = *No* and 1 = *Yes*\n* pclass :The variable indicates which class the individual took on the boat with the modality 1 = *1st class*, 2 = *2nd class* and 3 = *3rd class*\n* sex : The variable indicates the sex of the individual with male = *man* , female= *Woman*\n* Age: The variable indicates the age of the individual  \n* sibsp: The variable indicates if the individual has his brother, sister, or spouse on board   \n* parch : The variable indicates whether the individual has parents or children on board.\n* ticket: The variable indicates the ticket purchased by the passenger\n* fare: the variable indicates the amount that the individual paid\n* cabin: The variable indicates the cabin number of the individual\n* embarked: The variable indicates the port of embarkation with C = *Cherbourg*, Q = *Queenstown*, S = *Southampton*","metadata":{}},{"cell_type":"markdown","source":"We identify our *target variable* as the variable *survived* which is the variable to be predicted. \n\nWe identify 10 other variables of which 8 are qualitative (*Sex*, *Sibsp*, *Parch*, *Pclass*, *Ticket*, *Cabin*, *Embarked* and *Name*) *and* \n2 quantitative (*Age* and *Fare*)\n","metadata":{}},{"cell_type":"markdown","source":"__check if our dataset contains missing values__ ","metadata":{}},{"cell_type":"code","source":"# Afficons les varibles ayant les valeurs manquante par ordres decroissant de valeur manquant \ndataf.isnull().sum().sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T19:54:43.009988Z","iopub.execute_input":"2022-08-07T19:54:43.010891Z","iopub.status.idle":"2022-08-07T19:54:43.024343Z","shell.execute_reply.started":"2022-08-07T19:54:43.010823Z","shell.execute_reply":"2022-08-07T19:54:43.023302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Cabin and age variables contain missing values : *687* and *177* respectively.\n\nThe variable *Cabin* contains 77% of missing values and therefore the variable cannot give us positive information so the best solution would be to delete it.\n\nSimilarly the variable *Name* does not help us to explain the dependent variable (target variable) so we can freely delete it too.","metadata":{}},{"cell_type":"code","source":"#Deleting of Name, Cabin, and Ticket variables\ndataf = dataf.drop([\"Name\", \"Cabin\", \"Ticket\"], axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T19:54:43.026111Z","iopub.execute_input":"2022-08-07T19:54:43.026923Z","iopub.status.idle":"2022-08-07T19:54:43.041195Z","shell.execute_reply.started":"2022-08-07T19:54:43.026852Z","shell.execute_reply":"2022-08-07T19:54:43.039731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The variable *Age* only contains a few missing values, so we can impute those missing values. As such we can replace them with the mean of the variable.","metadata":{}},{"cell_type":"code","source":"# Imputing the missing values \ndataf[\"Age\"] = dataf[\"Age\"].fillna(dataf[\"Age\"].mean())","metadata":{"execution":{"iopub.status.busy":"2022-08-07T19:54:43.043443Z","iopub.execute_input":"2022-08-07T19:54:43.044571Z","iopub.status.idle":"2022-08-07T19:54:43.052267Z","shell.execute_reply.started":"2022-08-07T19:54:43.044520Z","shell.execute_reply":"2022-08-07T19:54:43.051154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The *Embarked* variable is the only one left with missing values (2), we will replace these missing values with the values of the variable that repeats the most.","metadata":{}},{"cell_type":"code","source":"dataf = dataf.apply(lambda x:x.fillna(x.value_counts().index[0]))","metadata":{"execution":{"iopub.status.busy":"2022-08-07T19:54:43.053948Z","iopub.execute_input":"2022-08-07T19:54:43.054712Z","iopub.status.idle":"2022-08-07T19:54:43.075623Z","shell.execute_reply.started":"2022-08-07T19:54:43.054675Z","shell.execute_reply":"2022-08-07T19:54:43.074172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### __Descriptive Analysis around to understand our target variable__","metadata":{}},{"cell_type":"markdown","source":"We will start by studing our target variable.","metadata":{}},{"cell_type":"code","source":"dataf.Survived.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T19:54:43.077045Z","iopub.execute_input":"2022-08-07T19:54:43.077420Z","iopub.status.idle":"2022-08-07T19:54:43.087096Z","shell.execute_reply.started":"2022-08-07T19:54:43.077388Z","shell.execute_reply":"2022-08-07T19:54:43.085637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The *Survived* variable tells us that out of the *891* individuals, *549* did not survive while *342* did.\n\nWe can visualize the variable target graphically as follows:","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(14,6))\nsns.countplot(dataf.Survived)\nOui = dataf.Survived.value_counts()[0]/len(dataf.Survived)\nNon = dataf.Survived.value_counts()[1]/len(dataf.Survived)\nprint(f\"The percentage of surviving passengers is : {Oui}\")\nprint(f\"The percentage of passengers who did not survive is : {Non}\")\n","metadata":{"execution":{"iopub.status.busy":"2022-08-07T19:54:43.088411Z","iopub.execute_input":"2022-08-07T19:54:43.090192Z","iopub.status.idle":"2022-08-07T19:54:43.319817Z","shell.execute_reply.started":"2022-08-07T19:54:43.090140Z","shell.execute_reply":"2022-08-07T19:54:43.318455Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The graph confirms our result obtained above with *38.38%* of individuals having survived and *61.61%* did not.","metadata":{}},{"cell_type":"markdown","source":"### Visualization of target - features relationship","metadata":{}},{"cell_type":"markdown","source":"__Target variable and quantitative variables__","metadata":{}},{"cell_type":"code","source":"plt.figure()\nsns.boxplot(data=dataf, y = \"Age\", x=\"Survived\" )\nplt.figure()\nsns.boxplot(data=dataf, y = \"Fare\", x=\"Survived\" )","metadata":{"execution":{"iopub.status.busy":"2022-08-07T19:54:43.322952Z","iopub.execute_input":"2022-08-07T19:54:43.323393Z","iopub.status.idle":"2022-08-07T19:54:43.719978Z","shell.execute_reply.started":"2022-08-07T19:54:43.323342Z","shell.execute_reply":"2022-08-07T19:54:43.718813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The Boxplot reveals to us:\n\n* The higher the amount paid by the individual, the more likely he was to survive.\n* Many of the children and elderly individuals did not survive.","metadata":{}},{"cell_type":"markdown","source":"**Target variable and other qualitative variables** ","metadata":{}},{"cell_type":"markdown","source":"We are going to make a cross table between the target variable and the other qualitative variables.","metadata":{}},{"cell_type":"code","source":"pd.crosstab(dataf[\"Survived\"], dataf[\"Sex\"])","metadata":{"execution":{"iopub.status.busy":"2022-08-07T19:54:43.721492Z","iopub.execute_input":"2022-08-07T19:54:43.721915Z","iopub.status.idle":"2022-08-07T19:54:43.749371Z","shell.execute_reply.started":"2022-08-07T19:54:43.721857Z","shell.execute_reply":"2022-08-07T19:54:43.748070Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To make it simple, we can write a function to allow us to sort all our variables two by two with our target variable.","metadata":{}},{"cell_type":"code","source":"var_qualitative = [col for col in dataf.select_dtypes(\"object\")]\nfor i in var_qualitative:\n    if i == \"Survived\":\n        pass\n    else:\n        print(pd.crosstab(dataf[\"Survived\"], dataf[i]))\n           ","metadata":{"execution":{"iopub.status.busy":"2022-08-07T19:54:44.031244Z","iopub.execute_input":"2022-08-07T19:54:44.032247Z","iopub.status.idle":"2022-08-07T19:54:44.065533Z","shell.execute_reply.started":"2022-08-07T19:54:44.032204Z","shell.execute_reply":"2022-08-07T19:54:44.064326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* __Survived__ and __Pclass__\n\nThe proportion of survivors is higher for people who are in the first class. The proportion of survivors is lower for people in the 3rd class.   \n\n* __Survived__ and __Sex__ \n\nThe proportion of women who survived is much higher than that of men, we can conclude that women were prioritized during the evacuation\n\n* __Survived__ and __SibSp__\n\nWe can conclude that the more the person has a high number of brothers and sisters, the more his survival rate decreases.\n\n* __Survived__ and __Embarked__\n\nThe proportion of surviving individuals is higher for individuals who embarked at the port of _Cherbourg_","metadata":{}},{"cell_type":"markdown","source":"## Part II : __Pre-Processing of our Data Set__ ","metadata":{}},{"cell_type":"markdown","source":"Our model does not include qualitative variables, for this we will encode our qualitative variables into quantitative variables.\n\nExample the variable sex will be encoded as follows: male = 0, female = 1.","metadata":{}},{"cell_type":"code","source":"# Function to encode our variable sex\ndef encod_sex(data):\n    if data[\"Sex\"]== \"male\":\n        data[\"Sex\"] = 0\n        return data\n    else:\n        data[\"Sex\"] = 1\n        return data\n    ","metadata":{"execution":{"iopub.status.busy":"2022-08-07T19:54:44.067965Z","iopub.execute_input":"2022-08-07T19:54:44.068460Z","iopub.status.idle":"2022-08-07T19:54:44.074380Z","shell.execute_reply.started":"2022-08-07T19:54:44.068426Z","shell.execute_reply":"2022-08-07T19:54:44.073381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Function to encode our variable Embarked \ndef encod_embar(data):\n    if data[\"Embarked\"]== \"Q\":\n        data[\"Embarked\"] = 0\n        return data\n    elif data[\"Embarked\"] == \"C\":\n        data[\"Embarked\"] = 1\n        return data\n    else:\n        data[\"Embarked\"] = 2\n        return data\n    ","metadata":{"execution":{"iopub.status.busy":"2022-08-07T19:54:44.075654Z","iopub.execute_input":"2022-08-07T19:54:44.076791Z","iopub.status.idle":"2022-08-07T19:54:44.088761Z","shell.execute_reply.started":"2022-08-07T19:54:44.076756Z","shell.execute_reply":"2022-08-07T19:54:44.087514Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Encoding of the sex variable\ndataf = dataf.apply(encod_sex, axis=\"columns\")\n\n# Encoding of the embarked variable \ndataf = dataf.apply(encod_embar, axis=\"columns\")","metadata":{"execution":{"iopub.status.busy":"2022-08-07T19:54:44.090491Z","iopub.execute_input":"2022-08-07T19:54:44.091178Z","iopub.status.idle":"2022-08-07T19:54:44.246907Z","shell.execute_reply.started":"2022-08-07T19:54:44.091070Z","shell.execute_reply":"2022-08-07T19:54:44.245491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can visualize our new dataset.","metadata":{}},{"cell_type":"code","source":"dataf.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T19:54:44.249983Z","iopub.execute_input":"2022-08-07T19:54:44.250663Z","iopub.status.idle":"2022-08-07T19:54:44.267331Z","shell.execute_reply.started":"2022-08-07T19:54:44.250612Z","shell.execute_reply":"2022-08-07T19:54:44.266158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Part III : __Model training with the algorithms: Logistic Regression / Naive bayes / KNN / Decision Tree / Neural Network__","metadata":{}},{"cell_type":"markdown","source":"Let's start by dividing our dataset into training data and test data.","metadata":{}},{"cell_type":"code","source":"y = dataf[\"Survived\"]\nX = dataf.drop(\"Survived\", axis=1 )\n\nX_train, X_test, y_train, y_test = train_test_split(X,y,test_size=0.2, train_size=0.8, \n                                                   random_state = 77, shuffle =True, \n                                                    stratify=y)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-07T19:54:44.268747Z","iopub.execute_input":"2022-08-07T19:54:44.269948Z","iopub.status.idle":"2022-08-07T19:54:44.282617Z","shell.execute_reply.started":"2022-08-07T19:54:44.269900Z","shell.execute_reply":"2022-08-07T19:54:44.281516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*Then lets define the algorithms that we will use*","metadata":{}},{"cell_type":"code","source":"models = {\n    'LogisticRegression':LogisticRegression(random_state=77, max_iter=200),\n    'DecisionTreeClassifier':DecisionTreeClassifier(max_depth=1, random_state=77),\n    'KNeighborsClassifier':KNeighborsClassifier(),\n    'GaussianNB': GaussianNB(),\n    'MLPClassifier':MLPClassifier(hidden_layer_sizes=(100,),random_state=77, \n                                  max_iter=300)\n    \n}","metadata":{"execution":{"iopub.status.busy":"2022-08-07T19:54:44.284283Z","iopub.execute_input":"2022-08-07T19:54:44.285307Z","iopub.status.idle":"2022-08-07T19:54:44.291886Z","shell.execute_reply.started":"2022-08-07T19:54:44.285266Z","shell.execute_reply":"2022-08-07T19:54:44.291065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"_Then we define function that will help us determine the precision of our models._","metadata":{}},{"cell_type":"code","source":"def precision(y_true, y_pred, retu =False):\n    acc = accuracy_score(y_true, y_pred)\n    if retu:\n        return acc\n    else:\n        print(f\"The precision of our model is: {acc}\") \n","metadata":{"execution":{"iopub.status.busy":"2022-08-07T19:54:44.293112Z","iopub.execute_input":"2022-08-07T19:54:44.293808Z","iopub.status.idle":"2022-08-07T19:54:44.305469Z","shell.execute_reply.started":"2022-08-07T19:54:44.293773Z","shell.execute_reply":"2022-08-07T19:54:44.304316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Fiting our models\n\ndef train_test_eval(models, X_train, y_train, X_test, y_test):\n    for name, model in models.items():\n        print(name,\":\")\n        model.fit(X_train,y_train)\n        precision(y_test,model.predict(X_test))\n        print(\"-\"*30)\n    ","metadata":{"execution":{"iopub.status.busy":"2022-08-07T19:54:44.616130Z","iopub.execute_input":"2022-08-07T19:54:44.616741Z","iopub.status.idle":"2022-08-07T19:54:44.622340Z","shell.execute_reply.started":"2022-08-07T19:54:44.616706Z","shell.execute_reply":"2022-08-07T19:54:44.621531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Launch the training and prediction \ntrain_test_eval(models, X_train, y_train, X_test, y_test)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-07T19:54:44.623997Z","iopub.execute_input":"2022-08-07T19:54:44.624592Z","iopub.status.idle":"2022-08-07T19:54:46.074580Z","shell.execute_reply.started":"2022-08-07T19:54:44.624557Z","shell.execute_reply":"2022-08-07T19:54:46.072942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can conclude that our four algorithms in particular: _Neural Network_(83.24%), _LogisticRegression_(82.12%), _DecisionTreeClassifier_(82.68%) and _Naive Bayes_(81.01%) have an accuracy of more than 80%.\n\n## The Metric of 80% has been successfully achieved !!!!!!! \n\nLets look foward to how to improve the prediction accuracy of our models. \n\n# See You Soon ...","metadata":{}}]}