{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Problem Statement\n\nThe sinking of the Titanic is one of the most infamous shipwrecks in history.\n\nOn April 15, 1912, during her maiden voyage, the widely considered “unsinkable” RMS Titanic sank after colliding with an iceberg. Unfortunately, there weren’t enough lifeboats for everyone onboard, resulting in the death of 1502 out of 2224 passengers and crew.\n\nWhile there was some element of luck involved in surviving, it seems some groups of people were more likely to survive than others.\n\nIn this challenge, we ask you to build a predictive model that answers the question: “what sorts of people were more likely to survive?” using passenger data (ie name, age, gender, socio-economic class, etc).\n\n### Datasets\nTrain.csv will contain the details of a subset of the passengers on board (891 to be exact) and importantly, will reveal whether they survived or not, also known as the “ground truth”.\n\nThe `test.csv` dataset contains similar information but does not disclose the “ground truth” for each passenger. It’s your job to predict these outcomes.\n\nYou can download the dataset from [here](https://www.kaggle.com/competitions/titanic/data)","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport sklearn\nfrom sklearn.preprocessing import OneHotEncoder, MinMaxScaler\nfrom sklearn.model_selection import train_test_split, KFold,StratifiedKFold\nfrom sklearn.metrics import accuracy_score,confusion_matrix,classification_report,roc_curve,roc_auc_score\nfrom sklearn.model_selection import RandomizedSearchCV, GridSearchCV\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.naive_bayes import BernoulliNB\nfrom sklearn.tree import DecisionTreeClassifier\nimport statsmodels.api as sm\nfrom statsmodels.formula.api import ols","metadata":{"execution":{"iopub.status.busy":"2022-07-27T10:34:42.144814Z","iopub.execute_input":"2022-07-27T10:34:42.146375Z","iopub.status.idle":"2022-07-27T10:34:42.155981Z","shell.execute_reply.started":"2022-07-27T10:34:42.146315Z","shell.execute_reply":"2022-07-27T10:34:42.154976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.style.use('seaborn-bright')","metadata":{"execution":{"iopub.status.busy":"2022-07-27T09:45:17.832143Z","iopub.execute_input":"2022-07-27T09:45:17.832684Z","iopub.status.idle":"2022-07-27T09:45:17.838324Z","shell.execute_reply.started":"2022-07-27T09:45:17.832644Z","shell.execute_reply":"2022-07-27T09:45:17.837192Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Importing the Data and Understanding the features in the data","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/titanic/train.csv')\ntest_df = pd.read_csv('/kaggle/input/titanic/test.csv')\nsubmission_df = pd.read_csv('/kaggle/input/titanic/gender_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-07-27T09:18:01.687260Z","iopub.execute_input":"2022-07-27T09:18:01.688187Z","iopub.status.idle":"2022-07-27T09:18:01.711282Z","shell.execute_reply.started":"2022-07-27T09:18:01.688139Z","shell.execute_reply":"2022-07-27T09:18:01.710416Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-27T09:18:01.991376Z","iopub.execute_input":"2022-07-27T09:18:01.991875Z","iopub.status.idle":"2022-07-27T09:18:02.009964Z","shell.execute_reply.started":"2022-07-27T09:18:01.991834Z","shell.execute_reply":"2022-07-27T09:18:02.009238Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-27T09:18:02.279544Z","iopub.execute_input":"2022-07-27T09:18:02.280297Z","iopub.status.idle":"2022-07-27T09:18:02.298138Z","shell.execute_reply.started":"2022-07-27T09:18:02.280245Z","shell.execute_reply":"2022-07-27T09:18:02.297056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Features in the dataset\n\n\n**survival**\tSurvival\t0 = No, 1 = Yes\n\n**pclass**\tTicket class\t1 = 1st, 2 = 2nd, 3 = 3rd\n            A proxy for socio-economic status (SES)\n            1st = Upper\n            2nd = Middle\n            3rd = Lower\n            \n**sex**  ->  Sex\n\n**Age**  ->  Age in years\n             Age is fractional if less than 1. If the age is estimated, is it in the form of xx.5\n            \n**sibsp** -> # of siblings / spouses aboard the Titanic\t\n             sibsp: The dataset defines family relations in this way...\n             Sibling = brother, sister, stepbrother, stepsister\n             Spouse = husband, wife (mistresses and fiancés were ignored)\n\n**parch** -> # of parents / children aboard the Titanic\t\n             parch: The dataset defines family relations in this way...\n             Parent = mother, father\n             Child = daughter, son, stepdaughter, stepson\n             Some children travelled only with a nanny, therefore parch=0 for them.\n\n**ticket**-> Ticket number\t\n\n**fare**  -> Passenger fare\t\n\n**cabin** -> Cabin number\t\n\n**embarked**-> Port of Embarkation\tC = Cherbourg, Q = Queenstown, S = Southampton\n","metadata":{}},{"cell_type":"markdown","source":"# Cleaning and Imputing data","metadata":{}},{"cell_type":"code","source":"def impute_data(dataframe):\n    #Filling in the na values of Embarked column as S as \n                                    #It is where most of the passengers boarded the ship\n    dataframe['Embarked'].fillna('S', inplace = True)\n    \n    #Filling in na values of Age column by mean of \n    dataframe.Age.fillna(train_df.Age.mean(), inplace = True)\n    dataframe.Fare.fillna(test_df.Fare.mean(), inplace=True)\n    return dataframe","metadata":{"execution":{"iopub.status.busy":"2022-07-27T09:18:02.463401Z","iopub.execute_input":"2022-07-27T09:18:02.463884Z","iopub.status.idle":"2022-07-27T09:18:02.471098Z","shell.execute_reply.started":"2022-07-27T09:18:02.463838Z","shell.execute_reply":"2022-07-27T09:18:02.470114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"impute_data(train_df).isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-27T09:18:02.643295Z","iopub.execute_input":"2022-07-27T09:18:02.643993Z","iopub.status.idle":"2022-07-27T09:18:02.657894Z","shell.execute_reply.started":"2022-07-27T09:18:02.643950Z","shell.execute_reply":"2022-07-27T09:18:02.657035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exploratory Data Analysis","metadata":{}},{"cell_type":"code","source":"eda_df = impute_data(train_df).fillna('Uncategorized')\npie_df = train_df[['Survived','PassengerId']].groupby(by = 'Survived').count()\nplt.pie(x=pie_df['PassengerId'], labels=['Not-Survived','Survived'],explode=[0.05,0.05], autopct='%.2f');","metadata":{"execution":{"iopub.status.busy":"2022-07-27T09:45:25.925153Z","iopub.execute_input":"2022-07-27T09:45:25.925675Z","iopub.status.idle":"2022-07-27T09:45:26.029858Z","shell.execute_reply.started":"2022-07-27T09:45:25.925636Z","shell.execute_reply":"2022-07-27T09:45:26.028426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems that less than 40% of people survived the clash.","metadata":{}},{"cell_type":"code","source":"pivot = train_df[['Pclass','Survived','Sex']].groupby(['Pclass','Sex']).mean().unstack()\nsns.heatmap(pivot, annot=True,cmap='cubehelix_r');","metadata":{"execution":{"iopub.status.busy":"2022-07-27T09:18:03.514109Z","iopub.execute_input":"2022-07-27T09:18:03.514993Z","iopub.status.idle":"2022-07-27T09:18:03.822132Z","shell.execute_reply.started":"2022-07-27T09:18:03.514921Z","shell.execute_reply":"2022-07-27T09:18:03.821096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The above heatmap denotes that the women had a higher chance of survival when compared to men during the disaster.","metadata":{}},{"cell_type":"code","source":"eda_df['Relatives'] = eda_df['SibSp']+eda_df['Parch']\neda_df[['PassengerId','Relatives','Survived']].groupby('Relatives').agg({'PassengerId':'count','Survived':'mean'})","metadata":{"execution":{"iopub.status.busy":"2022-07-27T09:18:03.823358Z","iopub.execute_input":"2022-07-27T09:18:03.823743Z","iopub.status.idle":"2022-07-27T09:18:03.843034Z","shell.execute_reply.started":"2022-07-27T09:18:03.823706Z","shell.execute_reply":"2022-07-27T09:18:03.841970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The above table represents that people who have lesser relatives on board the ship had more chance of survival than the people with more number of relatives.","metadata":{}},{"cell_type":"code","source":"eda_df.groupby(by = 'Embarked').mean()['Survived']","metadata":{"execution":{"iopub.status.busy":"2022-07-27T09:18:03.844917Z","iopub.execute_input":"2022-07-27T09:18:03.845296Z","iopub.status.idle":"2022-07-27T09:18:03.856980Z","shell.execute_reply.started":"2022-07-27T09:18:03.845263Z","shell.execute_reply":"2022-07-27T09:18:03.856203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The table shows that the people who embarked from Cherbourg had more chance of survival when compared to the others.","metadata":{}},{"cell_type":"markdown","source":"# Preparing Data to build the model","metadata":{}},{"cell_type":"code","source":"numerical_cols = ['PassengerId','Pclass','Age','SibSp','Parch','Fare']\ncategorical_cols = ['Embarked','Sex']\nencoded_cols = pd.get_dummies(train_df[categorical_cols]).columns\ndef encode_categorical(dataframe):\n    dataframe = impute_data(dataframe)\n    dataframe[encoded_cols] = pd.get_dummies(train_df[categorical_cols])\n    return dataframe\ntrain_df = encode_categorical(train_df)\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-27T09:18:03.858364Z","iopub.execute_input":"2022-07-27T09:18:03.858764Z","iopub.status.idle":"2022-07-27T09:18:03.895379Z","shell.execute_reply.started":"2022-07-27T09:18:03.858730Z","shell.execute_reply":"2022-07-27T09:18:03.894045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"input_cols = list(numerical_cols) + list(encoded_cols)\ninput_cols","metadata":{"execution":{"iopub.status.busy":"2022-07-27T09:20:04.197600Z","iopub.execute_input":"2022-07-27T09:20:04.198050Z","iopub.status.idle":"2022-07-27T09:20:04.206818Z","shell.execute_reply.started":"2022-07-27T09:20:04.198013Z","shell.execute_reply":"2022-07-27T09:20:04.205499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"inputs = train_df[input_cols]\ntarget = train_df['Survived']","metadata":{"execution":{"iopub.status.busy":"2022-07-27T09:18:04.184306Z","iopub.execute_input":"2022-07-27T09:18:04.185656Z","iopub.status.idle":"2022-07-27T09:18:04.192997Z","shell.execute_reply.started":"2022-07-27T09:18:04.185604Z","shell.execute_reply":"2022-07-27T09:18:04.191870Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Splitting the data into train data and test data","metadata":{}},{"cell_type":"code","source":"train_x, test_x, train_y, test_y = train_test_split(inputs,target, train_size = 0.33, random_state = 42)","metadata":{"execution":{"iopub.status.busy":"2022-07-27T09:18:04.440489Z","iopub.execute_input":"2022-07-27T09:18:04.441014Z","iopub.status.idle":"2022-07-27T09:18:04.450677Z","shell.execute_reply.started":"2022-07-27T09:18:04.440974Z","shell.execute_reply":"2022-07-27T09:18:04.449491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Building Machine Learning Models","metadata":{}},{"cell_type":"markdown","source":"### 1. Naive Bayes Classifier\n\nTrain Macro F1 ~ 0.80\n\nTest Macro F1 ~ 0.80\n\nUnknown Data Score ~ 0.60","metadata":{}},{"cell_type":"code","source":"bnb = BernoulliNB(alpha=0.05, fit_prior=False)\nbnb.fit(train_x, train_y)\nprint(classification_report(train_y, bnb.predict(train_x)))\nprint(\"------------------------------------------------------------\")\nprint(classification_report(test_y, bnb.predict(test_x)))","metadata":{"execution":{"iopub.status.busy":"2022-07-27T09:18:04.596298Z","iopub.execute_input":"2022-07-27T09:18:04.597577Z","iopub.status.idle":"2022-07-27T09:18:04.624784Z","shell.execute_reply.started":"2022-07-27T09:18:04.597346Z","shell.execute_reply":"2022-07-27T09:18:04.623616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\nThe score that I got for the model was 0.60 which is very less when compared to the accuracy that I got above.\n\nThis is the disadvantage of Naive Bayes since this model assumes that all the features are independent of each other but in reality they are not.\n\nI am going to build other models and tune the hyper parameters as well to achieve the maximum accuracy.","metadata":{}},{"cell_type":"markdown","source":"### 2. Building a Decision Tree Model\nTrain Macro F1 ~ 0.77\n\nTest Macro F1 ~ 0.77\n\nUnknown Data Score ~ 0.58","metadata":{}},{"cell_type":"code","source":"dt = DecisionTreeClassifier(random_state=42)\ndt.fit(train_x, train_y)\nprint(classification_report(train_y, dt.predict(train_x)))\nprint(\"------------------------------------------------------------\")\nprint(classification_report(test_y, dt.predict(test_x)))","metadata":{"execution":{"iopub.status.busy":"2022-07-27T09:18:04.664702Z","iopub.execute_input":"2022-07-27T09:18:04.665637Z","iopub.status.idle":"2022-07-27T09:18:04.692802Z","shell.execute_reply.started":"2022-07-27T09:18:04.665593Z","shell.execute_reply":"2022-07-27T09:18:04.691946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"roc_auc_score(test_y, dt.predict_proba(test_x)[:,1])","metadata":{"execution":{"iopub.status.busy":"2022-07-27T11:28:20.042116Z","iopub.execute_input":"2022-07-27T11:28:20.042586Z","iopub.status.idle":"2022-07-27T11:28:20.055283Z","shell.execute_reply.started":"2022-07-27T11:28:20.042549Z","shell.execute_reply":"2022-07-27T11:28:20.054276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This model is severely over fitted since train accuracy in almost 98 and the test accuracy is not even 80. To overcome overfitting, hyper parameters are tuned.","metadata":{}},{"cell_type":"markdown","source":"##### Tuning Hyperparameters of Decision Tree Classifier","metadata":{}},{"cell_type":"code","source":"def test_params_dt(**params):\n    model = DecisionTreeClassifier(random_state=42, **params).fit(train_x, train_y)\n    train_accuracy = accuracy_score(model.predict(train_x), train_y)\n    val_accuracy = accuracy_score(model.predict(test_x), test_y)\n    return train_accuracy, val_accuracy","metadata":{"execution":{"iopub.status.busy":"2022-07-27T11:06:49.997115Z","iopub.execute_input":"2022-07-27T11:06:49.997763Z","iopub.status.idle":"2022-07-27T11:06:50.005518Z","shell.execute_reply.started":"2022-07-27T11:06:49.997640Z","shell.execute_reply":"2022-07-27T11:06:50.003847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def test_param_and_plot_dt(param_name, param_values):\n    train_accuracies, test_accuracies = [], [] \n    for value in param_values:\n        params = {param_name: value}\n        train_accuracy, test_accuracy = test_params_dt(**params)\n        train_accuracies.append(train_accuracy)\n        test_accuracies.append(test_accuracy)\n    plt.figure(figsize=(10,6))\n    plt.title('Overfitting curve: ' + param_name)\n    plt.plot(param_values, train_accuracies, 'b-o')\n    plt.plot(param_values, test_accuracies, 'r-o')\n    plt.xlabel(param_name)\n    plt.ylabel('Accuracy')\n    plt.legend(['Training', 'Validation'])","metadata":{"execution":{"iopub.status.busy":"2022-07-27T11:07:56.691572Z","iopub.execute_input":"2022-07-27T11:07:56.691999Z","iopub.status.idle":"2022-07-27T11:07:56.700584Z","shell.execute_reply.started":"2022-07-27T11:07:56.691965Z","shell.execute_reply":"2022-07-27T11:07:56.699517Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_param_and_plot_dt('min_samples_leaf',[2,4,8,16,32,64])","metadata":{"execution":{"iopub.status.busy":"2022-07-27T11:09:43.329155Z","iopub.execute_input":"2022-07-27T11:09:43.329684Z","iopub.status.idle":"2022-07-27T11:09:43.644902Z","shell.execute_reply.started":"2022-07-27T11:09:43.329642Z","shell.execute_reply":"2022-07-27T11:09:43.643588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_param_and_plot_dt('min_weight_fraction_leaf',np.arange(0.01,0.5,0.05))","metadata":{"execution":{"iopub.status.busy":"2022-07-27T11:10:44.564225Z","iopub.execute_input":"2022-07-27T11:10:44.565300Z","iopub.status.idle":"2022-07-27T11:10:44.896235Z","shell.execute_reply.started":"2022-07-27T11:10:44.565243Z","shell.execute_reply":"2022-07-27T11:10:44.894943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_param_and_plot_dt('min_impurity_decrease',np.arange(0.001,0.015,0.001))","metadata":{"execution":{"iopub.status.busy":"2022-07-27T11:11:45.206261Z","iopub.execute_input":"2022-07-27T11:11:45.207052Z","iopub.status.idle":"2022-07-27T11:11:45.548771Z","shell.execute_reply.started":"2022-07-27T11:11:45.206981Z","shell.execute_reply":"2022-07-27T11:11:45.547316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"params = {\n    'min_impurity_decrease':[0.01,0.09,0.11],\n    'min_samples_leaf':np.arange(25,35),\n    'min_weight_fraction_leaf': np.arange(0.1,0.2,0.01)\n}\n\ndtrandomCV = RandomizedSearchCV(\n                estimator=DecisionTreeClassifier(random_state=42),\n                                param_distributions=params,\n                                 n_jobs = -1,\n                                scoring='roc_auc',\n                                error_score=np.nan,\n                                random_state=42)\n\ndtrandomCV.fit(train_x, train_y)\ndtrandomCV.best_params_","metadata":{"execution":{"iopub.status.busy":"2022-07-27T11:26:54.857138Z","iopub.execute_input":"2022-07-27T11:26:54.857591Z","iopub.status.idle":"2022-07-27T11:26:55.107369Z","shell.execute_reply.started":"2022-07-27T11:26:54.857554Z","shell.execute_reply":"2022-07-27T11:26:55.106611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"kfold = StratifiedKFold(n_splits=3,random_state=42, shuffle=True)\nfor train_index, test_index in kfold.split(train_df[input_cols],train_df['Survived']):\n    X_train, X_test = train_df.loc[train_index,:], train_df.loc[test_index,:]\n    y_train, y_test = train_df.loc[train_index,:]['Survived'], train_df.loc[test_index,:]['Survived']\n    model = dtrandomCV\n    model.fit(X_train[input_cols],y_train)\n    print(\"Train acc:\",accuracy_score(y_train,model.predict(X_train[input_cols])),\n                                     \"\\nTest acc:\",accuracy_score(y_test,model.predict(X_test[input_cols])),\n                                     \"\\n-------------------------------------------------\")","metadata":{"execution":{"iopub.status.busy":"2022-07-27T11:26:57.560719Z","iopub.execute_input":"2022-07-27T11:26:57.561174Z","iopub.status.idle":"2022-07-27T11:26:58.316559Z","shell.execute_reply.started":"2022-07-27T11:26:57.561137Z","shell.execute_reply":"2022-07-27T11:26:58.315780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"roc_auc_score(test_y, dtrandomCV.predict_proba(test_x)[:,1])","metadata":{"execution":{"iopub.status.busy":"2022-07-27T11:27:52.959049Z","iopub.execute_input":"2022-07-27T11:27:52.959860Z","iopub.status.idle":"2022-07-27T11:27:52.972538Z","shell.execute_reply.started":"2022-07-27T11:27:52.959817Z","shell.execute_reply":"2022-07-27T11:27:52.971241Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(classification_report(train_y, dtrandomCV.predict(train_x)))\nprint(\"------------------------------------------------------------\")\nprint(classification_report(test_y, dtrandomCV.predict(test_x)))\nsns.heatmap(confusion_matrix(test_y, dtrandomCV.predict(test_x)), annot=True);","metadata":{"execution":{"iopub.status.busy":"2022-07-27T11:27:04.505050Z","iopub.execute_input":"2022-07-27T11:27:04.506104Z","iopub.status.idle":"2022-07-27T11:27:04.769208Z","shell.execute_reply.started":"2022-07-27T11:27:04.506056Z","shell.execute_reply":"2022-07-27T11:27:04.767938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here is the convergence. Although the accuracies are converged, the False Negatives are too many.\n\nLet me train Random Forest Model for better performance.","metadata":{}},{"cell_type":"markdown","source":"### 3. Building Random Forest Classifier\n\nTrain Macro F1 ~ 0.77\n\nTest Macro F1 ~ 0.77\n\nUnknown Data Score ~ 0.787","metadata":{}},{"cell_type":"code","source":"rf = RandomForestClassifier(random_state=42)\n\nrf.fit(train_x, train_y)\nprint(classification_report(train_y, rf.predict(train_x)))\nprint(\"------------------------------------------------------------\")\nprint(classification_report(test_y, rf.predict(test_x)))","metadata":{"execution":{"iopub.status.busy":"2022-07-27T11:22:22.692680Z","iopub.execute_input":"2022-07-27T11:22:22.693107Z","iopub.status.idle":"2022-07-27T11:22:22.986345Z","shell.execute_reply.started":"2022-07-27T11:22:22.693066Z","shell.execute_reply":"2022-07-27T11:22:22.985060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"roc_auc_score(test_y, rf.predict_proba(test_x)[:,1])","metadata":{"execution":{"iopub.status.busy":"2022-07-27T10:22:17.905008Z","iopub.execute_input":"2022-07-27T10:22:17.905493Z","iopub.status.idle":"2022-07-27T10:22:17.939798Z","shell.execute_reply.started":"2022-07-27T10:22:17.905433Z","shell.execute_reply":"2022-07-27T10:22:17.938996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Tuning the Hyper Parameters","metadata":{}},{"cell_type":"code","source":"def test_params_rf(**params):\n    model = RandomForestClassifier(random_state=42, **params).fit(train_x, train_y)\n    train_accuracy = accuracy_score(model.predict(train_x), train_y)\n    val_accuracy = accuracy_score(model.predict(test_x), test_y)\n    return train_accuracy, val_accuracy","metadata":{"execution":{"iopub.status.busy":"2022-07-27T10:46:57.181203Z","iopub.execute_input":"2022-07-27T10:46:57.181697Z","iopub.status.idle":"2022-07-27T10:46:57.188586Z","shell.execute_reply.started":"2022-07-27T10:46:57.181657Z","shell.execute_reply":"2022-07-27T10:46:57.187363Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def test_param_and_plot_rf(param_name, param_values):\n    train_accuracies, test_accuracies = [],[]\n    for value in param_values:\n        params = {param_name: value}\n        train_accuracy, test_accuracy = test_params_dt(**params)\n        train_accuracies.append(train_accuracy)\n        test_accuracies.append(test_accuracy)\n    plt.figure(figsize=(10,6))\n    plt.title('Overfitting curve: ' + param_name)\n    plt.plot(param_values, train_accuracies, 'b-o')\n    plt.plot(param_values, test_accuracies, 'r-o')\n    plt.xlabel(param_name)\n    plt.ylabel('Accuracy')\n    plt.legend(['Training', 'Validation'])","metadata":{"execution":{"iopub.status.busy":"2022-07-27T10:49:39.722999Z","iopub.execute_input":"2022-07-27T10:49:39.723474Z","iopub.status.idle":"2022-07-27T10:49:39.731614Z","shell.execute_reply.started":"2022-07-27T10:49:39.723423Z","shell.execute_reply":"2022-07-27T10:49:39.730568Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_param_and_plot_rf('min_samples_leaf',[1,3,8,12,15,20,25,30,40])","metadata":{"execution":{"iopub.status.busy":"2022-07-27T10:49:52.837183Z","iopub.execute_input":"2022-07-27T10:49:52.837663Z","iopub.status.idle":"2022-07-27T10:49:53.130180Z","shell.execute_reply.started":"2022-07-27T10:49:52.837625Z","shell.execute_reply":"2022-07-27T10:49:53.129138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_param_and_plot_rf('min_impurity_decrease', np.arange(0.001,0.1,0.005))","metadata":{"execution":{"iopub.status.busy":"2022-07-27T10:51:06.008509Z","iopub.execute_input":"2022-07-27T10:51:06.008974Z","iopub.status.idle":"2022-07-27T10:51:06.365797Z","shell.execute_reply.started":"2022-07-27T10:51:06.008938Z","shell.execute_reply":"2022-07-27T10:51:06.364916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_param_and_plot_rf('min_samples_split', np.arange(40,60))","metadata":{"execution":{"iopub.status.busy":"2022-07-27T10:55:12.179500Z","iopub.execute_input":"2022-07-27T10:55:12.179955Z","iopub.status.idle":"2022-07-27T10:55:12.658347Z","shell.execute_reply.started":"2022-07-27T10:55:12.179912Z","shell.execute_reply":"2022-07-27T10:55:12.656953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"params = {\n    'min_impurity_decrease':[0.02,0.03,0.04],\n    'min_samples_leaf':[25,30],\n    'class_weight':['balanced','balanced_subsample'],\n    'min_samples_split':[51,52,55,56],\n    'max_samples': np.arange(0.1,1,0.1)\n}\n\nrfrandomCV = RandomizedSearchCV(\n                estimator=RandomForestClassifier(n_jobs=-1,random_state=42),\n                                param_distributions=params,\n                                 n_jobs = -1,\n                                scoring='roc_auc',\n                                error_score=np.nan)\n\nrfrandomCV.fit(train_x, train_y)\nrfrandomCV.best_params_","metadata":{"execution":{"iopub.status.busy":"2022-07-27T11:28:52.041603Z","iopub.execute_input":"2022-07-27T11:28:52.044415Z","iopub.status.idle":"2022-07-27T11:28:59.558032Z","shell.execute_reply.started":"2022-07-27T11:28:52.044334Z","shell.execute_reply":"2022-07-27T11:28:59.556523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(classification_report(train_y, rfrandomCV.predict(train_x)))\nprint(\"------------------------------------------------------------\")\nprint(classification_report(test_y, rfrandomCV.predict(test_x)))\nsns.heatmap(confusion_matrix(test_y, rfrandomCV.predict(test_x)), annot=True);","metadata":{"execution":{"iopub.status.busy":"2022-07-27T11:28:59.560691Z","iopub.execute_input":"2022-07-27T11:28:59.561322Z","iopub.status.idle":"2022-07-27T11:29:00.158900Z","shell.execute_reply.started":"2022-07-27T11:28:59.561267Z","shell.execute_reply":"2022-07-27T11:29:00.157533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Validating the Model","metadata":{}},{"cell_type":"code","source":"kfold = StratifiedKFold(n_splits=3,random_state=42, shuffle=True)\nfor train_index, test_index in kfold.split(train_df[input_cols],train_df['Survived']):\n    X_train, X_test = train_df.loc[train_index,:], train_df.loc[test_index,:]\n    y_train, y_test = train_df.loc[train_index,:]['Survived'], train_df.loc[test_index,:]['Survived']\n    model = rfrandomCV\n    model.fit(X_train[input_cols],y_train)\n    print(\"Train acc:\",accuracy_score(y_train,model.predict(X_train[input_cols])),\n                                     \"\\nTest acc:\",accuracy_score(y_test,model.predict(X_test[input_cols])),\n                                     \"\\n-------------------------------------------------\")","metadata":{"execution":{"iopub.status.busy":"2022-07-27T11:29:03.067479Z","iopub.execute_input":"2022-07-27T11:29:03.068413Z","iopub.status.idle":"2022-07-27T11:29:25.560878Z","shell.execute_reply.started":"2022-07-27T11:29:03.068368Z","shell.execute_reply":"2022-07-27T11:29:25.559586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The model has finally converged and the accuracy is maintained approximately at 78% in any step during k-fold cross validation.","metadata":{}},{"cell_type":"code","source":"roc_auc_score(test_y, rfrandomCV.predict_proba(test_x)[:,1])","metadata":{"execution":{"iopub.status.busy":"2022-07-27T11:29:25.562678Z","iopub.execute_input":"2022-07-27T11:29:25.563064Z","iopub.status.idle":"2022-07-27T11:29:25.678920Z","shell.execute_reply.started":"2022-07-27T11:29:25.563030Z","shell.execute_reply":"2022-07-27T11:29:25.677675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model Predictions for test Data","metadata":{}},{"cell_type":"code","source":"validation = test_df.copy()","metadata":{"execution":{"iopub.status.busy":"2022-07-27T09:25:33.666711Z","iopub.execute_input":"2022-07-27T09:25:33.667573Z","iopub.status.idle":"2022-07-27T09:25:33.674675Z","shell.execute_reply.started":"2022-07-27T09:25:33.667525Z","shell.execute_reply":"2022-07-27T09:25:33.673167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"input = encode_categorical(validation)[input_cols].copy()\n\npredict_test = rfrandomCV.predict(input)","metadata":{"execution":{"iopub.status.busy":"2022-07-27T09:25:43.461011Z","iopub.execute_input":"2022-07-27T09:25:43.461889Z","iopub.status.idle":"2022-07-27T09:25:43.585363Z","shell.execute_reply.started":"2022-07-27T09:25:43.461827Z","shell.execute_reply":"2022-07-27T09:25:43.584213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"output = pd.DataFrame({'PassengerId': test_df.PassengerId, 'Survived': predict_test})\noutput.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-27T09:25:46.500679Z","iopub.execute_input":"2022-07-27T09:25:46.501168Z","iopub.status.idle":"2022-07-27T09:25:46.514283Z","shell.execute_reply.started":"2022-07-27T09:25:46.501124Z","shell.execute_reply":"2022-07-27T09:25:46.512922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"output.to_csv('submission.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-27T09:25:51.212686Z","iopub.execute_input":"2022-07-27T09:25:51.213121Z","iopub.status.idle":"2022-07-27T09:25:51.224327Z","shell.execute_reply.started":"2022-07-27T09:25:51.213088Z","shell.execute_reply":"2022-07-27T09:25:51.223273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# References\n1. [How to Use ROC Curves and Precision-Recall Curves for Classification in Python](https://machinelearningmastery.com/roc-curves-and-precision-recall-curves-for-classification-in-python/#:~:text=The%20AUC%20for%20the%20ROC,skill%20and%20perfect%20skill%20respectively.)\n\n2. [A Beginner’s Guide to Random Forest Hyperparameter Tuning](https://www.analyticsvidhya.com/blog/2020/03/beginners-guide-random-forest-hyperparameter-tuning/#:~:text=This%20Random%20Forest%20hyperparameter%20specifies,node%20after%20splitting%20a%20node.&text=The%20tree%20on%20the%20left%20represents%20an%20unconstrained%20tree.,a%20minimum%20of%205%20samples.)\n\n3. [A Comprehensive Guide to Decision trees](https://www.analyticsvidhya.com/blog/2021/07/a-comprehensive-guide-to-decision-trees/)\n\n4. [RandomizedSearchCV - scikit-learn](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.RandomizedSearchCV.html)\n\n","metadata":{}},{"cell_type":"markdown","source":"# Future Works\n\n1. I would like to train a Deep Learning Model to predict the survivors better.\n\n2. I would try to improve the predictions through linear models as well.","metadata":{}}]}