{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-05T15:26:19.269013Z","iopub.execute_input":"2022-07-05T15:26:19.269545Z","iopub.status.idle":"2022-07-05T15:26:19.284359Z","shell.execute_reply.started":"2022-07-05T15:26:19.269483Z","shell.execute_reply":"2022-07-05T15:26:19.283085Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data preparation\n\nWe start by loading the data into a `pandas.DataFrame` with the `read_csv()` function, and then displaying the first five rows with the `.head()` method.","metadata":{}},{"cell_type":"code","source":"train_data = pd.read_csv(\"/kaggle/input/titanic/train.csv\")\ntrain_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-05T15:26:19.287135Z","iopub.execute_input":"2022-07-05T15:26:19.289553Z","iopub.status.idle":"2022-07-05T15:26:19.319306Z","shell.execute_reply.started":"2022-07-05T15:26:19.289512Z","shell.execute_reply":"2022-07-05T15:26:19.318622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We then seperate the target from this `DataFrame`, that being the binary variable `Survived` that indiactes if the passenger survuved or not. This will be used to train our classification model later.","metadata":{}},{"cell_type":"code","source":"target = train_data['Survived']","metadata":{"execution":{"iopub.status.busy":"2022-07-05T15:26:19.320616Z","iopub.execute_input":"2022-07-05T15:26:19.321149Z","iopub.status.idle":"2022-07-05T15:26:19.325400Z","shell.execute_reply.started":"2022-07-05T15:26:19.321117Z","shell.execute_reply":"2022-07-05T15:26:19.324662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We then split features that should be informative for our target from the main dataset. Here we use `Pclass` (class of ticket), `Sex`, `Age`, `SibSp` (number of siblings or spouses aboard), `Parch` (number of parents or children aboard), `Fare`, and `Embarked`.","metadata":{}},{"cell_type":"code","source":"train_data_cut = train_data[['Pclass','Sex','Age','SibSp','Parch', 'Fare', 'Embarked']].copy()\ntrain_data_cut.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-05T15:26:19.326881Z","iopub.execute_input":"2022-07-05T15:26:19.327208Z","iopub.status.idle":"2022-07-05T15:26:19.350255Z","shell.execute_reply.started":"2022-07-05T15:26:19.327180Z","shell.execute_reply":"2022-07-05T15:26:19.349556Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Both `Sex` and `Embarked` are catagorical features so need to encoded if we are going to use them to do any model training. We can do this with the `.get_dummies()` function.","metadata":{}},{"cell_type":"code","source":"X = pd.get_dummies(train_data_cut)\nX.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-05T15:26:19.352747Z","iopub.execute_input":"2022-07-05T15:26:19.353589Z","iopub.status.idle":"2022-07-05T15:26:19.374321Z","shell.execute_reply.started":"2022-07-05T15:26:19.353556Z","shell.execute_reply":"2022-07-05T15:26:19.373282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once we have done this encoding we can calculate and plot the Pearson correlation coefficient and verify that most of our features show some correlation with our target.","metadata":{}},{"cell_type":"code","source":"X.join(target).corr()['Survived'].iloc[:-1].plot.bar()","metadata":{"execution":{"iopub.status.busy":"2022-07-05T15:26:19.377835Z","iopub.execute_input":"2022-07-05T15:26:19.378596Z","iopub.status.idle":"2022-07-05T15:26:19.603916Z","shell.execute_reply.started":"2022-07-05T15:26:19.378551Z","shell.execute_reply":"2022-07-05T15:26:19.603207Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We next use the `.info()` method to examine how many rows of the dataset have missing values.","metadata":{}},{"cell_type":"code","source":"X.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-05T15:26:19.604973Z","iopub.execute_input":"2022-07-05T15:26:19.605901Z","iopub.status.idle":"2022-07-05T15:26:19.622083Z","shell.execute_reply.started":"2022-07-05T15:26:19.605867Z","shell.execute_reply":"2022-07-05T15:26:19.620744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see from the output above that there are a signifcant number of missing `Age` values. There are two options for dealing with this: 1) Don't use age as an input feature, 2) impute the missing values. We will do both to compare the impact of prediction accuracy. \n\nIt should be noted that there are two passengers in the training set with `nan` in the `Embarked` column in the training set. For these cases using `get_dummies()` simply sets 0 for `Embarked_C`, `Embarked_Q`, and `Embarked_S`. This is why it appears there are 0 non-null entries for these features. As it impacts such a small portion of the training set (and doesn't impact the test set at all) we accept this.","metadata":{}},{"cell_type":"code","source":"from sklearn.impute import SimpleImputer, KNNImputer\n\nX_noAge = X[['Pclass','SibSp','Parch', 'Sex_female', 'Sex_male']]\n\nimp_mean = SimpleImputer(missing_values=np.nan, strategy='mean')\nX_ImpMean = imp_mean.fit_transform(X)\n\nimp_KNN = KNNImputer(n_neighbors=5)\nX_ImpKNN = imp_KNN.fit_transform(X)","metadata":{"execution":{"iopub.status.busy":"2022-07-05T15:26:19.623589Z","iopub.execute_input":"2022-07-05T15:26:19.623936Z","iopub.status.idle":"2022-07-05T15:26:19.662557Z","shell.execute_reply.started":"2022-07-05T15:26:19.623905Z","shell.execute_reply":"2022-07-05T15:26:19.661378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we are ready to train our classifier!","metadata":{}},{"cell_type":"markdown","source":"# Train classifier\n\nWe will be training a random forest classifier (RFC); we start by training a baseline model with the same hyperparameters as in [Alexis Cook's tutorial](https://www.kaggle.com/code/alexisbcook/titanic-tutorial/notebook).","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier\n\nmodel = RandomForestClassifier(n_estimators=100, max_depth=5, random_state=1)\nmodel.fit(X_noAge, target);","metadata":{"execution":{"iopub.status.busy":"2022-07-05T15:26:19.664846Z","iopub.execute_input":"2022-07-05T15:26:19.665729Z","iopub.status.idle":"2022-07-05T15:26:19.904363Z","shell.execute_reply.started":"2022-07-05T15:26:19.665674Z","shell.execute_reply":"2022-07-05T15:26:19.903674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We evaluate the performance of this RFC by making predictions on the training set, and from these predictions computing the precision, recall, and their harmonic mean the $F_1$-score.","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import recall_score, precision_score, f1_score, accuracy_score\n\npreds_train = model.predict(X_noAge)\nprecision_train = precision_score(target, preds_train)\nrecall_train = recall_score(target, preds_train)\nf1_train = f1_score(target, preds_train)\naccuracy_train = accuracy_score(target, preds_train)\n\nprint(\"Precision (train): \", np.round(precision_train, decimals=2))\nprint(\"Recall (train): \", np.round(recall_train, decimals=2))\nprint(\"f1 (train): \", np.round(f1_train, decimals=2))\nprint(\"accuracy (train): \", np.round(accuracy_train, decimals=2))","metadata":{"execution":{"iopub.status.busy":"2022-07-05T15:26:19.905629Z","iopub.execute_input":"2022-07-05T15:26:19.906361Z","iopub.status.idle":"2022-07-05T15:26:19.944339Z","shell.execute_reply.started":"2022-07-05T15:26:19.906322Z","shell.execute_reply":"2022-07-05T15:26:19.943371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that these scores are already pretty good! But we can try tuning the hyperparameters of the RFC to see if we can improve these scores.\n\nWe will be doing grid search cross-validation to do the tuning. We will define the extent of a 2D grid of possible values for the maximum depth of each tree in the forest (`max_depth`) and the minimum number of samples required to create a split in each tree (`min_samples_split`).","metadata":{}},{"cell_type":"code","source":"param_dict = {'max_depth':np.arange(2,12,step=2), 'min_samples_split':np.arange(2,12,step=2)}","metadata":{"execution":{"iopub.status.busy":"2022-07-05T15:26:19.945632Z","iopub.execute_input":"2022-07-05T15:26:19.945965Z","iopub.status.idle":"2022-07-05T15:26:19.951483Z","shell.execute_reply.started":"2022-07-05T15:26:19.945936Z","shell.execute_reply":"2022-07-05T15:26:19.950616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we do the tuning with the tuning with the `GridSearchCV` class. The parameters will be tuned based on the $F_1$-score, this is the harmonic mean of the precision and recall.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV\n\n\ncv = GridSearchCV(RandomForestClassifier(random_state=1), param_dict, scoring='f1')\ncv.fit(X_noAge, target);\n\nbest_params = cv.best_params_\nbest_params","metadata":{"execution":{"iopub.status.busy":"2022-07-05T15:26:19.954879Z","iopub.execute_input":"2022-07-05T15:26:19.955535Z","iopub.status.idle":"2022-07-05T15:26:45.143690Z","shell.execute_reply.started":"2022-07-05T15:26:19.955450Z","shell.execute_reply":"2022-07-05T15:26:45.142657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We then train the RFC with the tuned parameters and the full dataset, and then compute the precision and recall again.","metadata":{}},{"cell_type":"code","source":"model = RandomForestClassifier(max_depth=best_params['max_depth'],\n                               min_samples_split=best_params['min_samples_split'],\n                               random_state=1)\nmodel.fit(X_noAge, target)\n\npreds_train = model.predict(X_noAge)\nprecision_train = precision_score(target, preds_train)\nrecall_train = recall_score(target, preds_train)\nf1_train = f1_score(target, preds_train)\naccuracy_train = accuracy_score(target, preds_train)\n\nprint(\"Precision (train): \", np.round(precision_train, decimals=2))\nprint(\"Recall (train): \", np.round(recall_train, decimals=2))\nprint(\"f1 (train): \", np.round(f1_train, decimals=2))\nprint(\"accuracy (train): \", np.round(accuracy_train, decimals=2))","metadata":{"execution":{"iopub.status.busy":"2022-07-05T15:26:45.145159Z","iopub.execute_input":"2022-07-05T15:26:45.145570Z","iopub.status.idle":"2022-07-05T15:26:45.368590Z","shell.execute_reply.started":"2022-07-05T15:26:45.145538Z","shell.execute_reply":"2022-07-05T15:26:45.367583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cv = GridSearchCV(RandomForestClassifier(random_state=1), param_dict, scoring='f1')\ncv.fit(X_ImpKNN, target);\n\nbest_params = cv.best_params_\nbest_params\n\nmodel = RandomForestClassifier(max_depth=best_params['max_depth'],\n                               min_samples_split=best_params['min_samples_split'],\n                               random_state=1)\nmodel.fit(X_ImpKNN, target)\n\npreds_train = model.predict(X_ImpKNN)\nprecision_train = precision_score(target, preds_train)\nrecall_train = recall_score(target, preds_train)\nf1_train = f1_score(target, preds_train)\naccuracy_train = accuracy_score(target, preds_train)\n\nprint(\"Precision (train): \", np.round(precision_train, decimals=2))\nprint(\"Recall (train): \", np.round(recall_train, decimals=2))\nprint(\"f1 (train): \", np.round(f1_train, decimals=2))\nprint(\"accuracy (train): \", np.round(accuracy_train, decimals=2))","metadata":{"execution":{"iopub.status.busy":"2022-07-05T15:26:45.369996Z","iopub.execute_input":"2022-07-05T15:26:45.371076Z","iopub.status.idle":"2022-07-05T15:27:12.011542Z","shell.execute_reply.started":"2022-07-05T15:26:45.371027Z","shell.execute_reply":"2022-07-05T15:27:12.010515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cv = GridSearchCV(RandomForestClassifier(random_state=1), param_dict, scoring='f1')\ncv.fit(X_ImpMean, target);\n\nbest_params = cv.best_params_\nbest_params\n\nmodel = RandomForestClassifier(max_depth=best_params['max_depth'],\n                               min_samples_split=best_params['min_samples_split'],\n                               random_state=1)\nmodel.fit(X_ImpMean, target)\n\npreds_train = model.predict(X_ImpMean)\nprecision_train = precision_score(target, preds_train)\nrecall_train = recall_score(target, preds_train)\nf1_train = f1_score(target, preds_train)\naccuracy_train = accuracy_score(target, preds_train)\n\nprint(\"Precision (train): \", np.round(precision_train, decimals=2))\nprint(\"Recall (train): \", np.round(recall_train, decimals=2))\nprint(\"f1 (train): \", np.round(f1_train, decimals=2))\nprint(\"accuracy (train): \", np.round(accuracy_train, decimals=2))","metadata":{"execution":{"iopub.status.busy":"2022-07-05T15:27:12.013068Z","iopub.execute_input":"2022-07-05T15:27:12.013405Z","iopub.status.idle":"2022-07-05T15:27:38.451530Z","shell.execute_reply.started":"2022-07-05T15:27:12.013375Z","shell.execute_reply":"2022-07-05T15:27:38.450554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"proba_train = model.predict_proba(X_ImpMean)[:,1]","metadata":{"execution":{"iopub.status.busy":"2022-07-05T15:27:38.453007Z","iopub.execute_input":"2022-07-05T15:27:38.454137Z","iopub.status.idle":"2022-07-05T15:27:38.486013Z","shell.execute_reply.started":"2022-07-05T15:27:38.454088Z","shell.execute_reply":"2022-07-05T15:27:38.484945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We then evaluate the precision and recall for a series of thresholds.","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import precision_recall_curve\n\nprecision, recall, thresholds = precision_recall_curve(target, proba_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-05T15:27:38.487057Z","iopub.execute_input":"2022-07-05T15:27:38.487367Z","iopub.status.idle":"2022-07-05T15:27:38.493698Z","shell.execute_reply.started":"2022-07-05T15:27:38.487339Z","shell.execute_reply":"2022-07-05T15:27:38.492842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To decide which threshold is best, we find that that maximises the $F_1$-score.","metadata":{}},{"cell_type":"code","source":"f1 = (2*precision*recall)/(precision+recall)\n\nbest_thresh = thresholds[np.where(f1 == f1.max())][0]\nbest_thresh","metadata":{"execution":{"iopub.status.busy":"2022-07-05T15:27:38.494939Z","iopub.execute_input":"2022-07-05T15:27:38.495401Z","iopub.status.idle":"2022-07-05T15:27:38.508190Z","shell.execute_reply.started":"2022-07-05T15:27:38.495365Z","shell.execute_reply":"2022-07-05T15:27:38.507228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preds_train_thresh = proba_train >= best_thresh\n\nprecision_train = precision_score(target, preds_train_thresh)\nrecall_train = recall_score(target, preds_train_thresh)\nf1_train = f1_score(target, preds_train_thresh)\naccuracy_train = accuracy_score(target, preds_train_thresh)\n\nprint(\"Precision (train): \", np.round(precision_train, decimals=2))\nprint(\"Recall (train): \", np.round(recall_train, decimals=2))\nprint(\"f1 (train): \", np.round(f1_train, decimals=2))\nprint(\"accuracy (train): \", np.round(accuracy_train, decimals=2))","metadata":{"execution":{"iopub.status.busy":"2022-07-05T15:27:38.509464Z","iopub.execute_input":"2022-07-05T15:27:38.509943Z","iopub.status.idle":"2022-07-05T15:27:38.527816Z","shell.execute_reply.started":"2022-07-05T15:27:38.509900Z","shell.execute_reply":"2022-07-05T15:27:38.526745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data = pd.read_csv(\"/kaggle/input/titanic/test.csv\")\ntest_data_cut = test_data[['Pclass','Sex','Age','SibSp','Parch', 'Fare', 'Embarked']].copy()\n\nX_test = pd.get_dummies(test_data_cut)\nX_test_ImpMean = imp_mean.transform(X_test)\n\noutput = pd.DataFrame({'PassengerId': test_data.PassengerId,\n                       'Survived': (model.predict_proba(X_test_ImpMean)[:,1]>= best_thresh).astype(int)})\noutput.to_csv('submission.csv', index=False)\nprint(\"Your submission was successfully saved!\")","metadata":{"execution":{"iopub.status.busy":"2022-07-05T15:27:38.529075Z","iopub.execute_input":"2022-07-05T15:27:38.529793Z","iopub.status.idle":"2022-07-05T15:27:38.574577Z","shell.execute_reply.started":"2022-07-05T15:27:38.529761Z","shell.execute_reply":"2022-07-05T15:27:38.573835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}