{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Introduction**\n\nThis kernel was inspired in part by the work of [SarahG](https://www.kaggle.com/sgus1318/titanic-analysis-learning-to-swim-with-python)'s analysis that I thank very much for the quality of her analysis. This work represents a deeper analysis by playing on several parameters while using only logistic regression estimator. In a future work, I will discuss other techniques. I am open to any criticism and proposal. You do not hesitate to evaluate this analysis.\nThe following kernel contains the steps enumerated below for assessing the Titanic survival dataset:\n\n1. [Import data and python packages](#t1.)\n2. [Assess Data Quality & Missing Values](#t2.)\n    * 2.1. [Age - Missing Values](#t2.1.)\n    * 2.2. [Cabin - Missing Values](#t2.2.)\n    * 2.3. [Embarked - Missing Values](#t2.3.)\n    * 2.4. [Final Adjustments to Data](#t2.4.)\n        * 2.4.1 [Additional Variables](#t2.4.1.)\n3. [Exploratory Data Analysis](#t3.)\n4. [Logistic Regression and Results](#t4.)\n    * 4.1. [Feature selection](#t4.1.)\n        * 4.1.1. [Recursive feature elimination](#t4.1.1.)\n        * 4.1.2. [Feature ranking with recursive feature elimination and cross-validation](#t4.1.2.)\n    * 4.2. [Review of model evaluation procedures](#t4.2.)\n        * 4.2.1. [Model evaluation based on simple train/test split using `train_test_split()`](#t4.2.1.)\n        * 4.2.2. [Model evaluation based on K-fold cross-validation using `cross_val_score()`](#t4.2.2.)\n        * 4.2.3. [Model evaluation based on K-fold cross-validation using `cross_validate()`](#t4.2.3.)\n    * 4.3. [GridSearchCV evaluating using multiple scorers simultaneously](#t4.3.)\n    * 4.4. [GridSearchCV evaluating using multiple scorers, RepeatedStratifiedKFold and pipeline for preprocessing simultaneously](#t4.4.)","metadata":{"_cell_guid":"cb19f71d-51c8-417f-829e-3179d3319dcd","_uuid":"c3b9226e142667d6b96e34daf7d6e42bea0ea1e2"}},{"cell_type":"markdown","source":"<a id=\"t1.\"></a>\n# 1. Import Data & Python Packages","metadata":{"_cell_guid":"33c91cae-2ff8-45a6-b8cb-671619e9c933","_uuid":"0a395fd25f20834b070ef55cb8987c8c1f9b55f9"}},{"cell_type":"markdown","source":"VARIABLE DESCRIPTIONS\nPclass Passenger Class (1 = 1st; 2 = 2nd; 3 = 3rd)\nsurvival Survival (0 = No; 1 = Yes)\nname Name\nsex Sex\nage Age\nsibsp Number of Siblings/Spouses Aboard\nparch Number of Parents/Children Aboard\nticket Ticket Number\nfare Passenger Fare (British pound)\ncabin Cabin\nembarked Port of Embarkation (C = Cherbourg; Q = Queenstown; S = Southampton)\nboat Lifeboat\nbody Body Identification Number\nhome.dest Home/Destination","metadata":{}},{"cell_type":"markdown","source":"Pclass Passenger Class (1 = 1st; 2 = 2nd; 3 = 3rd)\nsurvival Survival (0 = No; 1 = Yes)\nname Name\nsex Sex\nage Age\nsibsp Number of Siblings/Spouses Aboard\nparch Number of Parents/Children Aboard\nticket Ticket Number\nfare Passenger Fare (British pound)\ncabin Cabin\nembarked Port of Embarkation (C = Cherbourg; Q = Queenstown; S = Southampton)\nboat Lifeboat\nbody Body Identification Number\nhome.dest Home/Destination","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \n\nfrom sklearn import preprocessing\nimport matplotlib.pyplot as plt \nplt.rc(\"font\", size=14)\nimport seaborn as sns\nsns.set(style=\"white\") #white background style for seaborn plots\nsns.set(style=\"whitegrid\", color_codes=True)\n\nimport warnings\nwarnings.simplefilter(action='ignore')","metadata":{"_cell_guid":"de05512e-6991-44df-9599-da92a7e459ac","_uuid":"d8bdd5f0320e244e4702ed8ec1c2482b022c51cd","execution":{"iopub.status.busy":"2022-07-07T13:19:28.162116Z","iopub.execute_input":"2022-07-07T13:19:28.162781Z","iopub.status.idle":"2022-07-07T13:19:29.444607Z","shell.execute_reply.started":"2022-07-07T13:19:28.162714Z","shell.execute_reply":"2022-07-07T13:19:29.443816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Read CSV train data file into DataFrame\ntrain_df = pd.read_csv(\"../input/train.csv\")\n\n# Read CSV test data file into DataFrame\ntest_df = pd.read_csv(\"../input/test.csv\")\n\n# preview train data\ntrain_df.head()","metadata":{"_cell_guid":"e0a17223-f682-45fc-89a5-667af9782bbe","_uuid":"7964157913fbcff581fc1929eed487708e81ac9c","execution":{"iopub.status.busy":"2022-07-07T13:19:29.445802Z","iopub.execute_input":"2022-07-07T13:19:29.446304Z","iopub.status.idle":"2022-07-07T13:19:29.521795Z","shell.execute_reply.started":"2022-07-07T13:19:29.446249Z","shell.execute_reply":"2022-07-07T13:19:29.520912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:19:29.523131Z","iopub.execute_input":"2022-07-07T13:19:29.523698Z","iopub.status.idle":"2022-07-07T13:19:29.531408Z","shell.execute_reply.started":"2022-07-07T13:19:29.523416Z","shell.execute_reply":"2022-07-07T13:19:29.530313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('The number of samples into the train data is {}.'.format(train_df.shape[0]))","metadata":{"_cell_guid":"872d0de9-a873-4b60-b1ee-d557ee39d8a1","_uuid":"d38222a64d4dfd1d1ee1a7ee1f58c4aa54560de3","execution":{"iopub.status.busy":"2022-07-07T13:19:29.732173Z","iopub.execute_input":"2022-07-07T13:19:29.732505Z","iopub.status.idle":"2022-07-07T13:19:29.737593Z","shell.execute_reply.started":"2022-07-07T13:19:29.732449Z","shell.execute_reply":"2022-07-07T13:19:29.736833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# preview test data\ntest_df.head()","metadata":{"_cell_guid":"1d969b76-ea88-4d32-a58e-f22a070258bf","_uuid":"bff38fcf31baf67493513c06f0c2f6e50576ff09","execution":{"iopub.status.busy":"2022-07-07T13:19:30.155577Z","iopub.execute_input":"2022-07-07T13:19:30.155853Z","iopub.status.idle":"2022-07-07T13:19:30.186473Z","shell.execute_reply.started":"2022-07-07T13:19:30.155817Z","shell.execute_reply":"2022-07-07T13:19:30.185687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('The number of samples into the test data is {}.'.format(test_df.shape[0]))","metadata":{"_cell_guid":"254dd074-e07e-49f2-9184-80046b10b481","_uuid":"62de7ddd73fed8d88ccbe1ba79e59b8e596cbb13","execution":{"iopub.status.busy":"2022-07-07T13:19:30.542255Z","iopub.execute_input":"2022-07-07T13:19:30.542570Z","iopub.status.idle":"2022-07-07T13:19:30.549333Z","shell.execute_reply.started":"2022-07-07T13:19:30.542513Z","shell.execute_reply":"2022-07-07T13:19:30.548397Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font color=red>  Note: there is no target variable into test data (i.e. \"Survival\" column is missing), so the goal is to predict this target using different machine learning algorithms such as logistic regression. </font>","metadata":{"_cell_guid":"4cd08f1e-9cb9-4d8e-99d0-3a9ce91b91c3","_uuid":"2b1d45128663b9466fe9ac0059a13cfd4bd43657"}},{"cell_type":"markdown","source":"<a id=\"t2.\"></a>\n# 2. Data Quality & Missing Value Assessment","metadata":{"_cell_guid":"6578c0da-7bcf-433d-9f28-a66d8dfa6fa3","_uuid":"8660e63a62c2fcdb4f7633380166438caf5edae9"}},{"cell_type":"code","source":"# check missing values in train data\ntrain_df.isnull().sum()","metadata":{"_cell_guid":"29dddd33-d995-4b0f-92ea-a361b368cc42","_uuid":"d4fe22ead7e187724ca6f3ba7ba0e6412ae0e874","execution":{"iopub.status.busy":"2022-07-07T13:19:31.657339Z","iopub.execute_input":"2022-07-07T13:19:31.658022Z","iopub.status.idle":"2022-07-07T13:19:31.666066Z","shell.execute_reply.started":"2022-07-07T13:19:31.657971Z","shell.execute_reply":"2022-07-07T13:19:31.665455Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[['Age','Fare']].isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:19:31.994304Z","iopub.execute_input":"2022-07-07T13:19:31.994799Z","iopub.status.idle":"2022-07-07T13:19:32.002459Z","shell.execute_reply.started":"2022-07-07T13:19:31.994731Z","shell.execute_reply":"2022-07-07T13:19:32.001713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"t2.1.\"></a>\n## 2.1.    Age - Missing Values","metadata":{"_cell_guid":"7776faeb-6a8f-4460-a367-4b087d2cc089","_uuid":"696b428bd3ca49421f650665267ce7ca1b358814"}},{"cell_type":"code","source":"# percent of missing \"Age\" \nprint('Percent of missing \"Age\" records is %.2f%%' %((train_df['Age'].isnull().sum()/train_df.shape[0])*100))","metadata":{"_cell_guid":"d4ee6559-6d0c-409d-9dca-1d105a4ccd8a","_uuid":"129cf984d05d9ce97c54548145e65f9e4b9b0c37","execution":{"iopub.status.busy":"2022-07-07T13:19:32.801405Z","iopub.execute_input":"2022-07-07T13:19:32.801878Z","iopub.status.idle":"2022-07-07T13:19:32.809846Z","shell.execute_reply.started":"2022-07-07T13:19:32.801815Z","shell.execute_reply":"2022-07-07T13:19:32.808694Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"~20% of entries for passenger age are missing. Let's see what the 'Age' variable looks like in general.","metadata":{"_cell_guid":"951f7bb8-779c-4eac-85a2-3fdcfdcd293e","_uuid":"c8fff460fb532a063f6944450809014ca831ca52"}},{"cell_type":"code","source":"ax = train_df[\"Age\"].hist(bins=15, density=True, stacked=True, color='teal', alpha=0.6)\ntrain_df[\"Age\"].plot(kind='density', color='teal')\nax.set(xlabel='Age')\nplt.xlim(-10,85)\nplt.show()","metadata":{"_cell_guid":"6d65fcfa-52bf-45ab-b959-64a32c1c1976","_uuid":"c6fd60f15d5e803d4dffc89e782c6fbc72445a83","execution":{"iopub.status.busy":"2022-07-07T13:19:33.487588Z","iopub.execute_input":"2022-07-07T13:19:33.488254Z","iopub.status.idle":"2022-07-07T13:19:33.866738Z","shell.execute_reply.started":"2022-07-07T13:19:33.488198Z","shell.execute_reply":"2022-07-07T13:19:33.865701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since \"Age\" is (right) skewed, using the mean might give us biased results by filling in ages that are older than desired. To deal with this, we'll use the median to impute the missing values. ","metadata":{"_cell_guid":"e62d6951-d968-43ba-aabf-add90524d042","_uuid":"24c201948b9c8c8076ab01271a4790d9db9096b5"}},{"cell_type":"code","source":"train_df[\"Age\"].median(skipna=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:19:34.215843Z","iopub.execute_input":"2022-07-07T13:19:34.216506Z","iopub.status.idle":"2022-07-07T13:19:34.223552Z","shell.execute_reply.started":"2022-07-07T13:19:34.216330Z","shell.execute_reply":"2022-07-07T13:19:34.222222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# mean age\nprint('The mean of \"Age\" is %.2f' %(train_df[\"Age\"].mean(skipna=True)))\n# median age\nprint('The median of \"Age\" is %.2f' %(train_df[\"Age\"].median(skipna=True)))","metadata":{"_cell_guid":"1d70c27b-1e4d-4d5e-8a39-c134389d436c","_uuid":"4f13840d4f9bf1b4331523c99274aa0627485e6c","execution":{"iopub.status.busy":"2022-07-07T13:19:34.562620Z","iopub.execute_input":"2022-07-07T13:19:34.563063Z","iopub.status.idle":"2022-07-07T13:19:34.570219Z","shell.execute_reply.started":"2022-07-07T13:19:34.563023Z","shell.execute_reply":"2022-07-07T13:19:34.569030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"t2.2.\"></a>\n## 2.2. Cabin - Missing Values","metadata":{"_cell_guid":"dea7b01c-c8c1-401f-a336-36ee73de2222","_uuid":"e1a08114e302ddc90266e5f065b3f0b5a200bc89"}},{"cell_type":"code","source":"# percent of missing \"Cabin\" \nprint('Percent of missing \"Cabin\" records is %.2f%%' %((train_df['Cabin'].isnull().sum()/train_df.shape[0])*100))","metadata":{"_cell_guid":"1a1ad808-0a63-43ac-b757-71195880ed4f","_uuid":"1acbce9c6bc5d586dda3e47b7506067a85524e66","execution":{"iopub.status.busy":"2022-07-07T13:19:35.266728Z","iopub.execute_input":"2022-07-07T13:19:35.267433Z","iopub.status.idle":"2022-07-07T13:19:35.274826Z","shell.execute_reply.started":"2022-07-07T13:19:35.267021Z","shell.execute_reply":"2022-07-07T13:19:35.273900Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"77% of records are missing, which means that imputing information and using this variable for prediction is probably not wise.  We'll ignore this variable in our model.","metadata":{"_cell_guid":"eda8c434-63ff-4875-8566-2e194c0d3f66","_uuid":"b6e037c7ac5ec476516031a06b042d8b9999ba44"}},{"cell_type":"markdown","source":"<a id=\"t2.3.\"></a>\n## 2.3. Embarked - Missing Values","metadata":{"_cell_guid":"0e696cff-ca80-4cb5-862c-ee80f4b1ab1f","_uuid":"d575319b1f528c7a153d8ab680282048cb163b14"}},{"cell_type":"code","source":"# percent of missing \"Embarked\" \nprint('Percent of missing \"Embarked\" records is %.2f%%' %((train_df['Embarked'].isnull().sum()/train_df.shape[0])*100))","metadata":{"_cell_guid":"f21c2b55-2126-439d-8b1d-e96dafc97d81","_uuid":"92ab9e62fb62f2a0fb9972baf6ada444187540e6","execution":{"iopub.status.busy":"2022-07-07T13:19:36.244946Z","iopub.execute_input":"2022-07-07T13:19:36.245487Z","iopub.status.idle":"2022-07-07T13:19:36.252911Z","shell.execute_reply.started":"2022-07-07T13:19:36.245441Z","shell.execute_reply":"2022-07-07T13:19:36.251517Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are only 2 (0.22%) missing values for \"Embarked\", so we can just impute with the port where most people boarded.","metadata":{"_cell_guid":"d03a4187-c527-4f71-8260-0495f4523e9e","_uuid":"dc97b80524057522f024d0ae6f1abe77cb994903"}},{"cell_type":"code","source":"train_df['Embarked'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:19:36.945593Z","iopub.execute_input":"2022-07-07T13:19:36.946151Z","iopub.status.idle":"2022-07-07T13:19:36.954659Z","shell.execute_reply.started":"2022-07-07T13:19:36.946072Z","shell.execute_reply":"2022-07-07T13:19:36.953692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Boarded passengers grouped by port of embarkation (C = Cherbourg, Q = Queenstown, S = Southampton):')\nprint(train_df['Embarked'].value_counts())\nsns.countplot(x='Embarked', data=train_df, palette='Set2')\nplt.show()","metadata":{"_cell_guid":"22924bc4-5dfa-4df7-b0d0-de3ede9c58b7","_uuid":"f2a915f45264f8a580de6cc382d96b370eb75730","execution":{"iopub.status.busy":"2022-07-07T13:19:37.328815Z","iopub.execute_input":"2022-07-07T13:19:37.329425Z","iopub.status.idle":"2022-07-07T13:19:37.485854Z","shell.execute_reply.started":"2022-07-07T13:19:37.329374Z","shell.execute_reply":"2022-07-07T13:19:37.484041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('The most common boarding port of embarkation is %s.' %train_df['Embarked'].value_counts().idxmax())","metadata":{"_cell_guid":"def67427-3257-4dce-872e-7f5b4202d18a","_uuid":"c57a9f8a54efa382bc94b695c9664330d01709ea","execution":{"iopub.status.busy":"2022-07-07T13:19:37.671362Z","iopub.execute_input":"2022-07-07T13:19:37.671879Z","iopub.status.idle":"2022-07-07T13:19:37.681384Z","shell.execute_reply.started":"2022-07-07T13:19:37.671834Z","shell.execute_reply":"2022-07-07T13:19:37.680223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"By far the most passengers boarded in Southhampton, so we'll impute those 2 NaN's w/ \"S\".","metadata":{"_cell_guid":"c4c55f99-ce99-44f9-b7a8-d4d623ae9295","_uuid":"19cfaae8c484dcb1d00f69b2771e86dc249e9793"}},{"cell_type":"markdown","source":"<a id=\"t2.4.\"></a>\n## 2.4. Final Adjustments to Data (Train & Test)","metadata":{"_cell_guid":"684c308f-25ae-4039-9332-ddb58953a054","_uuid":"3609e785d210d5a8110f7ce550e61007d066449b"}},{"cell_type":"markdown","source":"Based on my assessment of the missing values in the dataset, I'll make the following changes to the data:\n* If \"Age\" is missing for a given row, I'll impute with 28 (median age).\n* If \"Embarked\" is missing for a riven row, I'll impute with \"S\" (the most common boarding port).\n* I'll ignore \"Cabin\" as a variable. There are too many missing values for imputation. Based on the information available, it appears that this value is associated with the passenger's class and fare paid.","metadata":{"_cell_guid":"b3025cdc-fe9f-43b6-bda1-e45c1f25e77c","_uuid":"06d2762ccec3f11564870fe941fc9ac45d71662f"}},{"cell_type":"code","source":"train_data = train_df.copy()  \n\ntrain_data[\"Age\"].fillna(train_df[\"Age\"].median(skipna=True), inplace=True)\n\ntrain_data[\"Embarked\"].fillna(train_df['Embarked'].value_counts().idxmax(), inplace=True)\n\ntrain_data.drop('Cabin', axis=1, inplace=True)","metadata":{"_cell_guid":"bc0d7121-1008-4890-9043-07eba1524e15","_uuid":"feeed4b6775f88edf5de12b0ee6ee73c16eba61d","execution":{"iopub.status.busy":"2022-07-07T13:19:39.375684Z","iopub.execute_input":"2022-07-07T13:19:39.376297Z","iopub.status.idle":"2022-07-07T13:19:39.387838Z","shell.execute_reply.started":"2022-07-07T13:19:39.376239Z","shell.execute_reply":"2022-07-07T13:19:39.386876Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check missing values in adjusted train data\ntrain_data.isnull().sum()","metadata":{"_cell_guid":"0cfe1c08-71a6-493e-803d-db255af01697","_uuid":"d6be29651bb903964e02d3a7bcc7033513eb76c9","execution":{"iopub.status.busy":"2022-07-07T13:19:39.761347Z","iopub.execute_input":"2022-07-07T13:19:39.761884Z","iopub.status.idle":"2022-07-07T13:19:39.771882Z","shell.execute_reply.started":"2022-07-07T13:19:39.761820Z","shell.execute_reply":"2022-07-07T13:19:39.770854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# preview adjusted train data\ntrain_data.head()","metadata":{"_cell_guid":"10dcfe1b-34f1-4bd8-b937-5ae8daf4a378","_uuid":"3ee37b1151416aeeec8ebd7b94bb0184aabc57cd","execution":{"iopub.status.busy":"2022-07-07T13:19:40.147364Z","iopub.execute_input":"2022-07-07T13:19:40.147747Z","iopub.status.idle":"2022-07-07T13:19:40.183610Z","shell.execute_reply.started":"2022-07-07T13:19:40.147685Z","shell.execute_reply":"2022-07-07T13:19:40.182698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15,8))\nax = train_df[\"Age\"].hist(bins=15, density=True, stacked=True, color='teal', alpha=0.6)\ntrain_df[\"Age\"].plot(kind='density', color='teal')\n\nax = train_data[\"Age\"].hist(bins=15, density=True, stacked=True, color='orange', alpha=0.5)\ntrain_data[\"Age\"].plot(kind='density', color='orange')\n\nax.legend(['Raw Age', 'Adjusted Age'])\nax.set(xlabel='Age')\nplt.xlim(-10,85)\nplt.show()","metadata":{"_cell_guid":"dda26046-b93b-49ee-a52e-35355ecb425c","_uuid":"293aec20df86ef529d10ae1f051dfe921ba07b88","execution":{"iopub.status.busy":"2022-07-07T13:19:40.525916Z","iopub.execute_input":"2022-07-07T13:19:40.526650Z","iopub.status.idle":"2022-07-07T13:19:41.004401Z","shell.execute_reply.started":"2022-07-07T13:19:40.526594Z","shell.execute_reply":"2022-07-07T13:19:41.003615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"t2.4.1.\"></a>\n## 2.4.1. Additional Variables","metadata":{"_cell_guid":"6925fcc2-977b-4369-85e1-77a9210326a7","_uuid":"d8280757e6bc627821fb0540c87ccd6ca110f1e0"}},{"cell_type":"markdown","source":"According to the Kaggle data dictionary, both SibSp and Parch relate to traveling with family.  For simplicity's sake (and to account for possible multicollinearity), I'll combine the effect of these variables into one categorical predictor: whether or not that individual was traveling alone.","metadata":{"_cell_guid":"5cf98f33-fdd5-4a16-b6bf-fa36bc8b84e0","_uuid":"3bfdee842f11d27ca490f466c45ef9bf3673e7ae"}},{"cell_type":"code","source":"## Create categorical variable for traveling alone\nimport numpy as np\ntrain_data['TravelAlone']=np.where((train_data[\"SibSp\"]+train_data[\"Parch\"])>0, 0, 1)\ntrain_data.drop('SibSp', axis=1, inplace=True)\ntrain_data.drop('Parch', axis=1, inplace=True)","metadata":{"_cell_guid":"759c3c8e-8db6-41d9-a1a2-058a15b338a6","_uuid":"d1f5815ba663f7e8cc17d7efcff73653af5b1bdb","execution":{"iopub.status.busy":"2022-07-07T13:19:41.838670Z","iopub.execute_input":"2022-07-07T13:19:41.839124Z","iopub.status.idle":"2022-07-07T13:19:41.892737Z","shell.execute_reply.started":"2022-07-07T13:19:41.839065Z","shell.execute_reply":"2022-07-07T13:19:41.891789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.head(2)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:19:42.230465Z","iopub.execute_input":"2022-07-07T13:19:42.230884Z","iopub.status.idle":"2022-07-07T13:19:42.261872Z","shell.execute_reply.started":"2022-07-07T13:19:42.230815Z","shell.execute_reply":"2022-07-07T13:19:42.260646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I'll also create categorical variables for Passenger Class (\"Pclass\"), Gender (\"Sex\"), and Port Embarked (\"Embarked\"). ","metadata":{"_cell_guid":"e4a22367-b719-4204-952f-d2e9a3b8075e","_uuid":"ca53796bf788bd3b015f1a79a97e050bafa2c770"}},{"cell_type":"code","source":"#create categorical variables and drop some variables\ntraining=pd.get_dummies(train_data, columns=[\"Pclass\",\"Embarked\",\"Sex\"])\ntraining.head()","metadata":{"_cell_guid":"f95361e8-2533-4731-a7ab-a99cf686ed50","_uuid":"4494fcbf9faa90151e20042f74d73395fac3cc8e","execution":{"iopub.status.busy":"2022-07-07T13:19:42.890026Z","iopub.execute_input":"2022-07-07T13:19:42.890393Z","iopub.status.idle":"2022-07-07T13:19:42.939291Z","shell.execute_reply.started":"2022-07-07T13:19:42.890347Z","shell.execute_reply":"2022-07-07T13:19:42.938483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"training.drop('Sex_female', axis=1, inplace=True)\ntraining.drop('PassengerId', axis=1, inplace=True)\ntraining.drop('Name', axis=1, inplace=True)\ntraining.drop('Ticket', axis=1, inplace=True)\n\nfinal_train = training\nfinal_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:19:43.414941Z","iopub.execute_input":"2022-07-07T13:19:43.415630Z","iopub.status.idle":"2022-07-07T13:19:43.449020Z","shell.execute_reply.started":"2022-07-07T13:19:43.415564Z","shell.execute_reply":"2022-07-07T13:19:43.448124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_train.dtypes","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:19:46.239784Z","iopub.execute_input":"2022-07-07T13:19:46.240289Z","iopub.status.idle":"2022-07-07T13:19:46.248616Z","shell.execute_reply.started":"2022-07-07T13:19:46.240222Z","shell.execute_reply":"2022-07-07T13:19:46.247520Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Now, apply the same changes to the test data. <br>\nI will apply to same imputation for \"Age\" in the Test data as I did for my Training data (if missing, Age = 28).  <br> I'll also remove the \"Cabin\" variable from the test data, as I've decided not to include it in my analysis. <br> There were no missing values in the \"Embarked\" port variable. <br> I'll add the dummy variables to finalize the test set.  <br> Finally, I'll impute the 1 missing value for \"Fare\" with the median, 14.45.","metadata":{"_cell_guid":"768cf074-6ecb-47eb-9b42-8b9079ffb811","_uuid":"6a8e533e77c7f1f1a68d136119f93972447a31cf"}},{"cell_type":"code","source":"test_df.isnull().sum()","metadata":{"_cell_guid":"501f9a53-881d-4440-9366-7aae67eb358b","_uuid":"d80416a026d17ccac3bf793408dd5f4f1e17bf63","execution":{"iopub.status.busy":"2022-07-07T13:19:49.704251Z","iopub.execute_input":"2022-07-07T13:19:49.704579Z","iopub.status.idle":"2022-07-07T13:19:49.716585Z","shell.execute_reply.started":"2022-07-07T13:19:49.704499Z","shell.execute_reply":"2022-07-07T13:19:49.715286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data = test_df.copy()\ntest_data[\"Age\"].fillna(train_df[\"Age\"].median(skipna=True), inplace=True)\ntest_data[\"Fare\"].fillna(train_df[\"Fare\"].median(skipna=True), inplace=True)\ntest_data.drop('Cabin', axis=1, inplace=True)\n\ntest_data['TravelAlone']=np.where((test_data[\"SibSp\"]+test_data[\"Parch\"])>0, 0, 1)\n\ntest_data.drop('SibSp', axis=1, inplace=True)\ntest_data.drop('Parch', axis=1, inplace=True)\n\ntesting = pd.get_dummies(test_data, columns=[\"Pclass\",\"Embarked\",\"Sex\"])\ntesting.drop('Sex_female', axis=1, inplace=True)\ntesting.drop('PassengerId', axis=1, inplace=True)\ntesting.drop('Name', axis=1, inplace=True)\ntesting.drop('Ticket', axis=1, inplace=True)\n\nfinal_test = testing\nfinal_test.head()","metadata":{"_cell_guid":"8b9ef076-3669-4339-8d10-0d8783a92e07","_uuid":"145675b90aa2befa533c640aaedd4bf8069b12d4","execution":{"iopub.status.busy":"2022-07-07T13:19:50.115307Z","iopub.execute_input":"2022-07-07T13:19:50.115712Z","iopub.status.idle":"2022-07-07T13:19:50.164146Z","shell.execute_reply.started":"2022-07-07T13:19:50.115672Z","shell.execute_reply":"2022-07-07T13:19:50.163377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"t3.\"></a>\n# 3. Exploratory Data Analysis","metadata":{"_cell_guid":"1430d510-1c8d-4544-8009-3911fff7afbb","_uuid":"4e26c19bf719b7086addc0e1981c00836a19f189"}},{"cell_type":"markdown","source":"<a id=\"t3.1.\"></a>\n## 3.1. Exploration of Age","metadata":{"_cell_guid":"2655428b-d69d-4c0f-85ff-e31ada8e37b9","_uuid":"32e9c04a3281fb1aa8c77e1406c56cd820459202"}},{"cell_type":"code","source":"plt.figure(figsize=(15,8))\nax = sns.kdeplot(final_train[\"Age\"][final_train.Survived == 1], color=\"darkturquoise\", shade=True)\nsns.kdeplot(final_train[\"Age\"][final_train.Survived == 0], color=\"lightcoral\", shade=True)\nplt.legend(['Survived', 'Died'])\nplt.title('Density Plot of Age for Surviving Population and Deceased Population')\nax.set(xlabel='Age')\nplt.xlim(-10,85)\nplt.show()","metadata":{"_cell_guid":"9f9ca9e5-50a0-4487-ba53-815dda90af1c","_uuid":"790e8d7ca89d19e276b3398e299c42893a796b79","execution":{"iopub.status.busy":"2022-07-07T13:19:54.194780Z","iopub.execute_input":"2022-07-07T13:19:54.195496Z","iopub.status.idle":"2022-07-07T13:19:54.537657Z","shell.execute_reply.started":"2022-07-07T13:19:54.195425Z","shell.execute_reply":"2022-07-07T13:19:54.536908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The age distribution for survivors and deceased is actually very similar.  One notable difference is that, of the survivors, a larger proportion were children.  The passengers evidently made an attempt to save children by giving them a place on the life rafts. ","metadata":{"_cell_guid":"8e304d72-27f3-41cf-863f-63872f4c37df","_uuid":"6c5625b454f5e01dd6b6d843d801851c14c64d1e"}},{"cell_type":"markdown","source":"Considering the survival rate of passengers under 16, I'll also include another categorical variable in my dataset: \"Minor\"","metadata":{"_cell_guid":"0b636440-ab38-46a8-8cc9-9421683d5c0b","_uuid":"67051cf653243b3103c9f8015c501d89d92bd3bc"}},{"cell_type":"code","source":"final_train['IsMinor']=np.where(final_train['Age']<=16, 1, 0)\n\nfinal_test['IsMinor']=np.where(final_test['Age']<=16, 1, 0)","metadata":{"_cell_guid":"1655b49b-b33f-4236-8b31-d995ef26c6f6","_uuid":"8918defa6e17b83c700ea45357ebd67a3a22f02f","execution":{"iopub.status.busy":"2022-07-07T13:19:55.228782Z","iopub.execute_input":"2022-07-07T13:19:55.229426Z","iopub.status.idle":"2022-07-07T13:19:55.236520Z","shell.execute_reply.started":"2022-07-07T13:19:55.229347Z","shell.execute_reply":"2022-07-07T13:19:55.235745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"t3.2.\"></a>\n## 3.2. Exploration of Fare","metadata":{"_cell_guid":"a643b196-91c6-4b12-9463-0f984fbfc91a","_uuid":"337b3ced0c6423cf1d126f23a7e60c0181af6a47"}},{"cell_type":"code","source":"plt.figure(figsize=(15,8))\nax = sns.kdeplot(final_train[\"Fare\"][final_train.Survived == 1], color=\"darkturquoise\", shade=True)\nsns.kdeplot(final_train[\"Fare\"][final_train.Survived == 0], color=\"lightcoral\", shade=True)\nplt.legend(['Survived', 'Died'])\nplt.title('Density Plot of Fare for Surviving Population and Deceased Population')\nax.set(xlabel='Fare')\nplt.xlim(-20,200)\nplt.show()","metadata":{"_cell_guid":"9f31ffe1-7cd8-4169-b193-ed44e56d0bd4","_uuid":"4a1c521f08460f6983eca0c4e01294fb7c86e4f9","execution":{"iopub.status.busy":"2022-07-07T13:19:55.929735Z","iopub.execute_input":"2022-07-07T13:19:55.930032Z","iopub.status.idle":"2022-07-07T13:19:56.257688Z","shell.execute_reply.started":"2022-07-07T13:19:55.929981Z","shell.execute_reply":"2022-07-07T13:19:56.256562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As the distributions are clearly different for the fares of survivors vs. deceased, it's likely that this would be a significant predictor in our final model.  Passengers who paid lower fare appear to have been less likely to survive.  This is probably strongly correlated with Passenger Class, which we'll look at next.","metadata":{"_cell_guid":"346b7322-a3e4-48df-bbb1-4d8ec7716f3f","_uuid":"2717310b6c443d675c7342be0c2c18b265723273"}},{"cell_type":"markdown","source":"<a id=\"t3.3.\"></a>\n## 3.3. Exploration of Passenger Class","metadata":{"_cell_guid":"cf585311-4029-4be4-8af2-3eea8258801a","_uuid":"4524affda51265ea23fa923e2ea7f93d7bb91875"}},{"cell_type":"code","source":"sns.barplot('Pclass', 'Survived', data=train_df, color=\"darkturquoise\")\nplt.show()","metadata":{"_cell_guid":"676548e8-6dd4-4180-800c-7b164acb3877","_uuid":"08fd677214959e0b938a0f8a94b63ab548673ea5","execution":{"iopub.status.busy":"2022-07-07T13:19:57.075364Z","iopub.execute_input":"2022-07-07T13:19:57.075651Z","iopub.status.idle":"2022-07-07T13:19:57.313477Z","shell.execute_reply.started":"2022-07-07T13:19:57.075602Z","shell.execute_reply":"2022-07-07T13:19:57.311897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Unsurprisingly, being a first class passenger was safest.","metadata":{"_cell_guid":"193233f8-b220-4cae-aa0f-f822316d5623","_uuid":"8ddb19191253a6e09dfcb0beff2b3690f1052d52"}},{"cell_type":"markdown","source":"<a id=\"t3.4.\"></a>\n## 3.4. Exploration of Embarked Port","metadata":{"_cell_guid":"c59f8e8f-e8c2-40fb-b9c8-12dddd6d318f","_uuid":"2fc06b75321946b721852f78431435f9ba5fef39"}},{"cell_type":"code","source":"sns.barplot('Embarked', 'Survived', data=train_df, color=\"teal\")\nplt.show()","metadata":{"_cell_guid":"6e5bec50-2f5e-433e-9130-c56956fddad3","_uuid":"a9f0598701c7c5224eaa73dafa869af73beffe18","execution":{"iopub.status.busy":"2022-07-07T13:19:58.149201Z","iopub.execute_input":"2022-07-07T13:19:58.149511Z","iopub.status.idle":"2022-07-07T13:19:58.394609Z","shell.execute_reply.started":"2022-07-07T13:19:58.149462Z","shell.execute_reply":"2022-07-07T13:19:58.393717Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Passengers who boarded in Cherbourg, France, appear to have the highest survival rate.  Passengers who boarded in Southhampton were marginally less likely to survive than those who boarded in Queenstown.  This is probably related to passenger class, or maybe even the order of room assignments (e.g. maybe earlier passengers were more likely to have rooms closer to deck). <br> It's also worth noting the size of the whiskers in these plots.  Because the number of passengers who boarded at Southhampton was highest, the confidence around the survival rate is the highest.  The whisker of the Queenstown plot includes the Southhampton average, as well as the lower bound of its whisker.  It's possible that Queenstown passengers were equally, or even more, ill-fated than their Southhampton counterparts.","metadata":{"_cell_guid":"88d78820-35a5-48fd-a234-9f3ca3fca779","_uuid":"2f6a0329cf0c7b771a707ec790efc065924e1ee2"}},{"cell_type":"markdown","source":"<a id=\"t3.5.\"></a>\n## 3.5. Exploration of Traveling Alone vs. With Family","metadata":{"_cell_guid":"9e6dc87e-ba59-4004-8145-79709328fe27","_uuid":"92bacce85a7dec5509217b9570bc2a2fea6a8452"}},{"cell_type":"code","source":"sns.barplot('TravelAlone', 'Survived', data=final_train, color=\"mediumturquoise\")\nplt.show()","metadata":{"_cell_guid":"67017a88-93d4-412b-9adf-8b4d1d9b9db0","_uuid":"e0c3dc16292ef0bcabf0fc680d821ef654084ab4","execution":{"iopub.status.busy":"2022-07-07T13:19:59.316253Z","iopub.execute_input":"2022-07-07T13:19:59.316582Z","iopub.status.idle":"2022-07-07T13:19:59.517975Z","shell.execute_reply.started":"2022-07-07T13:19:59.316524Z","shell.execute_reply":"2022-07-07T13:19:59.516988Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Individuals traveling without family were more likely to die in the disaster than those with family aboard. Given the era, it's likely that individuals traveling alone were likely male.","metadata":{"_cell_guid":"e9e68cef-5e74-46aa-8343-39afbbf00efe","_uuid":"f160bd7399e024ae669d55f09caf6e7902768851"}},{"cell_type":"markdown","source":"<a id=\"t3.6.\"></a>\n## 3.6. Exploration of Gender Variable","metadata":{"_cell_guid":"201b4c9d-b9f0-4ae9-8580-0b4e24ee62be","_uuid":"693c25c25f3590f0b027725471ddd74d56f154af"}},{"cell_type":"code","source":"sns.barplot('Sex', 'Survived', data=train_df, color=\"aquamarine\")\nplt.show()","metadata":{"_cell_guid":"7b416e59-8616-4a44-93e1-a8005eff78a9","_uuid":"354794315925dff1e96229cc737eaf299aaea17a","execution":{"iopub.status.busy":"2022-07-07T13:20:00.371852Z","iopub.execute_input":"2022-07-07T13:20:00.372301Z","iopub.status.idle":"2022-07-07T13:20:00.595366Z","shell.execute_reply.started":"2022-07-07T13:20:00.372250Z","shell.execute_reply":"2022-07-07T13:20:00.594223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is a very obvious difference.  Clearly being female greatly increased your chances of survival.","metadata":{"_cell_guid":"490ed298-f0e4-466b-acc8-81280315e6a2","_uuid":"80c02b9fe2151c443f189cbe44c9cacf7e5c44a4"}},{"cell_type":"markdown","source":"<a id=\"t4.\"></a>\n# 4. Logistic Regression and Results","metadata":{"_cell_guid":"c833cbf5-74db-44ff-90fa-b600ff0a09d7","_uuid":"39dbc095f99dcec6d25a7a4561e81bb641078622"}},{"cell_type":"code","source":"Selected_features = ['Age', 'TravelAlone', 'Pclass_1', 'Pclass_2', 'Embarked_C', \n                     'Embarked_S', 'Sex_male']\nX = final_train[Selected_features]\n\nplt.subplots(figsize=(8, 5))\nsns.heatmap(X.corr(), annot=True, cmap=\"RdYlGn\")\nplt.show()","metadata":{"_cell_guid":"08986ec4-79ff-466b-b763-61bf84a0879b","_uuid":"3f6950a7c24c629b72e17e54c556f3c183b3f779","execution":{"iopub.status.busy":"2022-07-07T13:20:04.570306Z","iopub.execute_input":"2022-07-07T13:20:04.570846Z","iopub.status.idle":"2022-07-07T13:20:05.050730Z","shell.execute_reply.started":"2022-07-07T13:20:04.570784Z","shell.execute_reply":"2022-07-07T13:20:05.049371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"t4.2.\"></a>\n## 4.2. Review of model evaluation procedures\n\nMotivation: Need a way to choose between machine learning models\n* Goal is to estimate likely performance of a model on out-of-sample data\n\nInitial idea: Train and test on the same data\n* But, maximizing training accuracy rewards overly complex models which overfit the training data\n\nAlternative idea: Train/test split\n* Split the dataset into two pieces, so that the model can be trained and tested on different data\n* Testing accuracy is a better estimate than training accuracy of out-of-sample performance\n* Problem with train/test split\n    * It provides a high variance estimate since changing which observations happen to be in the testing set can significantly change testing accuracy\n    * Testing accuracy can change a lot depending on a which observation happen to be in the testing set\n\nReference: <br>\nhttp://www.ritchieng.com/machine-learning-cross-validation/ <br>","metadata":{"_cell_guid":"a7455afe-9716-4189-b207-f1cc9facce12","_uuid":"46b76691c5f109b17f805f233a5ad5ba900b353b"}},{"cell_type":"markdown","source":"## Model ","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nfrom sklearn.datasets import make_classification\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.neighbors import KNeighborsClassifier \nfrom sklearn import metrics","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:20:29.020469Z","iopub.execute_input":"2022-07-07T13:20:29.020943Z","iopub.status.idle":"2022-07-07T13:20:29.183221Z","shell.execute_reply.started":"2022-07-07T13:20:29.020899Z","shell.execute_reply":"2022-07-07T13:20:29.181864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split, cross_val_score\nfrom sklearn.metrics import accuracy_score, classification_report, precision_score, recall_score \nfrom sklearn.metrics import confusion_matrix, precision_recall_curve, roc_curve, auc, log_loss\nfrom sklearn.linear_model import LogisticRegression\n\ncols = [\"Age\",\"Fare\",\"TravelAlone\",\"Pclass_1\",\"Pclass_2\",\"Embarked_C\",\"Embarked_S\",\"Sex_male\"] \nX = final_train[cols]\ny = final_train['Survived']\n\n# create X (features) and y (response)\n#X = final_train[Selected_features]\n#y = final_train['Survived']\n\n# use train/test split with different random_state values\n# we can change the random_state values that changes the accuracy scores\n# the scores change a lot, this is why testing scores is a high-variance estimate\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=2)\n\nerror1= []\nerror2= []\nfor k in range(1,15):\n    knn= KNeighborsClassifier(n_neighbors=k)\n    knn.fit(X_train,y_train)\n    y_pred1= knn.predict(X_train)\n    error1.append(np.mean(y_train!= y_pred1))\n    y_pred2= knn.predict(X_test)\n    error2.append(np.mean(y_test!= y_pred2))\n# plt.figure(figsize(10,5))\nplt.plot(range(1,15),error1,label=\"train\")\nplt.plot(range(1,15),error2,label=\"test\")\nplt.xlabel('k Value')\nplt.ylabel('Error')\nplt.legend()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:21:15.657656Z","iopub.execute_input":"2022-07-07T13:21:15.658008Z","iopub.status.idle":"2022-07-07T13:21:16.220586Z","shell.execute_reply.started":"2022-07-07T13:21:15.657955Z","shell.execute_reply":"2022-07-07T13:21:16.219639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"knn= KNeighborsClassifier(n_neighbors=7)\nknn.fit(X_train,y_train)\ny_pred= knn.predict(X_test)\nmetrics.accuracy_score(y_test,y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:21:55.926143Z","iopub.execute_input":"2022-07-07T13:21:55.926773Z","iopub.status.idle":"2022-07-07T13:21:55.939589Z","shell.execute_reply.started":"2022-07-07T13:21:55.926412Z","shell.execute_reply":"2022-07-07T13:21:55.938623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred = knn.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:24:29.448353Z","iopub.execute_input":"2022-07-07T13:24:29.448716Z","iopub.status.idle":"2022-07-07T13:24:29.456122Z","shell.execute_reply.started":"2022-07-07T13:24:29.448659Z","shell.execute_reply":"2022-07-07T13:24:29.455128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import classification_report\nprint(classification_report(y_test,y_pred))","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:24:47.334782Z","iopub.execute_input":"2022-07-07T13:24:47.335403Z","iopub.status.idle":"2022-07-07T13:24:47.343872Z","shell.execute_reply.started":"2022-07-07T13:24:47.335350Z","shell.execute_reply":"2022-07-07T13:24:47.343266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 10-fold cross-validation KNN\nknn = KNeighborsClassifier(n_neighbors=7)\n# Use cross_val_score function\n# We are passing the entirety of X and y, not X_train or y_train, it takes care of splitting the data\n# cv=10 for 10 folds\n# scoring = {'accuracy', 'neg_log_loss', 'roc_auc'} for evaluation metric - althought they are many\nscores_accuracy = cross_val_score(knn, X, y, cv=10, scoring='accuracy')\nscores_auc = cross_val_score(knn, X, y, cv=10, scoring='roc_auc')\nprint('K-fold cross-validation results:')\nprint(knn.__class__.__name__+\" average accuracy is %2.3f\" % scores_accuracy.mean())\nprint(knn.__class__.__name__+\" average auc is %2.3f\" % scores_auc.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:26:01.391714Z","iopub.execute_input":"2022-07-07T13:26:01.392123Z","iopub.status.idle":"2022-07-07T13:26:01.504556Z","shell.execute_reply.started":"2022-07-07T13:26:01.392044Z","shell.execute_reply":"2022-07-07T13:26:01.503883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV\ngrid={\"n_neighbors\":[2,3,4,5,6,7,8]}# l1 lasso l2 ridge\nknn = KNeighborsClassifier()\nknn_cv=GridSearchCV(knn,grid,cv=10)\nknn_cv.fit(X,y)\n\nprint(\"tuned hpyerparameters :(best parameters) \",knn_cv.best_params_)\nprint(\"accuracy :\",knn_cv.best_score_)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:27:15.432106Z","iopub.execute_input":"2022-07-07T13:27:15.432772Z","iopub.status.idle":"2022-07-07T13:27:16.145873Z","shell.execute_reply.started":"2022-07-07T13:27:15.432416Z","shell.execute_reply":"2022-07-07T13:27:16.145234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split, cross_val_score\nfrom sklearn.metrics import accuracy_score, classification_report, precision_score, recall_score \nfrom sklearn.metrics import confusion_matrix, precision_recall_curve, roc_curve, auc, log_loss\nfrom sklearn.linear_model import LogisticRegression\n\ncols = [\"Age\",\"Fare\",\"TravelAlone\",\"Pclass_1\",\"Pclass_2\",\"Embarked_C\",\"Embarked_S\",\"Sex_male\"] \nX = final_train[cols]\ny = final_train['Survived']\n\n# create X (features) and y (response)\n#X = final_train[Selected_features]\n#y = final_train['Survived']\n\n# use train/test split with different random_state values\n# we can change the random_state values that changes the accuracy scores\n# the scores change a lot, this is why testing scores is a high-variance estimate\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=2)\n\n# check classification scores of logistic regression\nlogreg = LogisticRegression()\nlogreg.fit(X_train, y_train)\n\ny_pred = logreg.predict(X_test)\n\ny_pred_proba = logreg.predict_proba(X_test)[:, 1]\n\n[fpr, tpr, thr] = roc_curve(y_test, y_pred_proba)\n\nprint('Train/Test split results:')\nprint(logreg.__class__.__name__+\" accuracy is %2.3f\" % accuracy_score(y_test, y_pred))\nprint(logreg.__class__.__name__+\" log_loss is %2.3f\" % log_loss(y_test, y_pred_proba))\nprint(logreg.__class__.__name__+\" auc is %2.3f\" % auc(fpr, tpr))\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-04T12:12:37.886929Z","iopub.execute_input":"2022-07-04T12:12:37.887226Z","iopub.status.idle":"2022-07-04T12:12:38.109436Z","shell.execute_reply.started":"2022-07-04T12:12:37.887187Z","shell.execute_reply":"2022-07-04T12:12:38.108426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"t4.2.1.\"></a>\n### 4.2.1. Model evaluation based on simple train/test split using `train_test_split()` function","metadata":{"_cell_guid":"b894002e-07cf-4d02-b708-a2ac387eed54","_uuid":"e35125f8aa230d4875541aa4f6b5964d2f14a6a3"}},{"cell_type":"code","source":"y_pred","metadata":{"execution":{"iopub.status.busy":"2022-07-04T12:12:41.174054Z","iopub.execute_input":"2022-07-04T12:12:41.174433Z","iopub.status.idle":"2022-07-04T12:12:41.181862Z","shell.execute_reply.started":"2022-07-04T12:12:41.17436Z","shell.execute_reply":"2022-07-04T12:12:41.180669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import classification_report\nprint(classification_report(y_test,y_pred))","metadata":{"execution":{"iopub.status.busy":"2022-07-04T12:12:41.538408Z","iopub.execute_input":"2022-07-04T12:12:41.539068Z","iopub.status.idle":"2022-07-04T12:12:41.547665Z","shell.execute_reply.started":"2022-07-04T12:12:41.539006Z","shell.execute_reply":"2022-07-04T12:12:41.546907Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split, cross_val_score\nfrom sklearn.metrics import accuracy_score, classification_report, precision_score, recall_score \nfrom sklearn.metrics import confusion_matrix, precision_recall_curve, roc_curve, auc, log_loss\n\ncols = [\"Age\",\"Fare\",\"TravelAlone\",\"Pclass_1\",\"Pclass_2\",\"Embarked_C\",\"Embarked_S\",\"Sex_male\"] \nX = final_train[cols]\ny = final_train['Survived']\n# Build a logreg and compute the feature importances\nmodel = LogisticRegression() \n\n# use train/test split with different random_state values\n# we can change the random_state values that changes the accuracy scores\n# the scores change a lot, this is why testing scores is a high-variance estimate\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=2)\n\n# check classification scores of logistic regression\nlogreg = LogisticRegression()\nlogreg.fit(X_train, y_train)\ny_pred = logreg.predict(X_test)\ny_pred_proba = logreg.predict_proba(X_test)[:, 1]\n[fpr, tpr, thr] = roc_curve(y_test, y_pred_proba)\nprint('Train/Test split results:')\nprint(logreg.__class__.__name__+\" accuracy is %2.3f\" % accuracy_score(y_test, y_pred))\nprint(logreg.__class__.__name__+\" log_loss is %2.3f\" % log_loss(y_test, y_pred_proba))\nprint(logreg.__class__.__name__+\" auc is %2.3f\" % auc(fpr, tpr))\n\nidx = np.min(np.where(tpr > 0.95)) # index of the first threshold for which the sensibility > 0.95\n\nplt.figure()\nplt.plot(fpr, tpr, color='coral', label='ROC curve (area = %0.3f)' % auc(fpr, tpr))\nplt.plot([0, 1], [0, 1], 'k--')\nplt.plot([0,fpr[idx]], [tpr[idx],tpr[idx]], 'k--', color='blue')\nplt.plot([fpr[idx],fpr[idx]], [0,tpr[idx]], 'k--', color='blue')\nplt.xlim([0.0, 1.0])\nplt.ylim([0.0, 1.05])\nplt.xlabel('False Positive Rate (1 - specificity)', fontsize=14)\nplt.ylabel('True Positive Rate (recall)', fontsize=14)\nplt.title('Receiver operating characteristic (ROC) curve')\nplt.legend(loc=\"lower right\")\nplt.show()\n\nprint(\"Using a threshold of %.3f \" % thr[idx] + \"guarantees a sensitivity of %.3f \" % tpr[idx] +  \n      \"and a specificity of %.3f\" % (1-fpr[idx]) + \n      \", i.e. a false positive rate of %.2f%%.\" % (np.array(fpr[idx])*100))","metadata":{"_cell_guid":"84233f59-f3c7-4ea0-884d-96f8ad4d5b10","_uuid":"46336228eeb864bc82e6739768122579d1c9634c","execution":{"iopub.status.busy":"2022-07-04T12:13:17.312041Z","iopub.execute_input":"2022-07-04T12:13:17.312441Z","iopub.status.idle":"2022-07-04T12:13:17.561707Z","shell.execute_reply.started":"2022-07-04T12:13:17.31237Z","shell.execute_reply":"2022-07-04T12:13:17.560755Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"t4.2.2.\"></a>\n### 4.2.2. Model evaluation based on K-fold cross-validation using `cross_val_score()` function ","metadata":{"_cell_guid":"6292c3f2-6be9-45e4-be33-1dca41e604a7","_uuid":"ba0017b461cea0b8849746e76475598bbba7c9ce"}},{"cell_type":"code","source":"# 10-fold cross-validation logistic regression\nlogreg = LogisticRegression()\n# Use cross_val_score function\n# We are passing the entirety of X and y, not X_train or y_train, it takes care of splitting the data\n# cv=10 for 10 folds\n# scoring = {'accuracy', 'neg_log_loss', 'roc_auc'} for evaluation metric - althought they are many\nscores_accuracy = cross_val_score(logreg, X, y, cv=10, scoring='accuracy')\nscores_log_loss = cross_val_score(logreg, X, y, cv=10, scoring='neg_log_loss')\nscores_auc = cross_val_score(logreg, X, y, cv=10, scoring='roc_auc')\nprint('K-fold cross-validation results:')\nprint(logreg.__class__.__name__+\" average accuracy is %2.3f\" % scores_accuracy.mean())\nprint(logreg.__class__.__name__+\" average log_loss is %2.3f\" % -scores_log_loss.mean())\nprint(logreg.__class__.__name__+\" average auc is %2.3f\" % scores_auc.mean())","metadata":{"_cell_guid":"32a611ae-b2b7-43e0-8fa8-3cc56e351bf6","_uuid":"7f0aba7b861c3fa1748060b4733778851fb00a31","execution":{"iopub.status.busy":"2022-07-04T12:14:01.519315Z","iopub.execute_input":"2022-07-04T12:14:01.519661Z","iopub.status.idle":"2022-07-04T12:14:01.740769Z","shell.execute_reply.started":"2022-07-04T12:14:01.519594Z","shell.execute_reply":"2022-07-04T12:14:01.739764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scores_accuracy","metadata":{"execution":{"iopub.status.busy":"2022-07-04T12:14:08.052322Z","iopub.execute_input":"2022-07-04T12:14:08.052668Z","iopub.status.idle":"2022-07-04T12:14:08.058995Z","shell.execute_reply.started":"2022-07-04T12:14:08.0526Z","shell.execute_reply":"2022-07-04T12:14:08.058039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"t4.2.3.\"></a>\n### 4.2.3. Model evaluation based on K-fold cross-validation using `cross_validate()` function ","metadata":{"_cell_guid":"bf358561-2762-4b24-b33b-bfda3c5d725f","_uuid":"9debd9d4281cf7893159aa2cb1f7993652d60228"}},{"cell_type":"code","source":"from sklearn.model_selection import cross_validate\n\nscoring = {'accuracy': 'accuracy', 'log_loss': 'neg_log_loss', 'auc': 'roc_auc'}\n\nmodelCV = LogisticRegression()\n\nresults = cross_validate(modelCV, X, y, cv=10, scoring=list(scoring.values()), \n                         return_train_score=False)\n\nprint('K-fold cross-validation results:')\nfor sc in range(len(scoring)):\n    print(modelCV.__class__.__name__+\" average %s: %.3f (+/-%.3f)\" % (list(scoring.keys())[sc], -results['test_%s' % list(scoring.values())[sc]].mean()\n                               if list(scoring.values())[sc]=='neg_log_loss' \n                               else results['test_%s' % list(scoring.values())[sc]].mean(), \n                               results['test_%s' % list(scoring.values())[sc]].std()))","metadata":{"_cell_guid":"9ea95aac-00b2-413e-ab86-c6b8782c40ef","_uuid":"90840c89d42e9284c9480a256dc259778c3a5b1b","execution":{"iopub.status.busy":"2022-07-04T12:14:22.604598Z","iopub.execute_input":"2022-07-04T12:14:22.605077Z","iopub.status.idle":"2022-07-04T12:14:22.709743Z","shell.execute_reply.started":"2022-07-04T12:14:22.605016Z","shell.execute_reply":"2022-07-04T12:14:22.708814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"results","metadata":{"execution":{"iopub.status.busy":"2022-07-04T12:14:38.200207Z","iopub.execute_input":"2022-07-04T12:14:38.200577Z","iopub.status.idle":"2022-07-04T12:14:38.207656Z","shell.execute_reply.started":"2022-07-04T12:14:38.200528Z","shell.execute_reply":"2022-07-04T12:14:38.206702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"t4.3.\"></a>\n## 4.3. GridSearchCV evaluating using multiple scorers simultaneously","metadata":{"_cell_guid":"f485ca4c-2172-4383-979e-21e78a66192c","_uuid":"ec44fbaddbc23f03a8ac470391f41ce35f40cf72"}},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV\ngrid={\"C\":np.arange(1e-05, 3, 0.1), \"penalty\":[\"l1\",\"l2\"]}# l1 lasso l2 ridge\nlogreg=LogisticRegression()\nlogreg_cv=GridSearchCV(logreg,grid,cv=10)\nlogreg_cv.fit(X,y)\n\nprint(\"tuned hpyerparameters :(best parameters) \",logreg_cv.best_params_)\nprint(\"accuracy :\",logreg_cv.best_score_)","metadata":{"execution":{"iopub.status.busy":"2022-07-04T12:19:00.071611Z","iopub.execute_input":"2022-07-04T12:19:00.071995Z","iopub.status.idle":"2022-07-04T12:19:04.512091Z","shell.execute_reply.started":"2022-07-04T12:19:00.071909Z","shell.execute_reply":"2022-07-04T12:19:04.51119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import RandomizedSearchCV\ngrid={\"C\":np.arange(1e-05, 3, 0.1), \"penalty\":[\"l1\",\"l2\"]}# l1 lasso l2 ridge\nlogreg=LogisticRegression()\nlogreg_cv=RandomizedSearchCV(logreg,grid,cv=10)\nlogreg_cv.fit(X,y)\n\nprint(\"tuned hpyerparameters :(best parameters) \",logreg_cv.best_params_)\nprint(\"accuracy :\",logreg_cv.best_score_)","metadata":{"execution":{"iopub.status.busy":"2022-07-04T12:20:15.82986Z","iopub.execute_input":"2022-07-04T12:20:15.830583Z","iopub.status.idle":"2022-07-04T12:20:16.599889Z","shell.execute_reply.started":"2022-07-04T12:20:15.830228Z","shell.execute_reply":"2022-07-04T12:20:16.598909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}