{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Titanic: Survival Prediction Using Machine Learning\nIn this notebook we are going to be doing the **Titanic - Machine Learning from Disaster** competition by Kaggle.\nThe objective of the notebook is:\n* Predict if the passengers **survived** or **not**.\n\n**IMPORTANT**: This dataset is not the only way of aproaching the competition, is how i thought as a trainee data scientist. I hope it will be helpful for all of you, all constructive comments are welcome. Thank you! \n\n**Lets start!**","metadata":{}},{"cell_type":"markdown","source":"# Importing Libraries","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport xgboost as xgb\n\nfrom sklearn import preprocessing\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.model_selection import RandomizedSearchCV\nfrom sklearn.metrics import accuracy_score\nfrom sklearn.model_selection import GridSearchCV\nfrom sklearn.model_selection import RandomizedSearchCV\nfrom sklearn.metrics import f1_score\nfrom sklearn.metrics import recall_score","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:54.153157Z","iopub.execute_input":"2022-07-25T17:02:54.153944Z","iopub.status.idle":"2022-07-25T17:02:54.162213Z","shell.execute_reply.started":"2022-07-25T17:02:54.153891Z","shell.execute_reply":"2022-07-25T17:02:54.160867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loading the Data","metadata":{}},{"cell_type":"code","source":"train_data = pd.read_csv('../input/c/titanic/train.csv')\ntrain_data","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:54.368490Z","iopub.execute_input":"2022-07-25T17:02:54.369135Z","iopub.status.idle":"2022-07-25T17:02:54.407211Z","shell.execute_reply.started":"2022-07-25T17:02:54.369100Z","shell.execute_reply":"2022-07-25T17:02:54.406313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data = pd.read_csv('../input/c/titanic/test.csv')\ntest_data","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:54.609970Z","iopub.execute_input":"2022-07-25T17:02:54.610599Z","iopub.status.idle":"2022-07-25T17:02:54.643981Z","shell.execute_reply.started":"2022-07-25T17:02:54.610565Z","shell.execute_reply":"2022-07-25T17:02:54.642952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Analysis\nLets have a look to the data, in order to see what we are dealing with.","metadata":{}},{"cell_type":"code","source":"train_data.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:54.800135Z","iopub.execute_input":"2022-07-25T17:02:54.800589Z","iopub.status.idle":"2022-07-25T17:02:54.820946Z","shell.execute_reply.started":"2022-07-25T17:02:54.800555Z","shell.execute_reply":"2022-07-25T17:02:54.819960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that we have **Nan values** in three variables:\n* 'Age'\n* 'Cabin' (Most values are Nan, we will have to analyse why, if is a missing data problem or maybe they did not have a cabin.\n* 'Embarked' (Just 2)","metadata":{}},{"cell_type":"code","source":"cols = ['Embarked', 'Sex', 'Pclass', 'Survived']\n\nfor col in cols:\n    print(train_data[col].value_counts(normalize=True).mul(100).round(1).astype(str) + '%')\n    print('\\n')","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:54.991786Z","iopub.execute_input":"2022-07-25T17:02:54.992634Z","iopub.status.idle":"2022-07-25T17:02:55.012141Z","shell.execute_reply.started":"2022-07-25T17:02:54.992600Z","shell.execute_reply":"2022-07-25T17:02:55.010513Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"With those simple lines of code, we can observe that:\n* The data set is **unbalanced** (having 23% between survived and non survived passengers),\n* Most people embarked on **S**(Southampton), the second busiest was **C**(Cherbourg) and the one where less people embarked is **Q** (Queenstown),\n* There were 30% more of male passengers on board the Titanic.\n* Over half of the passengers belonged to 3rd class. And the rest was split almost equally between 1st and 2nd class.","metadata":{}},{"cell_type":"code","source":"print(train_data['Age'].describe())\nprint('\\n')\nprint(train_data['Fare'].describe())","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:55.179913Z","iopub.execute_input":"2022-07-25T17:02:55.180506Z","iopub.status.idle":"2022-07-25T17:02:55.199112Z","shell.execute_reply.started":"2022-07-25T17:02:55.180467Z","shell.execute_reply":"2022-07-25T17:02:55.197427Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**1.** As mentioned before, in 'Age' we are missing some values. We can highlight that the youngest person on board must have been a new born, as it has less than 1 year of age. While the oldest was 80 years old. \n\n**2.** We can see in the second chart that the max Fare is a lot bigger that the mean, so it may be an outlier or will see after analysing that variable.","metadata":{}},{"cell_type":"code","source":"print(train_data['Name'].head(20))","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:55.374126Z","iopub.execute_input":"2022-07-25T17:02:55.375532Z","iopub.status.idle":"2022-07-25T17:02:55.383211Z","shell.execute_reply.started":"2022-07-25T17:02:55.375489Z","shell.execute_reply":"2022-07-25T17:02:55.382108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that all passengers have a 'Status' inside their names. We are going to extract it and try to get some info out of them.","metadata":{}},{"cell_type":"code","source":"train_data['Status'] = train_data['Name'].str.extract(' ([A-Za-z]+)\\.', expand=False)\nprint(train_data['Status'].value_counts())","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:55.559751Z","iopub.execute_input":"2022-07-25T17:02:55.560459Z","iopub.status.idle":"2022-07-25T17:02:55.574081Z","shell.execute_reply.started":"2022-07-25T17:02:55.560423Z","shell.execute_reply":"2022-07-25T17:02:55.573069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''\nSome of the 'Status' are very similar, we are going to group them in order to have a more clear idea of what we are dealing with.\n'''\ntrain_data['Status'] = train_data['Status'].replace(['Lady', 'Countess', 'Don', 'Sir', 'Jonkheer', 'Dona'], 'Royal')\n\ntrain_data['Status'] = train_data['Status'].replace(['Capt', 'Col','Major', 'Rev', 'Dr'], 'Officers')\n\ntrain_data['Status'] = train_data['Status'].replace('Mlle', 'Miss')\ntrain_data['Status'] = train_data['Status'].replace('Ms', 'Miss')\ntrain_data['Status'] = train_data['Status'].replace('Mme', 'Mrs')\n\nprint(train_data['Status'].value_counts())","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:55.748042Z","iopub.execute_input":"2022-07-25T17:02:55.748814Z","iopub.status.idle":"2022-07-25T17:02:55.770611Z","shell.execute_reply.started":"2022-07-25T17:02:55.748770Z","shell.execute_reply":"2022-07-25T17:02:55.768438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train_data.loc[train_data['Status'] == 'Master', 'Age'].describe())\nprint('\\n')\nprint(train_data.loc[train_data['Status'] == 'Master', 'Sex'].value_counts())","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:55.995987Z","iopub.execute_input":"2022-07-25T17:02:55.996509Z","iopub.status.idle":"2022-07-25T17:02:56.012527Z","shell.execute_reply.started":"2022-07-25T17:02:55.996461Z","shell.execute_reply":"2022-07-25T17:02:56.011010Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can observe that ALL passengers with the **'Status == Master'** are **under 12** years old and are **male**. \nWe are going to use this info to try to fill some missing data in the most accurate way.","metadata":{}},{"cell_type":"code","source":"train_data.groupby(['Status','Sex']).Age.mean()","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:56.309999Z","iopub.execute_input":"2022-07-25T17:02:56.311761Z","iopub.status.idle":"2022-07-25T17:02:56.328320Z","shell.execute_reply.started":"2022-07-25T17:02:56.311699Z","shell.execute_reply":"2022-07-25T17:02:56.327223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train_data[train_data['Embarked'].isnull()])","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:56.591710Z","iopub.execute_input":"2022-07-25T17:02:56.592182Z","iopub.status.idle":"2022-07-25T17:02:56.611110Z","shell.execute_reply.started":"2022-07-25T17:02:56.592148Z","shell.execute_reply":"2022-07-25T17:02:56.609600Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After doing some futher analysis on the two missing values in 'Embarked', i run into a forum that talked about it. It explained that 'Icard, Miss. Amelie' boarded the titanic as the maid of 'Stone, Mrs. George Nelson (Martha Evelyn). They both embarked in 'Southampton'. (This is the only external information used in this notebook)","metadata":{}},{"cell_type":"code","source":"train_data['Cabin'] = train_data['Cabin'].str.extract('([A-Za-z]+)', expand=False)\nprint(train_data['Cabin'].value_counts())","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:56.778020Z","iopub.execute_input":"2022-07-25T17:02:56.778966Z","iopub.status.idle":"2022-07-25T17:02:56.790281Z","shell.execute_reply.started":"2022-07-25T17:02:56.778911Z","shell.execute_reply":"2022-07-25T17:02:56.788757Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We extracted the letter from the 'Cabin' variable. I want to see if there is any posibility of filling the missing values or getting something out of this column. Or if it is just useless.","metadata":{}},{"cell_type":"markdown","source":"# Nan Values","metadata":{}},{"cell_type":"code","source":"'''\nFilling all the Males with rank master the mean we obtain before\n'''\ntrain_data['Age'] = np.where((train_data['Status'] == 'Master') & (train_data['Age'].isnull()), train_data['Age'].fillna(4.5), train_data['Age'])\n'''\nWe will fill the rest with the average of each social status\n'''\ntrain_data['Age'] = np.where((train_data['Status'] == 'Miss') & (train_data['Age'].isnull()), train_data['Age'].fillna(21.8), train_data['Age'])\n\ntrain_data['Age'] = np.where((train_data['Status'] == 'Mr') & (train_data['Age'].isnull()), train_data['Age'].fillna(32.3), train_data['Age'])\n\ntrain_data['Age'] = np.where((train_data['Status'] == 'Mrs') & (train_data['Age'].isnull()), train_data['Age'].fillna(35.8), train_data['Age'])\n\ntrain_data['Age'] = np.where((train_data['Status'] == 'Officers') & (train_data['Age'].isnull()), train_data['Age'].fillna(47.5), train_data['Age'])\n\ntrain_data['Age'] = np.where((train_data['Status'] == 'Royal') & (train_data['Age'].isnull()), train_data['Age'].fillna(41), train_data['Age'])\n\n'''\nFilling the 'Embarked' with S (explained before)\n'''\ntrain_data['Embarked'] = train_data['Embarked'].fillna('S')","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:56.964758Z","iopub.execute_input":"2022-07-25T17:02:56.965403Z","iopub.status.idle":"2022-07-25T17:02:56.987595Z","shell.execute_reply.started":"2022-07-25T17:02:56.965369Z","shell.execute_reply":"2022-07-25T17:02:56.985582Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''\nWith Cabin we are going to try two methods,assigning if they had or not a cabin. \nAnd we will also try removing the variable and see what happends. \nFirst we are going to say if they had or not a cabin, and see how the model reacts. If we see that\nit is useful we will leave it, else we will drop the column or check other more apropiate methods.\n'''\ntrain_data['Has_cabin'] = np.where(train_data['Cabin'].isnull(), 0, 1)\ntrain_data.drop('Cabin', axis = 1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:57.151790Z","iopub.execute_input":"2022-07-25T17:02:57.152226Z","iopub.status.idle":"2022-07-25T17:02:57.163028Z","shell.execute_reply.started":"2022-07-25T17:02:57.152192Z","shell.execute_reply":"2022-07-25T17:02:57.161129Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"columns = train_data.columns\nfor column in columns:\n    has_nan = train_data[column].isnull().sum(axis = 0)\n    print(f'{column}: {has_nan}')\n    \nprint(train_data.isnull().values.any())","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:57.352555Z","iopub.execute_input":"2022-07-25T17:02:57.353726Z","iopub.status.idle":"2022-07-25T17:02:57.369916Z","shell.execute_reply.started":"2022-07-25T17:02:57.353686Z","shell.execute_reply":"2022-07-25T17:02:57.368421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have taken care of all the Nan values in the data set.","metadata":{}},{"cell_type":"markdown","source":"# Feature Engineering","metadata":{}},{"cell_type":"code","source":"'''\nCreating variable to detect if it was considered a child or not. Acording to the oficial Titanic data, 12 or under was \nconsider a child. Older it was an adult.\n'''\ntrain_data['Is_child'] = np.where(train_data['Age'] <= 12, 1, 0)\n\n'''\nTransform Sex in numeric\n'''\ntrain_data['Sex'] = np.where(train_data['Sex'] == 'male', 1,0)\n\n'''\nFrom the Name we have already extracted the 'Status', maybe we could do something with the families, so just in case we will\nalso extract the surname in case we want to use it later.\n'''\ntrain_data['Surname'] = train_data['Name'].str.extract('([A-Za-z]+)\\,', expand=False)\nprint(train_data['Surname'].value_counts())\ntrain_data.drop('Name', axis = 1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:57.582295Z","iopub.execute_input":"2022-07-25T17:02:57.582826Z","iopub.status.idle":"2022-07-25T17:02:57.602159Z","shell.execute_reply.started":"2022-07-25T17:02:57.582789Z","shell.execute_reply":"2022-07-25T17:02:57.601168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:57.780049Z","iopub.execute_input":"2022-07-25T17:02:57.780516Z","iopub.status.idle":"2022-07-25T17:02:57.815686Z","shell.execute_reply.started":"2022-07-25T17:02:57.780482Z","shell.execute_reply":"2022-07-25T17:02:57.814066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.drop('Ticket', axis = 1, inplace = True)\n'''\nWe drop the Ticket variable because a large % of them are unique values and i will be hard to make conclusion out of it.\n'''","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:57.945620Z","iopub.execute_input":"2022-07-25T17:02:57.946938Z","iopub.status.idle":"2022-07-25T17:02:57.956845Z","shell.execute_reply.started":"2022-07-25T17:02:57.946899Z","shell.execute_reply":"2022-07-25T17:02:57.955307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''\nFamily organization\n'''\ntrain_data['Fam_size'] = train_data['SibSp'] + train_data['Parch']\n\ntrain_data['Solo'] = np.where(train_data['Fam_size'] == 0, 1, 0)\n\ntrain_data.drop(['SibSp', 'Parch'], axis = 1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:58.105857Z","iopub.execute_input":"2022-07-25T17:02:58.106762Z","iopub.status.idle":"2022-07-25T17:02:58.119231Z","shell.execute_reply.started":"2022-07-25T17:02:58.106727Z","shell.execute_reply":"2022-07-25T17:02:58.118122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''\nFare\nWe will transfor the Fare (which is per family/group) into a personal fare.\n'''\ntrain_data['Fare'] = train_data['Fare'] / (train_data['Fam_size'] + 1)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:58.256201Z","iopub.execute_input":"2022-07-25T17:02:58.256862Z","iopub.status.idle":"2022-07-25T17:02:58.264100Z","shell.execute_reply.started":"2022-07-25T17:02:58.256826Z","shell.execute_reply":"2022-07-25T17:02:58.262767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''\nI tried a couple variables using the surname column that we extracted before, but non of those were game changing so \nwe drop the 'Surname' column for now.\n'''\ntrain_data.drop('Surname', axis = 1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:58.425733Z","iopub.execute_input":"2022-07-25T17:02:58.426610Z","iopub.status.idle":"2022-07-25T17:02:58.433503Z","shell.execute_reply.started":"2022-07-25T17:02:58.426575Z","shell.execute_reply":"2022-07-25T17:02:58.432440Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Bins for 'Age'\nWe are going to try using bins. If it does not work as we want, we will normalize the data","metadata":{}},{"cell_type":"code","source":"train_data['Age'] = pd.cut(train_data['Age'], 6)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:58.583923Z","iopub.execute_input":"2022-07-25T17:02:58.584683Z","iopub.status.idle":"2022-07-25T17:02:58.595217Z","shell.execute_reply.started":"2022-07-25T17:02:58.584648Z","shell.execute_reply":"2022-07-25T17:02:58.593813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Dummies Variables","metadata":{}},{"cell_type":"code","source":"def createDummies(df, var_name):\n    dummy = pd.get_dummies(df[var_name], prefix=var_name)\n    df = df.drop(var_name, axis= 1)\n    df = pd.concat([df, dummy], axis = 1)\n    return df","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:58.730280Z","iopub.execute_input":"2022-07-25T17:02:58.731111Z","iopub.status.idle":"2022-07-25T17:02:58.737027Z","shell.execute_reply.started":"2022-07-25T17:02:58.731069Z","shell.execute_reply":"2022-07-25T17:02:58.736040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = createDummies(train_data, 'Status')\n\ntrain_data = createDummies(train_data, 'Embarked')\n\ntrain_data = createDummies(train_data, 'Age')\n","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:58.878390Z","iopub.execute_input":"2022-07-25T17:02:58.878847Z","iopub.status.idle":"2022-07-25T17:02:58.895873Z","shell.execute_reply.started":"2022-07-25T17:02:58.878814Z","shell.execute_reply":"2022-07-25T17:02:58.894681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Normalizing the Data\nFor now we are only going to normalize the Fare column","metadata":{}},{"cell_type":"code","source":"'''\nHandleling Outliers\n'''\n\ndef outliers_check(df, feature_1, feature_2, threshold):\n\n    fig, axes = plt.subplots(1, 2, figsize=(16, 6))\n    fig.suptitle(f'Outliers Before and After Handleling - {feature_1} - z = {threshold}', fontsize=20)\n    \n    sns.scatterplot(ax=axes[0], data=df, x=feature_1, y=feature_2)\n    axes[0].set_title('Before Handleling')\n    \n    mean_m2 = np.mean(df[feature_1])\n    \n    std_m2 = np.std(df[feature_1])\n        \n    df['z'] = (df[feature_1] - mean_m2)/std_m2\n    \n    df['z_filter'] = np.where(df['z'] >= threshold, 1,0)\n    \n    df.drop(df[train_data['z_filter'] == 1].index, inplace = True)\n    \n    df.drop('z', axis = 1, inplace = True)\n    df.drop('z_filter', axis = 1, inplace = True)\n    \n    sns.scatterplot(ax=axes[1], data=df, x=feature_1, y=feature_2)\n    axes[1].set_title('After Handleling')\n    plt.show()\n\n#%%\n\noutliers_check(train_data, 'Fare', 'Survived', 2)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:59.020231Z","iopub.execute_input":"2022-07-25T17:02:59.020685Z","iopub.status.idle":"2022-07-25T17:02:59.446743Z","shell.execute_reply.started":"2022-07-25T17:02:59.020651Z","shell.execute_reply":"2022-07-25T17:02:59.445373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''\nNormalizing some numeric data\n'''\n\nx_number = train_data[['Fare']]\n\nmin_max_scaler = preprocessing.MinMaxScaler()\nx_scaled = min_max_scaler.fit_transform(x_number)\n\ntrain_data[['Fare']] = x_scaled\n","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:59.449455Z","iopub.execute_input":"2022-07-25T17:02:59.449847Z","iopub.status.idle":"2022-07-25T17:02:59.462076Z","shell.execute_reply.started":"2022-07-25T17:02:59.449813Z","shell.execute_reply":"2022-07-25T17:02:59.460716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Separating ID","metadata":{}},{"cell_type":"code","source":"pass_id = train_data.pop('PassengerId')","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:59.463578Z","iopub.execute_input":"2022-07-25T17:02:59.463954Z","iopub.status.idle":"2022-07-25T17:02:59.474997Z","shell.execute_reply.started":"2022-07-25T17:02:59.463922Z","shell.execute_reply":"2022-07-25T17:02:59.473810Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Correlation check","metadata":{}},{"cell_type":"code","source":"def correlation(df, threshold):\n    col_corr = set()\n    corr_matrix = train_data.corr()\n    for i in range(len(corr_matrix.columns)):\n        for j in range(i):\n            if abs(corr_matrix.iloc[i, j]) > threshold:\n                colname = corr_matrix.columns[i]\n                col_corr.add(colname)\n    return col_corr\n\n\ncorr_try = correlation(train_data, 0.8)\nprint(corr_try)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:59.491727Z","iopub.execute_input":"2022-07-25T17:02:59.492607Z","iopub.status.idle":"2022-07-25T17:02:59.518833Z","shell.execute_reply.started":"2022-07-25T17:02:59.492568Z","shell.execute_reply":"2022-07-25T17:02:59.515903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''\nWe droped variables with high correlation.\n'''\ntrain_data.drop(corr_try, axis = 1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:59.682119Z","iopub.execute_input":"2022-07-25T17:02:59.682597Z","iopub.status.idle":"2022-07-25T17:02:59.691631Z","shell.execute_reply.started":"2022-07-25T17:02:59.682563Z","shell.execute_reply":"2022-07-25T17:02:59.690055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.rename(columns={'Age_(13.683, 26.947]': '13-27', 'Age_(26.947, 40.21]': '27-40',\n                           'Age_(40.21, 53.473]':'40-53', 'Age_(53.473, 66.737]':'53-67',\n                           'Age_(66.737, 80.0]': '67-80'}, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:59.796075Z","iopub.execute_input":"2022-07-25T17:02:59.797553Z","iopub.status.idle":"2022-07-25T17:02:59.805069Z","shell.execute_reply.started":"2022-07-25T17:02:59.797511Z","shell.execute_reply":"2022-07-25T17:02:59.803623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:02:59.920213Z","iopub.execute_input":"2022-07-25T17:02:59.920680Z","iopub.status.idle":"2022-07-25T17:02:59.951638Z","shell.execute_reply.started":"2022-07-25T17:02:59.920646Z","shell.execute_reply":"2022-07-25T17:02:59.950525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train & Test split","metadata":{}},{"cell_type":"code","source":"target = train_data.pop('Survived')\n\nX_train, X_test, y_train, y_test = train_test_split(train_data, target, test_size = 0.2, random_state = 22)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:00.044161Z","iopub.execute_input":"2022-07-25T17:03:00.045000Z","iopub.status.idle":"2022-07-25T17:03:00.053855Z","shell.execute_reply.started":"2022-07-25T17:03:00.044963Z","shell.execute_reply":"2022-07-25T17:03:00.052530Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:00.167740Z","iopub.execute_input":"2022-07-25T17:03:00.169193Z","iopub.status.idle":"2022-07-25T17:03:00.196666Z","shell.execute_reply.started":"2022-07-25T17:03:00.169150Z","shell.execute_reply":"2022-07-25T17:03:00.195396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_train","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:00.284393Z","iopub.execute_input":"2022-07-25T17:03:00.284834Z","iopub.status.idle":"2022-07-25T17:03:00.293678Z","shell.execute_reply.started":"2022-07-25T17:03:00.284800Z","shell.execute_reply":"2022-07-25T17:03:00.292416Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Balancing the Data","metadata":{}},{"cell_type":"code","source":"def undersampling(X_train, y_train, plus=1):\n\n    X_train['y'] = y_train\n    \n    minor = X_train[X_train['y']==1]\n    other = X_train[X_train['y']==0]\n    other = other.sample(n=int(len(minor)*plus), random_state=22)\n    X_train = (pd.concat([minor,other],axis=0)).sample(frac=1)\n    \n    y_train = X_train['y']\n    X_train.drop('y', axis=1,inplace=True)\n\n    return X_train, y_train\n\n\nX_train, y_train = undersampling(X_train, y_train, 1.1)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:00.388012Z","iopub.execute_input":"2022-07-25T17:03:00.388476Z","iopub.status.idle":"2022-07-25T17:03:00.407589Z","shell.execute_reply.started":"2022-07-25T17:03:00.388441Z","shell.execute_reply":"2022-07-25T17:03:00.406290Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(y_train.value_counts())","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:00.471983Z","iopub.execute_input":"2022-07-25T17:03:00.473151Z","iopub.status.idle":"2022-07-25T17:03:00.480533Z","shell.execute_reply.started":"2022-07-25T17:03:00.473109Z","shell.execute_reply":"2022-07-25T17:03:00.479276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:00.521341Z","iopub.execute_input":"2022-07-25T17:03:00.522549Z","iopub.status.idle":"2022-07-25T17:03:00.551152Z","shell.execute_reply.started":"2022-07-25T17:03:00.522506Z","shell.execute_reply":"2022-07-25T17:03:00.549810Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model Creation\nFirst we are going to use XGBoost Classifier to predict if the passengers survived or not. \nAfter obtaining a result with a basic model, we will search for the best sutable hyperparameters and seak for a better model.","metadata":{}},{"cell_type":"code","source":"#Model without parameters\nno_params_model = xgb.XGBClassifier()\n#Fitting the data to the model\nno_params_model.fit(X_train, y_train)\n#Predicting with the trained model\npreds = no_params_model.predict(X_test)\n#Printing the acuracy of the model\nprint(f'Predictions accuracy score: {accuracy_score(y_test,preds)}')","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:00.574637Z","iopub.execute_input":"2022-07-25T17:03:00.575097Z","iopub.status.idle":"2022-07-25T17:03:00.952306Z","shell.execute_reply.started":"2022-07-25T17:03:00.575062Z","shell.execute_reply":"2022-07-25T17:03:00.951171Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Hyperparameter search (Randomized Search)","metadata":{}},{"cell_type":"code","source":"#Parameters\n'''\nparams_XGB = {\n        'min_child_weight': [1, 5, 10],\n        'gamma': [0, 0.01, 0.5, 1, 1.5, 2],\n        'subsample': [0.6, 0.8, 1.0],\n        'learning_rate': [0.01, 0.02, 0.05, 0.1, 0.3],\n        'max_depth': [3, 4, 5, 6, 7],\n        'n_estimators': [100, 200, 400, 600, 800, 1000],\n        }\n\nparams_model_xgb = xgb.XGBClassifier()\n\n#Randomized Search\nRS_xgb = RandomizedSearchCV(params_model_xgb, param_distributions=params_XGB)\nRS_xgb.fit(X_train, y_train)\n\nprint(f'Best params for XGBoost Classifier using Random Search: {RS_xgb.best_params_}')\n'''","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:00.954293Z","iopub.execute_input":"2022-07-25T17:03:00.954668Z","iopub.status.idle":"2022-07-25T17:03:00.962618Z","shell.execute_reply.started":"2022-07-25T17:03:00.954636Z","shell.execute_reply":"2022-07-25T17:03:00.961600Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Hyperparameters\nbest_params = {'subsample': 1.0, 'n_estimators': 400, 'min_child_weight': 1, 'max_depth': 3, 'learning_rate': 0.02, 'gamma': 0}\n#Model with hyperparameters\nmodel = xgb.XGBClassifier(**best_params, random_state = 22)\n#Fitting the data to the model\nmodel.fit(X_train, y_train)\n#Predicting with the trained model\npreds = model.predict(X_test)\n#Printing the acuracy of the model\nprint(f'Predictions Accuracy Score: {accuracy_score(y_test,preds)}')\n#Printing the F1 score of the model\nf1 = f1_score(y_test, preds, average='macro')\nprint(f'Predictions F1 score: {f1}')\n#Printing the Recall score of the model\nrecall = recall_score(y_test, preds)\nprint('Predictions Recall Score: %f' % recall)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:00.964068Z","iopub.execute_input":"2022-07-25T17:03:00.964868Z","iopub.status.idle":"2022-07-25T17:03:02.282044Z","shell.execute_reply.started":"2022-07-25T17:03:00.964830Z","shell.execute_reply":"2022-07-25T17:03:02.280665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Feature Importance","metadata":{}},{"cell_type":"code","source":"FeatImp_xgb = pd.DataFrame(data = {'Cols': X_train.columns.tolist(),'FeatImp':model.feature_importances_})\nFeatImp_xgb.sort_values(by='FeatImp', inplace=True, ascending=False)\nFeatImp_xgb.reset_index(inplace = True, drop = True)\nsns.barplot(x='FeatImp', y='Cols', data=FeatImp_xgb.head(20))\nplt.title('Feature Importance XGB')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:02.284730Z","iopub.execute_input":"2022-07-25T17:03:02.285148Z","iopub.status.idle":"2022-07-25T17:03:03.133918Z","shell.execute_reply.started":"2022-07-25T17:03:02.285115Z","shell.execute_reply":"2022-07-25T17:03:03.132533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Test: Prediction\nIn this part of the code we are going to apply all the changes and tranformation made in the training set to the test set.","metadata":{}},{"cell_type":"code","source":"test_data.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:03.135767Z","iopub.execute_input":"2022-07-25T17:03:03.136145Z","iopub.status.idle":"2022-07-25T17:03:03.154922Z","shell.execute_reply.started":"2022-07-25T17:03:03.136105Z","shell.execute_reply":"2022-07-25T17:03:03.153408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''\nExtracting Status\n'''\n\ntest_data['Status'] = test_data['Name'].str.extract(' ([A-Za-z]+)\\.', expand=False)\n\ntest_data['Status'] = test_data['Status'].replace(['Lady', 'Countess', 'Don', 'Sir', 'Jonkheer', 'Dona'], 'Royal')\n\ntest_data['Status'] = test_data['Status'].replace(['Capt', 'Col','Major', 'Rev', 'Dr'], 'Officers')\n\ntest_data['Status'] = test_data['Status'].replace('Mlle', 'Miss')\ntest_data['Status'] = test_data['Status'].replace('Ms', 'Miss')\ntest_data['Status'] = test_data['Status'].replace('Mme', 'Mrs')\n\nprint(test_data['Status'].value_counts())","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:03.156707Z","iopub.execute_input":"2022-07-25T17:03:03.157227Z","iopub.status.idle":"2022-07-25T17:03:03.181636Z","shell.execute_reply.started":"2022-07-25T17:03:03.157175Z","shell.execute_reply":"2022-07-25T17:03:03.179967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"columns = test_data.columns\nfor column in columns:\n    has_nan = test_data[column].isnull().sum(axis = 0)\n    print(f'{column}: {has_nan}')\n    \nprint(test_data.isnull().values.any())","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:03.183587Z","iopub.execute_input":"2022-07-25T17:03:03.184886Z","iopub.status.idle":"2022-07-25T17:03:03.199638Z","shell.execute_reply.started":"2022-07-25T17:03:03.184845Z","shell.execute_reply":"2022-07-25T17:03:03.197987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''\nNan values\n'''\ntest_data['Age'] = np.where((test_data['Status'] == 'Master') & (test_data['Age'].isnull()), test_data['Age'].fillna(4.5), test_data['Age'])\n\ntest_data['Age'] = np.where((test_data['Status'] == 'Miss') & (test_data['Age'].isnull()), test_data['Age'].fillna(21.8), test_data['Age'])\n\ntest_data['Age'] = np.where((test_data['Status'] == 'Mr') & (test_data['Age'].isnull()), test_data['Age'].fillna(32.3), test_data['Age'])\n\ntest_data['Age'] = np.where((test_data['Status'] == 'Mrs') & (test_data['Age'].isnull()), test_data['Age'].fillna(35.8), test_data['Age'])\n\ntest_data['Age'] = np.where((test_data['Status'] == 'Officers') & (test_data['Age'].isnull()), test_data['Age'].fillna(47.5), test_data['Age'])\n\ntest_data['Age'] = np.where((test_data['Status'] == 'Royal') & (test_data['Age'].isnull()), test_data['Age'].fillna(41), test_data['Age'])\n","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:03.207074Z","iopub.execute_input":"2022-07-25T17:03:03.207697Z","iopub.status.idle":"2022-07-25T17:03:03.227963Z","shell.execute_reply.started":"2022-07-25T17:03:03.207649Z","shell.execute_reply":"2022-07-25T17:03:03.226746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.groupby(['Pclass','Sex']).Age.mean()","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:03.229632Z","iopub.execute_input":"2022-07-25T17:03:03.229954Z","iopub.status.idle":"2022-07-25T17:03:03.244174Z","shell.execute_reply.started":"2022-07-25T17:03:03.229923Z","shell.execute_reply":"2022-07-25T17:03:03.242735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.loc[test_data['Fare'].isnull()]\n\ntest_data['Fare'] = test_data['Fare'].fillna(26)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:03.246163Z","iopub.execute_input":"2022-07-25T17:03:03.246785Z","iopub.status.idle":"2022-07-25T17:03:03.259693Z","shell.execute_reply.started":"2022-07-25T17:03:03.246748Z","shell.execute_reply":"2022-07-25T17:03:03.258365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data['Has_cabin'] = np.where(test_data['Cabin'].isnull(), 0, 1)\ntest_data.drop('Cabin', axis = 1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:03.261373Z","iopub.execute_input":"2022-07-25T17:03:03.261864Z","iopub.status.idle":"2022-07-25T17:03:03.271431Z","shell.execute_reply.started":"2022-07-25T17:03:03.261829Z","shell.execute_reply":"2022-07-25T17:03:03.270449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data['Is_child'] = np.where(test_data['Age'] <= 12, 1, 0)\n\ntest_data['Sex'] = np.where(test_data['Sex'] == 'male', 1,0)\n\ntest_data.drop('Name', axis = 1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:03.273289Z","iopub.execute_input":"2022-07-25T17:03:03.274073Z","iopub.status.idle":"2022-07-25T17:03:03.288221Z","shell.execute_reply.started":"2022-07-25T17:03:03.274022Z","shell.execute_reply":"2022-07-25T17:03:03.286712Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.drop('Ticket', axis = 1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:03.289862Z","iopub.execute_input":"2022-07-25T17:03:03.290197Z","iopub.status.idle":"2022-07-25T17:03:03.302701Z","shell.execute_reply.started":"2022-07-25T17:03:03.290166Z","shell.execute_reply":"2022-07-25T17:03:03.301253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data['Fam_size'] = test_data['SibSp'] + test_data['Parch']\n\ntest_data['Solo'] = np.where(test_data['Fam_size'] == 0, 1, 0)\n\ntest_data.drop(['SibSp', 'Parch'], axis = 1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:03.304530Z","iopub.execute_input":"2022-07-25T17:03:03.305129Z","iopub.status.idle":"2022-07-25T17:03:03.321027Z","shell.execute_reply.started":"2022-07-25T17:03:03.305042Z","shell.execute_reply":"2022-07-25T17:03:03.319456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data['Fare'] = test_data['Fare'] / (test_data['Fam_size'] + 1)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:03.322633Z","iopub.execute_input":"2022-07-25T17:03:03.322971Z","iopub.status.idle":"2022-07-25T17:03:03.333657Z","shell.execute_reply.started":"2022-07-25T17:03:03.322938Z","shell.execute_reply":"2022-07-25T17:03:03.332617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data['0-13'] = np.where(test_data['Age'] < 13, 1, 0)\ntest_data['13-27'] = np.where((test_data['Age'] <= 13) & (test_data['Age'] > 27), 1, 0)\ntest_data['27-40'] = np.where((test_data['Age'] <= 27) & (test_data['Age'] > 40), 1, 0)\ntest_data['40-53'] = np.where((test_data['Age'] >= 40) & (test_data['Age'] < 53), 1, 0)\ntest_data['53-67'] = np.where((test_data['Age'] >= 53) & (test_data['Age'] < 67), 1, 0)\ntest_data['67-80'] = np.where((test_data['Age'] >= 67) & (test_data['Age'] <= 80), 1, 0)\n\ntest_data.drop('Age', axis = 1, inplace = True)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:03.335472Z","iopub.execute_input":"2022-07-25T17:03:03.336103Z","iopub.status.idle":"2022-07-25T17:03:03.356454Z","shell.execute_reply.started":"2022-07-25T17:03:03.336069Z","shell.execute_reply":"2022-07-25T17:03:03.355139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data = createDummies(test_data, 'Status')\n\ntest_data = createDummies(test_data, 'Embarked')\n","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:03.358495Z","iopub.execute_input":"2022-07-25T17:03:03.359231Z","iopub.status.idle":"2022-07-25T17:03:03.379361Z","shell.execute_reply.started":"2022-07-25T17:03:03.359175Z","shell.execute_reply":"2022-07-25T17:03:03.378005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:03.381081Z","iopub.execute_input":"2022-07-25T17:03:03.381520Z","iopub.status.idle":"2022-07-25T17:03:03.413545Z","shell.execute_reply.started":"2022-07-25T17:03:03.381486Z","shell.execute_reply":"2022-07-25T17:03:03.412361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.drop(['0-13', 'Status_Mr'], axis = 1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:03.415234Z","iopub.execute_input":"2022-07-25T17:03:03.415596Z","iopub.status.idle":"2022-07-25T17:03:03.426288Z","shell.execute_reply.started":"2022-07-25T17:03:03.415565Z","shell.execute_reply":"2022-07-25T17:03:03.424843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pass_id_test = test_data.pop('PassengerId')","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:03.433574Z","iopub.execute_input":"2022-07-25T17:03:03.434052Z","iopub.status.idle":"2022-07-25T17:03:03.441970Z","shell.execute_reply.started":"2022-07-25T17:03:03.434016Z","shell.execute_reply":"2022-07-25T17:03:03.440320Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x_scaled_test = min_max_scaler.transform(test_data[['Fare']])\n\ntest_data[['Fare']] = x_scaled_test\n","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:03.446958Z","iopub.execute_input":"2022-07-25T17:03:03.447553Z","iopub.status.idle":"2022-07-25T17:03:03.460248Z","shell.execute_reply.started":"2022-07-25T17:03:03.447515Z","shell.execute_reply":"2022-07-25T17:03:03.458991Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:03:03.462598Z","iopub.execute_input":"2022-07-25T17:03:03.463242Z","iopub.status.idle":"2022-07-25T17:03:03.496714Z","shell.execute_reply.started":"2022-07-25T17:03:03.463191Z","shell.execute_reply":"2022-07-25T17:03:03.495455Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission","metadata":{}},{"cell_type":"code","source":"test_submission = model.predict(test_data)\n\nsubmission = pd.DataFrame({'PassengerId':pass_id_test, 'Survived':test_submission})","metadata":{"execution":{"iopub.status.busy":"2022-07-25T17:06:13.350865Z","iopub.execute_input":"2022-07-25T17:06:13.351876Z","iopub.status.idle":"2022-07-25T17:06:13.365889Z","shell.execute_reply.started":"2022-07-25T17:06:13.351830Z","shell.execute_reply":"2022-07-25T17:06:13.364822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv(\"submission.csv\", index = False)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission\n","metadata":{},"execution_count":null,"outputs":[]}]}