{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# SS Titanic Modeling\n\nby David Zambrano\n\n<img width=\"300\" src=\"https://cdn.drawception.com/drawings/KO3D1hhjOk.png\">\n\ncredits to the author of the image.\n\n## **Competition context:**\n\nWelcome to the year 2912, where your data science skills are needed to solve a cosmic mystery. We've received a transmission from four lightyears away and things aren't looking good.\n\nThe Spaceship Titanic was an interstellar passenger liner launched a month ago. With almost 13,000 passengers on board, the vessel set out on its maiden voyage transporting emigrants from our solar system to three newly habitable exoplanets orbiting nearby stars.\n\nWhile rounding Alpha Centauri en route to its first destination—the torrid 55 Cancri E—the unwary Spaceship Titanic collided with a spacetime anomaly hidden within a dust cloud. Sadly, it met a similar fate as its namesake from 1000 years before. Though the ship stayed intact, almost half of the passengers were transported to an alternate dimension!\n\nTo help rescue crews and retrieve the lost passengers, you are challenged to predict which passengers were transported by the anomaly using records recovered from the spaceship’s damaged computer system.\n\nFor more information about this competition click [here](https://www.kaggle.com/competitions/spaceship-titanic).","metadata":{}},{"cell_type":"markdown","source":"## Data Description:\n\n**File name:** train.csv - Contains personal records for about two-thirds (~8700) of the passengers, to be used as training data.\n\n\n> **PassengerId** - A unique Id for each passenger. Each Id takes the form gggg_pp where gggg indicates a group the passenger is travelling with and pp is their number within the group. People in a group are often family members, but not always.\n\n> **HomePlanet** - The planet the passenger departed from, typically their planet of permanent residence.\n\n> **CryoSleep** - Indicates whether the passenger elected to be put into suspended animation for the duration of the voyage. Passengers in cryosleep are confined to their cabins.\n\n> **Cabin** - The cabin number where the passenger is staying. Takes the form deck/num/side, where side can be either P for Port or S for Starboard.\n\n> **Destination** - The planet the passenger will be debarking to.\n\n> **Age** - The age of the passenger.\n\n> **VIP** - Whether the passenger has paid for special VIP service during the voyage.\n\n> **RoomService, FoodCourt, ShoppingMall, Spa, VRDeck** - Amount the passenger has billed at each of the Spaceship Titanic's many luxury amenities.\n\n> **Name** - The first and last names of the passenger.\n\n> **Transported** - Whether the passenger was transported to another dimension. This is the target, the column you are trying to predict.","metadata":{}},{"cell_type":"markdown","source":"To review **Exploratory Data Analysis** click [here](https://www.kaggle.com/code/davidzambrano87/ss-titanic-eda-by-dz).\n\nTo review **Feature Engineering** click [here](https://www.kaggle.com/code/davidzambrano87/ss-titanic-fe-by-dz).","metadata":{}},{"cell_type":"markdown","source":"### Train Dataset:","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\ntrain_dataset = pd.read_csv('../input/spaceship-titanic/train.csv')\ntrain_dataset.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:41:51.404953Z","iopub.execute_input":"2022-08-11T18:41:51.405530Z","iopub.status.idle":"2022-08-11T18:41:51.517855Z","shell.execute_reply.started":"2022-08-11T18:41:51.405399Z","shell.execute_reply":"2022-08-11T18:41:51.516626Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Feature Engineering Implementation:","metadata":{}},{"cell_type":"code","source":"import numpy as np\n\n# Function to fill nan values\ndef fill_nan(dataset):\n    \n    fill_with_mean = ['Age','RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck',]\n    fill_with_mode = ['HomePlanet', 'CryoSleep', 'Destination', 'VIP']\n    fill_with_unkn = ['Cabin', 'Name']\n        \n    new_dataset = dataset.copy()\n    \n    for var in fill_with_mean:\n        new_dataset[var] = new_dataset[var].fillna(new_dataset[var].mean())   \n    \n    for var in fill_with_mode:\n        new_dataset[var] = new_dataset[var].fillna(new_dataset[var].mode().iloc[0])\n    \n    for var in fill_with_unkn:\n        new_dataset[var] = new_dataset[var].fillna('unkn/un kn/unkn')\n    \n    return new_dataset\n\n\n# Function to create PassangerGroup:\ndef passenger_group(dataset):\n    \n    new_dataset = dataset.copy()\n    new_dataset['PassengerGroup'] = new_dataset['PassengerId'].str.split('_',expand=True)[0].astype('float')\n    \n    return new_dataset\n\n\n# Function to identify number of members per family:\ndef family_size(dataset):\n    \n    new_dataset = dataset.copy()\n    new_dataset['FamilyName'] = new_dataset['Name'].str.split(' ',expand=True)[1]\n    family_name_size = new_dataset['FamilyName'].value_counts().to_dict()\n    new_dataset['FamilySize'] = new_dataset['FamilyName'].map(family_name_size)\n    \n    return new_dataset\n\n\n# Function to create Cabin-Deck-Side:\ndef cabin_deck_side(dataset):\n    \n    new_dataset = dataset.copy()\n    new_dataset['L_Cabin'] = new_dataset['Cabin'].str.split('/',expand=True)[0]\n    new_dataset['Side'] = new_dataset['Cabin'].str.split('/',expand=True)[2]\n    \n    return new_dataset\n\n\n# Function to estimate log values\ndef log_transform(dataset):\n\n    vars_to_transform = ['Age', 'RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']\n\n    new_dataset = dataset.copy()\n    \n    for var in vars_to_transform:\n        new_label = 'log_' + var\n        new_dataset[new_label] = np.log(new_dataset[var] + 1)\n        new_dataset[new_label] = new_dataset[new_label] / new_dataset[new_label].max()\n        \n    return new_dataset\n\n\n# Function to create new features:\ndef new_features(dataset):\n    \n    new_dataset = dataset.copy()\n    \n    new_dataset[\"exp_1\"] = (new_dataset[\"log_RoomService\"] > new_dataset[\"log_VRDeck\"]) | (new_dataset[\"log_RoomService\"] > new_dataset[\"log_Spa\"])\n    new_dataset[\"exp_2\"] = ~(new_dataset[\"log_RoomService\"] > new_dataset[\"log_VRDeck\"]) | ~(new_dataset[\"log_RoomService\"] > new_dataset[\"log_Spa\"])\n    new_dataset[\"exp_3\"] = (new_dataset[\"FoodCourt\"] > new_dataset[\"VRDeck\"]) | ~(new_dataset[\"RoomService\"] > new_dataset[\"VRDeck\"])\n    new_dataset[\"exp_4\"] = (new_dataset[\"FoodCourt\"] > new_dataset[\"Spa\"]) | ~(new_dataset[\"RoomService\"] > new_dataset[\"Spa\"])\n\n    for exp_var in new_dataset.columns:\n        if exp_var.startswith('exp') == True:\n            new_dataset[exp_var] = new_dataset[exp_var].astype(float)\n    \n    return new_dataset\n\n\n# Function for one-hot-encoding:\ndef one_hot_encode(dataset):\n    \n    new_dataset = dataset.copy()\n    \n    variables_to_encode = ['HomePlanet', 'Destination', 'L_Cabin', 'Side', 'FamilySize']\n    \n    for var in variables_to_encode:\n        cat_labels = list(new_dataset[var].unique())\n        dummy_label = var[:4]\n        dummy = pd.get_dummies(new_dataset[var], prefix = dummy_label, drop_first = True)\n        new_dataset = pd.merge(left = new_dataset, right = dummy, left_index = True, right_index = True)\n        \n        for cat in cat_labels:\n            column_name = dummy_label + '_' + str(cat)\n            try:\n                new_dataset[column_name] = new_dataset[column_name].astype(float)\n            except:\n                continue\n        \n    return new_dataset\n\n\n# Function to remove unused variables:\ndef remove_variables(dataset):\n    \n    new_dataset = dataset.copy()\n    \n    variables_to_drop = ['Age', 'RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck', \n                         'HomePlanet', 'Destination', 'Cabin', 'Name', 'L_Cabin', 'Side', \n                         'FamilyName', 'FamilySize']\n    \n    try:\n        new_dataset = new_dataset.drop(variables_to_drop, axis = 1)    \n    except:\n        new_dataset = new_dataset[variables_to_keep[:-1]]\n    \n    return new_dataset\n\n\n# Function to fix data types:\ndef fix_data_types(dataset):\n    \n    new_dataset = dataset.copy()\n    \n    vars_to_fix_data_type = ['CryoSleep', 'VIP']\n    \n    for var in vars_to_fix_data_type:\n        new_dataset[var] = new_dataset[var].astype(float)\n    \n    return new_dataset\n\n\n# Function to apply all the feature engineering to a dataset:\ndef apply_fe(dataset):\n    \n    new_dataset = dataset.copy()\n    new_dataset = fill_nan(new_dataset)\n    #new_dataset = passenger_group(new_dataset)\n    new_dataset = family_size(new_dataset)\n    new_dataset = cabin_deck_side(new_dataset)\n    new_dataset = log_transform(new_dataset)\n    new_dataset = new_features(new_dataset)\n    new_dataset = one_hot_encode(new_dataset)\n    new_dataset = remove_variables(new_dataset)\n    new_dataset = fix_data_types(new_dataset)\n    \n    return new_dataset","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:41:55.779793Z","iopub.execute_input":"2022-08-11T18:41:55.780166Z","iopub.status.idle":"2022-08-11T18:41:55.813329Z","shell.execute_reply.started":"2022-08-11T18:41:55.780136Z","shell.execute_reply":"2022-08-11T18:41:55.812384Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_dataset_clean = train_dataset.dropna()\ntrain_dataset_fe = apply_fe(train_dataset_clean)\ntrain_dataset_fe.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:42:00.404747Z","iopub.execute_input":"2022-08-11T18:42:00.405233Z","iopub.status.idle":"2022-08-11T18:42:00.573369Z","shell.execute_reply.started":"2022-08-11T18:42:00.405190Z","shell.execute_reply":"2022-08-11T18:42:00.572141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\n\nfe_vars = ['log_Age', 'log_RoomService', 'log_FoodCourt', 'log_ShoppingMall', 'log_Spa', 'log_VRDeck', \n           'exp_1', 'exp_2', 'exp_3', 'exp_4', 'VIP', 'CryoSleep', 'Transported']\n\nfig, axes = plt.subplots(1, 2, sharex=False, sharey=False, figsize=(17,7))\nfig.suptitle('original vs new features correlations')\nsns.heatmap(train_dataset.corr(), annot = True, cmap = 'coolwarm', ax = axes[0], fmt = \".2f\")\nsns.heatmap(train_dataset_fe[fe_vars].corr(), annot = True, cmap = 'coolwarm', ax = axes[1], fmt = \".2f\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:42:07.923642Z","iopub.execute_input":"2022-08-11T18:42:07.924037Z","iopub.status.idle":"2022-08-11T18:42:10.101801Z","shell.execute_reply.started":"2022-08-11T18:42:07.924004Z","shell.execute_reply":"2022-08-11T18:42:10.100654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"yeah! new features have higher correlation with target variable than the original ones.","metadata":{}},{"cell_type":"markdown","source":"### Train/Dev split dataset:\n\n","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\nX = train_dataset_fe.drop('Transported', axis = 1)\ny = train_dataset_fe['Transported']\n\nX_train, X_dev, y_train, y_dev = train_test_split(X, y, test_size=0.20, random_state=1)\nlen(X_dev), len(X_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:42:22.556220Z","iopub.execute_input":"2022-08-11T18:42:22.556637Z","iopub.status.idle":"2022-08-11T18:42:22.686612Z","shell.execute_reply.started":"2022-08-11T18:42:22.556605Z","shell.execute_reply":"2022-08-11T18:42:22.685766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:42:25.428733Z","iopub.execute_input":"2022-08-11T18:42:25.429146Z","iopub.status.idle":"2022-08-11T18:42:25.463042Z","shell.execute_reply.started":"2022-08-11T18:42:25.429111Z","shell.execute_reply":"2022-08-11T18:42:25.461680Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_dev.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:42:29.236559Z","iopub.execute_input":"2022-08-11T18:42:29.237299Z","iopub.status.idle":"2022-08-11T18:42:29.268194Z","shell.execute_reply.started":"2022-08-11T18:42:29.237231Z","shell.execute_reply":"2022-08-11T18:42:29.267064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_train = y_train.astype(float)\ny_dev = y_dev.astype(float)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:42:33.509494Z","iopub.execute_input":"2022-08-11T18:42:33.510474Z","iopub.status.idle":"2022-08-11T18:42:33.516102Z","shell.execute_reply.started":"2022-08-11T18:42:33.510434Z","shell.execute_reply":"2022-08-11T18:42:33.514704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Are X_dev and test_dataset balanced?","metadata":{}},{"cell_type":"code","source":"test_dataset = pd.read_csv('../input/spaceship-titanic/test.csv')\ntest_dataset.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:42:37.193345Z","iopub.execute_input":"2022-08-11T18:42:37.193736Z","iopub.status.idle":"2022-08-11T18:42:37.240571Z","shell.execute_reply.started":"2022-08-11T18:42:37.193703Z","shell.execute_reply":"2022-08-11T18:42:37.239237Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def apply_fe(dataset):\n    \n    new_dataset = dataset.copy()\n    new_dataset = fill_nan(new_dataset)\n    #new_dataset = passenger_group(new_dataset)\n    new_dataset = family_size(new_dataset)\n    new_dataset = cabin_deck_side(new_dataset)\n    new_dataset = log_transform(new_dataset)\n    new_dataset = new_features(new_dataset)\n    new_dataset = one_hot_encode(new_dataset)\n    new_dataset = fix_data_types(new_dataset)\n    new_dataset = remove_variables(new_dataset)\n    \n    return new_dataset\n\nX_test = apply_fe(test_dataset)\nX_test.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:42:43.213867Z","iopub.execute_input":"2022-08-11T18:42:43.214292Z","iopub.status.idle":"2022-08-11T18:42:43.332224Z","shell.execute_reply.started":"2022-08-11T18:42:43.214244Z","shell.execute_reply":"2022-08-11T18:42:43.331144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from scipy.stats import ttest_ind\n\ndef density_comparison(var1):\n    stat, p = ttest_ind(X_dev[var1], X_test[var1])\n    print(\"p-value for identical distribution:\", p)\n    sns.kdeplot(np.log(X_dev[var1] + 1), shade=True, color=\"r\")\n    sns.kdeplot(np.log(X_test[var1] + 1), shade=True, color=\"b\")\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:42:49.407522Z","iopub.execute_input":"2022-08-11T18:42:49.408656Z","iopub.status.idle":"2022-08-11T18:42:49.415597Z","shell.execute_reply.started":"2022-08-11T18:42:49.408609Z","shell.execute_reply":"2022-08-11T18:42:49.414713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for var in X_dev.columns:\n    try:\n        density_comparison(var)\n    except:\n        continue","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:42:58.886430Z","iopub.execute_input":"2022-08-11T18:42:58.887190Z","iopub.status.idle":"2022-08-11T18:43:06.654087Z","shell.execute_reply.started":"2022-08-11T18:42:58.887150Z","shell.execute_reply":"2022-08-11T18:43:06.653053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can go with the hipothesys that both datasets **X_dev** and **X_test** share the same distribution for the inspected variables.","metadata":{}},{"cell_type":"markdown","source":"## Models:","metadata":{}},{"cell_type":"markdown","source":"### 1. Logit witht statsmodels:","metadata":{}},{"cell_type":"code","source":"import statsmodels.api as sm\n\nlogit_model = sm.Logit(y_train, X_train.iloc[:,1:])\nlogit_results = logit_model.fit(method = 'newton')\nlogit_results.summary()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:43:34.676882Z","iopub.execute_input":"2022-08-11T18:43:34.677457Z","iopub.status.idle":"2022-08-11T18:43:35.631848Z","shell.execute_reply.started":"2022-08-11T18:43:34.677398Z","shell.execute_reply":"2022-08-11T18:43:35.630560Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It is noticeable that Home and Destination variables are not significant for the model, and so might be the case for our variables exp_2 and VIP. The rest of the variables explain in different ways the decision of being transported or not.","metadata":{}},{"cell_type":"code","source":"def confusion_matrix(model, X_dataset, y_dataset):\n    \n    temp_df = pd.concat([pd.DataFrame((model.predict(X_dataset.iloc[:,1:]) >= 0.5).astype(int)), y_dataset.astype(int)], axis = 1)\n    temp_df.columns = ['predicted', 'actual']\n    confusion_matrix = pd.crosstab(temp_df['predicted'], temp_df['actual'])\n    \n    return confusion_matrix","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:43:43.188894Z","iopub.execute_input":"2022-08-11T18:43:43.189285Z","iopub.status.idle":"2022-08-11T18:43:43.196803Z","shell.execute_reply.started":"2022-08-11T18:43:43.189233Z","shell.execute_reply":"2022-08-11T18:43:43.195348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"confusion_matrix(logit_results, X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:43:46.872806Z","iopub.execute_input":"2022-08-11T18:43:46.873221Z","iopub.status.idle":"2022-08-11T18:43:46.909649Z","shell.execute_reply.started":"2022-08-11T18:43:46.873188Z","shell.execute_reply":"2022-08-11T18:43:46.908472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def accuracy_model(model, X_dataset, y_dataset):\n    \n    conf_matrix = confusion_matrix(model, X_dataset, y_dataset)    \n    accuracy = (conf_matrix[0][0] + conf_matrix[1][1]) / len(y_dataset)\n\n    return np.round(accuracy, 4)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:43:50.371343Z","iopub.execute_input":"2022-08-11T18:43:50.372099Z","iopub.status.idle":"2022-08-11T18:43:50.377961Z","shell.execute_reply.started":"2022-08-11T18:43:50.372060Z","shell.execute_reply":"2022-08-11T18:43:50.376685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"accuracy_model(logit_results, X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:43:54.709598Z","iopub.execute_input":"2022-08-11T18:43:54.709973Z","iopub.status.idle":"2022-08-11T18:43:54.738871Z","shell.execute_reply.started":"2022-08-11T18:43:54.709943Z","shell.execute_reply":"2022-08-11T18:43:54.737556Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There's a significant bias problem with the logit problem.","metadata":{}},{"cell_type":"markdown","source":"### Validating Model:","metadata":{}},{"cell_type":"code","source":"confusion_matrix(logit_results, X_dev, y_dev)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:43:59.003322Z","iopub.execute_input":"2022-08-11T18:43:59.003708Z","iopub.status.idle":"2022-08-11T18:43:59.036174Z","shell.execute_reply.started":"2022-08-11T18:43:59.003676Z","shell.execute_reply":"2022-08-11T18:43:59.034973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"accuracy_model(logit_results, X_dev, y_dev)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:44:02.954642Z","iopub.execute_input":"2022-08-11T18:44:02.955381Z","iopub.status.idle":"2022-08-11T18:44:02.980773Z","shell.execute_reply.started":"2022-08-11T18:44:02.955344Z","shell.execute_reply":"2022-08-11T18:44:02.979665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Expanding dataset with interactions:","metadata":{}},{"cell_type":"code","source":"# Function to add interactions to dataset:\ndef interaction_exp(dataset):\n    \n    var_to_interact = ['VIP', 'exp_1', 'exp_2', 'exp_3', 'exp_4']\n    num_vars = ['log_RoomService', 'log_FoodCourt','log_ShoppingMall', 'log_Spa', 'log_VRDeck']\n\n    dataset_exp = dataset.copy()\n\n    for i in var_to_interact:\n        for j in num_vars:\n            new_label = 'int_' + i + '_' + j[4:7]\n            dataset_exp[new_label] = dataset_exp[i] * dataset_exp[j]\n            \n    return dataset_exp","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:48:04.711376Z","iopub.execute_input":"2022-08-11T18:48:04.711761Z","iopub.status.idle":"2022-08-11T18:48:04.718610Z","shell.execute_reply.started":"2022-08-11T18:48:04.711730Z","shell.execute_reply":"2022-08-11T18:48:04.717559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_exp = interaction_exp(X_train)\n\nlogit_model_interactions = sm.Logit(y_train, X_train_exp.iloc[:,1:])\nlogit_results_interactions = logit_model_interactions.fit()\nlogit_results_interactions.summary()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:48:06.020797Z","iopub.execute_input":"2022-08-11T18:48:06.021440Z","iopub.status.idle":"2022-08-11T18:48:06.232409Z","shell.execute_reply.started":"2022-08-11T18:48:06.021392Z","shell.execute_reply":"2022-08-11T18:48:06.231074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"confusion_matrix(logit_results_interactions, X_train_exp, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:48:12.794916Z","iopub.execute_input":"2022-08-11T18:48:12.795294Z","iopub.status.idle":"2022-08-11T18:48:12.829016Z","shell.execute_reply.started":"2022-08-11T18:48:12.795244Z","shell.execute_reply":"2022-08-11T18:48:12.827786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"accuracy_model(logit_results_interactions, X_train_exp, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:48:16.290499Z","iopub.execute_input":"2022-08-11T18:48:16.290915Z","iopub.status.idle":"2022-08-11T18:48:16.321654Z","shell.execute_reply.started":"2022-08-11T18:48:16.290873Z","shell.execute_reply":"2022-08-11T18:48:16.320600Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"interactions increased accuracy by 1%.","metadata":{}},{"cell_type":"markdown","source":"### Validating Expanded Model:","metadata":{}},{"cell_type":"code","source":"X_dev_exp = interaction_exp(X_dev)\n\nconfusion_matrix(logit_results_interactions, X_dev_exp, y_dev)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:48:43.541406Z","iopub.execute_input":"2022-08-11T18:48:43.541825Z","iopub.status.idle":"2022-08-11T18:48:43.590625Z","shell.execute_reply.started":"2022-08-11T18:48:43.541792Z","shell.execute_reply":"2022-08-11T18:48:43.589474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"accuracy_model(logit_results_interactions, X_dev_exp, y_dev)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:48:47.517661Z","iopub.execute_input":"2022-08-11T18:48:47.518052Z","iopub.status.idle":"2022-08-11T18:48:47.549282Z","shell.execute_reply.started":"2022-08-11T18:48:47.518019Z","shell.execute_reply":"2022-08-11T18:48:47.548098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Interactions increased accuracy on dev dataset by 1%. Notwithstanding, results of the logit model are not satisfying so far.","metadata":{}},{"cell_type":"markdown","source":"### 2. Naive Bayes:","metadata":{}},{"cell_type":"code","source":"from sklearn.naive_bayes import MultinomialNB\n\nmnb = MultinomialNB().fit(X_train_exp, y_train)\n\nprint(\"score on dev: \" + str(mnb.score(X_dev_exp, y_dev)))\nprint(\"score on train: \"+ str(mnb.score(X_train_exp, y_train)))","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:49:10.901199Z","iopub.execute_input":"2022-08-11T18:49:10.901586Z","iopub.status.idle":"2022-08-11T18:49:11.031196Z","shell.execute_reply.started":"2022-08-11T18:49:10.901554Z","shell.execute_reply":"2022-08-11T18:49:11.030046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This model perform worst than the logit.","metadata":{}},{"cell_type":"markdown","source":"### 3. LinearSVC:","metadata":{}},{"cell_type":"code","source":"from sklearn.svm import LinearSVC\n\nsvm=LinearSVC(C = 0.0001, max_iter = 10000)\nsvm.fit(X_train_exp.iloc[:,1:], y_train)\n\nprint(\"score on dev: \" + str(svm.score(X_dev_exp.iloc[:,1:], y_dev)))\nprint(\"score on train: \"+ str(svm.score(X_train_exp.iloc[:,1:], y_train)))","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:49:15.725972Z","iopub.execute_input":"2022-08-11T18:49:15.726365Z","iopub.status.idle":"2022-08-11T18:49:15.829742Z","shell.execute_reply.started":"2022-08-11T18:49:15.726333Z","shell.execute_reply":"2022-08-11T18:49:15.828569Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4. Decision Tree:","metadata":{}},{"cell_type":"code","source":"from sklearn.tree import DecisionTreeClassifier\n\nclf = DecisionTreeClassifier(max_depth = 50, criterion=\"entropy\", min_samples_leaf=50)\nclf.fit(X_train_exp.iloc[:,1:], y_train)\n\nprint(\"score on dev: \"  + str(clf.score(X_dev_exp.iloc[:,1:], y_dev)))\nprint(\"score on train: \" + str(clf.score(X_train_exp.iloc[:,1:], y_train)))","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:49:19.312483Z","iopub.execute_input":"2022-08-11T18:49:19.312880Z","iopub.status.idle":"2022-08-11T18:49:19.458147Z","shell.execute_reply.started":"2022-08-11T18:49:19.312847Z","shell.execute_reply":"2022-08-11T18:49:19.457294Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 5. Bagging Classifier:","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import BaggingClassifier\n\nbg=BaggingClassifier(DecisionTreeClassifier(),max_samples = 0.52, max_features = 0.8 ,n_estimators = 69, random_state = 50)\nbg.fit(X_train_exp.iloc[:,1:], y_train)\n\nprint(\"score on dev: \" + str(bg.score(X_dev_exp.iloc[:,1:], y_dev)))\nprint(\"score on train: \"+ str(bg.score(X_train_exp.iloc[:,1:], y_train)))","metadata":{"execution":{"iopub.status.busy":"2022-08-11T19:19:14.202149Z","iopub.execute_input":"2022-08-11T19:19:14.203243Z","iopub.status.idle":"2022-08-11T19:19:15.994822Z","shell.execute_reply.started":"2022-08-11T19:19:14.203188Z","shell.execute_reply":"2022-08-11T19:19:15.993692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.ensemble import BaggingClassifier\n\nbg=BaggingClassifier(DecisionTreeClassifier(), \n                     max_samples = 0.52, \n                     max_features = 0.8, \n                     n_estimators = 69,\n                     random_state = 50)\nbg.fit(X_train_exp.iloc[:,1:], y_train)\n\nprint(\"score on dev: \" + str(bg.score(X_dev_exp.iloc[:,1:], y_dev)))\nprint(\"score on train: \"+ str(bg.score(X_train_exp.iloc[:,1:], y_train)))","metadata":{"execution":{"iopub.status.busy":"2022-08-11T20:02:53.286417Z","iopub.execute_input":"2022-08-11T20:02:53.286920Z","iopub.status.idle":"2022-08-11T20:02:55.106904Z","shell.execute_reply.started":"2022-08-11T20:02:53.286872Z","shell.execute_reply":"2022-08-11T20:02:55.105767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_dataset_fe = apply_fe(test_dataset)\nX_test_exp = interaction_exp(test_dataset_fe)\n\nX_test_exp = X_test_exp.drop(['L_Ca_unkn', 'Side_unkn', 'Fami_14', 'Fami_94'], axis = 1)\nX_test_exp['Fami_12'] = 0\nX_test_exp['Fami_15'] = 0\nX_test_exp['Fami_17'] = 0\nX_test_exp = X_test_exp[X_train_exp.columns]\n\ny_pred_test = bg.predict(X_test_exp.iloc[:,1:])\npredictions_test = [round(value) for value in y_pred_test]\n\ntest_dataset['Transported'] = predictions_test\ntest_dataset['Transported'] = test_dataset['Transported'].astype(bool)\n\ntest_dataset.head()\n\nsubmission = test_dataset[['PassengerId', 'Transported']]\nsubmission.to_csv('submission.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T20:11:46.771365Z","iopub.execute_input":"2022-08-11T20:11:46.771752Z","iopub.status.idle":"2022-08-11T20:11:47.005771Z","shell.execute_reply.started":"2022-08-11T20:11:46.771720Z","shell.execute_reply":"2022-08-11T20:11:47.004524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 6. Boost Classifier:","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import AdaBoostClassifier\n\nadb = AdaBoostClassifier(DecisionTreeClassifier(min_samples_split=80,\n                                                max_depth=10,\n                                                criterion = 'gini'),\n                         n_estimators=30,\n                         learning_rate=0.05,\n                         random_state = 50)\nadb.fit(X_train_exp.iloc[:,1:], y_train)\n\nprint(\"score on dev: \" + str(adb.score(X_dev_exp.iloc[:,1:], y_dev)))\nprint(\"score on train: \"+ str(adb.score(X_train_exp.iloc[:,1:], y_train)))","metadata":{"execution":{"iopub.status.busy":"2022-08-11T20:13:44.325825Z","iopub.execute_input":"2022-08-11T20:13:44.326348Z","iopub.status.idle":"2022-08-11T20:13:46.021750Z","shell.execute_reply.started":"2022-08-11T20:13:44.326300Z","shell.execute_reply":"2022-08-11T20:13:46.020563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_dataset_fe = apply_fe(test_dataset)\nX_test_exp = interaction_exp(test_dataset_fe)\n\nX_test_exp = X_test_exp.drop(['L_Ca_unkn', 'Side_unkn', 'Fami_14', 'Fami_94'], axis = 1)\nX_test_exp['Fami_12'] = 0\nX_test_exp['Fami_15'] = 0\nX_test_exp['Fami_17'] = 0\nX_test_exp = X_test_exp[X_train_exp.columns]\n\ny_pred_test = adb.predict(X_test_exp.iloc[:,1:])\npredictions_test = [round(value) for value in y_pred_test]\n\ntest_dataset['Transported'] = predictions_test\ntest_dataset['Transported'] = test_dataset['Transported'].astype(bool)\n\nsubmission = test_dataset[['PassengerId', 'Transported']]\nsubmission.to_csv('submission.csv', index = False)\n\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T20:14:16.714179Z","iopub.execute_input":"2022-08-11T20:14:16.714610Z","iopub.status.idle":"2022-08-11T20:14:16.895072Z","shell.execute_reply.started":"2022-08-11T20:14:16.714576Z","shell.execute_reply":"2022-08-11T20:14:16.893916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 7. Random Forest Classifier:","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier\n\nrf = RandomForestClassifier(n_estimators=58, max_depth=20, criterion = 'entropy', random_state = 50)\nrf.fit(X_train_exp.iloc[:,1:], y_train)\n\nprint(\"score on dev: \" + str(rf.score(X_dev_exp.iloc[:,1:], y_dev)))\nprint(\"score on train: \"+ str(rf.score(X_train_exp.iloc[:,1:], y_train)))","metadata":{"execution":{"iopub.status.busy":"2022-08-11T20:16:52.547658Z","iopub.execute_input":"2022-08-11T20:16:52.548116Z","iopub.status.idle":"2022-08-11T20:16:53.215291Z","shell.execute_reply.started":"2022-08-11T20:16:52.548073Z","shell.execute_reply":"2022-08-11T20:16:53.214063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_dataset_fe = apply_fe(test_dataset)\nX_test_exp = interaction_exp(test_dataset_fe)\n\nX_test_exp = X_test_exp.drop(['L_Ca_unkn', 'Side_unkn', 'Fami_14', 'Fami_94'], axis = 1)\nX_test_exp['Fami_12'] = 0\nX_test_exp['Fami_15'] = 0\nX_test_exp['Fami_17'] = 0\nX_test_exp = X_test_exp[X_train_exp.columns]\n\ny_pred_test = rf.predict(X_test_exp.iloc[:,1:])\npredictions_test = [round(value) for value in y_pred_test]\n\ntest_dataset['Transported'] = predictions_test\ntest_dataset['Transported'] = test_dataset['Transported'].astype(bool)\n\nsubmission = test_dataset[['PassengerId', 'Transported']]\nsubmission.to_csv('submission.csv', index = False)\n\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T20:17:16.452244Z","iopub.execute_input":"2022-08-11T20:17:16.452731Z","iopub.status.idle":"2022-08-11T20:17:16.655768Z","shell.execute_reply.started":"2022-08-11T20:17:16.452694Z","shell.execute_reply":"2022-08-11T20:17:16.654604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 8. Neural Network:","metadata":{}},{"cell_type":"code","source":"from keras import layers\nfrom keras import models\nfrom keras import optimizers\nfrom keras import losses\nfrom keras import metrics\n\nmodel=models.Sequential()\nmodel.add(layers.Dense(128,activation='relu'))\nmodel.add(layers.Dense(32,activation='selu'))\nmodel.add(layers.Dense(16,activation='elu'))\nmodel.add(layers.Dense(4,activation='tanh'))\nmodel.add(layers.Dense(1,activation='sigmoid'))\nmodel.compile(optimizer='rmsprop',loss='binary_crossentropy',metrics=['accuracy'])\nmodel.fit(X_train_exp.iloc[:,1:],y_train,epochs=15,batch_size=20,validation_data=(X_dev_exp.iloc[:,1:], y_dev))\n\nprint(\"score on test: \" + str(model.evaluate(X_dev_exp.iloc[:,1:], y_dev)[1]))\nprint(\"score on train: \"+ str(model.evaluate(X_train_exp.iloc[:,1:],y_train)[1]))","metadata":{"execution":{"iopub.status.busy":"2022-08-11T20:22:03.917364Z","iopub.execute_input":"2022-08-11T20:22:03.917759Z","iopub.status.idle":"2022-08-11T20:22:16.774549Z","shell.execute_reply.started":"2022-08-11T20:22:03.917726Z","shell.execute_reply":"2022-08-11T20:22:16.773313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_dataset_fe = apply_fe(test_dataset)\nX_test_exp = interaction_exp(test_dataset_fe)\n\nX_test_exp = X_test_exp.drop(['L_Ca_unkn', 'Side_unkn', 'Fami_14', 'Fami_94'], axis = 1)\nX_test_exp['Fami_12'] = 0\nX_test_exp['Fami_15'] = 0\nX_test_exp['Fami_17'] = 0\nX_test_exp = X_test_exp[X_train_exp.columns]\n\ny_pred_test = model.predict(X_test_exp.iloc[:,1:])\npredictions_test = [np.round(x) for x in y_pred_test]\n\ntest_dataset['Transported'] = predictions_test\ntest_dataset['Transported'] = test_dataset['Transported'].astype(bool)\n\nsubmission = test_dataset[['PassengerId', 'Transported']]\nsubmission.to_csv('submission.csv', index = False)\n\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T20:29:56.044512Z","iopub.execute_input":"2022-08-11T20:29:56.044884Z","iopub.status.idle":"2022-08-11T20:29:56.414816Z","shell.execute_reply.started":"2022-08-11T20:29:56.044853Z","shell.execute_reply":"2022-08-11T20:29:56.412906Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 9. XGBoost:","metadata":{}},{"cell_type":"code","source":"import xgboost as xgb\nfrom sklearn.metrics import accuracy_score\n\nmodel = xgb.XGBRegressor(objective ='reg:squarederror', colsample_bytree = 0.8, learning_rate = 0.001,\n                max_depth = 30, alpha = 6, n_estimators = 25)\nmodel.fit(X_train_exp.iloc[:,1:],y_train)\n\ny_pred_train = model.predict(X_train_exp.iloc[:,1:])\npredictions_train = [round(value) for value in y_pred_train]\ny_pred_dev = model.predict(X_dev_exp.iloc[:,1:])\npredictions_dev = [round(value) for value in y_pred_dev]\n\naccuracy_train = accuracy_score(y_train, predictions_train)\naccuracy_dev = accuracy_score(y_dev, predictions_dev)\n\nprint(\"Score on dev: %.2f%%\" % (accuracy_dev * 100.0))\nprint(\"Score on train: %.2f%%\" % (accuracy_train * 100.0))","metadata":{"execution":{"iopub.status.busy":"2022-08-11T20:30:36.291383Z","iopub.execute_input":"2022-08-11T20:30:36.292748Z","iopub.status.idle":"2022-08-11T20:30:37.145996Z","shell.execute_reply.started":"2022-08-11T20:30:36.292704Z","shell.execute_reply":"2022-08-11T20:30:37.144903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So far, the XGBoost is the model with the best performance. To improve results further analysis on omitted variables should be done.","metadata":{}},{"cell_type":"markdown","source":"### Test Dataset Classification from model:","metadata":{}},{"cell_type":"code","source":"test_dataset_fe = apply_fe(test_dataset)\nX_test_exp = interaction_exp(test_dataset_fe)\n\nX_test_exp = X_test_exp.drop(['L_Ca_unkn', 'Side_unkn', 'Fami_14', 'Fami_94'], axis = 1)\nX_test_exp['Fami_12'] = 0\nX_test_exp['Fami_15'] = 0\nX_test_exp['Fami_17'] = 0\nX_test_exp = X_test_exp[X_train_exp.columns]\n\ny_pred_test = model.predict(X_test_exp.iloc[:,1:])\npredictions_test = [np.round(x) for x in y_pred_test]\n\ntest_dataset['Transported'] = predictions_test\ntest_dataset['Transported'] = test_dataset['Transported'].astype(bool)\n\nsubmission = test_dataset[['PassengerId', 'Transported']]\nsubmission.to_csv('submission.csv', index = False)\n\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T20:31:04.739706Z","iopub.execute_input":"2022-08-11T20:31:04.740124Z","iopub.status.idle":"2022-08-11T20:31:04.918436Z","shell.execute_reply.started":"2022-08-11T20:31:04.740090Z","shell.execute_reply":"2022-08-11T20:31:04.917184Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"execution":{"iopub.status.busy":"2022-08-11T18:12:00.434950Z","iopub.status.idle":"2022-08-11T18:12:00.435596Z","shell.execute_reply.started":"2022-08-11T18:12:00.435239Z","shell.execute_reply":"2022-08-11T18:12:00.435266Z"},"trusted":true},"execution_count":null,"outputs":[]}]}