{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# SS Titanic EDA\n\nby David Zambrano\n\n<img width=\"300\" src=\"https://cdn.drawception.com/drawings/KO3D1hhjOk.png\">\n\ncredits to the author of the image.\n\n## **Competition context:**\n\nWelcome to the year 2912, where your data science skills are needed to solve a cosmic mystery. We've received a transmission from four lightyears away and things aren't looking good.\n\nThe Spaceship Titanic was an interstellar passenger liner launched a month ago. With almost 13,000 passengers on board, the vessel set out on its maiden voyage transporting emigrants from our solar system to three newly habitable exoplanets orbiting nearby stars.\n\nWhile rounding Alpha Centauri en route to its first destination—the torrid 55 Cancri E—the unwary Spaceship Titanic collided with a spacetime anomaly hidden within a dust cloud. Sadly, it met a similar fate as its namesake from 1000 years before. Though the ship stayed intact, almost half of the passengers were transported to an alternate dimension!\n\nTo help rescue crews and retrieve the lost passengers, you are challenged to predict which passengers were transported by the anomaly using records recovered from the spaceship’s damaged computer system.\n\nFor more information about this competition click [here](https://www.kaggle.com/competitions/spaceship-titanic).","metadata":{}},{"cell_type":"markdown","source":"## Data Description:\n\n**File name:** train.csv - Contains personal records for about two-thirds (~8700) of the passengers, to be used as training data.\n\n\n> **PassengerId** - A unique Id for each passenger. Each Id takes the form gggg_pp where gggg indicates a group the passenger is travelling with and pp is their number within the group. People in a group are often family members, but not always.\n\n> **HomePlanet** - The planet the passenger departed from, typically their planet of permanent residence.\n\n> **CryoSleep** - Indicates whether the passenger elected to be put into suspended animation for the duration of the voyage. Passengers in cryosleep are confined to their cabins.\n\n> **Cabin** - The cabin number where the passenger is staying. Takes the form deck/num/side, where side can be either P for Port or S for Starboard.\n\n> **Destination** - The planet the passenger will be debarking to.\n\n> **Age** - The age of the passenger.\n\n> **VIP** - Whether the passenger has paid for special VIP service during the voyage.\n\n> **RoomService, FoodCourt, ShoppingMall, Spa, VRDeck** - Amount the passenger has billed at each of the Spaceship Titanic's many luxury amenities.\n\n> **Name** - The first and last names of the passenger.\n\n> **Transported** - Whether the passenger was transported to another dimension. This is the target, the column you are trying to predict.","metadata":{}},{"cell_type":"markdown","source":"This part of the project is to figure out what the frog happened with the bias of the models. Fortunately the variance is not a problem... yet. As a result, it is expected to bring feedback to improve EDA, Feature Engineering and Models, and off course improve the position in the leaderboard.\n\nHere I go again!\n\nTo know more about the project you can review:\n\n> [EDA](https://www.kaggle.com/code/davidzambrano87/ss-titanic-eda-by-dz)\n\n> [Feature Engineering](https://www.kaggle.com/code/davidzambrano87/ss-titanic-fe-by-dz) \n\n> [Models](https://www.kaggle.com/code/davidzambrano87/ss-titanic-modeling-by-dz)","metadata":{}},{"cell_type":"markdown","source":"## Data, FE, models, predictions and everything required to start Error Analysis:","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\ntrain_dataset = pd.read_csv('../input/spaceship-titanic/train.csv')\ntrain_dataset.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T17:17:54.162664Z","iopub.execute_input":"2022-08-09T17:17:54.163103Z","iopub.status.idle":"2022-08-09T17:17:54.216177Z","shell.execute_reply.started":"2022-08-09T17:17:54.163066Z","shell.execute_reply":"2022-08-09T17:17:54.215005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\n\n# Function to fill nan values\ndef fill_nan(dataset):\n    \n    fill_with_mean = ['Age','RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck',]\n    fill_with_mode = ['HomePlanet', 'CryoSleep', 'Cabin', 'Destination', 'VIP']\n        \n    new_dataset = dataset\n    \n    for var in fill_with_mean:\n        new_dataset[var] = new_dataset[var].fillna(0)   \n    \n    for var in fill_with_mode:\n        new_dataset[var] = new_dataset[var].fillna(new_dataset[var].mode().iloc[0])       \n    \n    return new_dataset\n\n\n# Function to estimate log values\ndef log_transform(dataset):\n\n    vars_to_transform = ['Age', 'RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']\n\n    new_dataset = dataset.copy()\n    \n    for var in vars_to_transform:\n        new_label = 'log_' + var\n        new_dataset[new_label] = np.log(new_dataset[var] + 1)\n        new_dataset[new_label] = new_dataset[new_label] / new_dataset[new_label].max()\n        \n    return new_dataset\n\n\n# Function to create new features:\ndef new_features(dataset):\n    \n    new_dataset = dataset.copy()\n    \n    new_dataset[\"exp_1\"] = (new_dataset[\"log_RoomService\"] > new_dataset[\"log_VRDeck\"]) | (new_dataset[\"log_RoomService\"] > new_dataset[\"log_Spa\"])\n    new_dataset[\"exp_2\"] = ~(new_dataset[\"log_RoomService\"] > new_dataset[\"log_VRDeck\"]) | ~(new_dataset[\"log_RoomService\"] > new_dataset[\"log_Spa\"])\n    \n    return new_dataset\n\n\n# Function for one-hot-encoding:\ndef one_hot_encode(dataset):\n    \n    new_dataset = dataset.copy()\n    \n    variables_to_encode = ['HomePlanet', 'Destination']\n    \n    for var in variables_to_encode:\n        dummy_label = var[:4]\n        dummy = pd.get_dummies(new_dataset[var], prefix = dummy_label)\n        new_dataset = pd.merge(left = new_dataset, right = dummy, left_index = True, right_index = True)\n        \n    return new_dataset\n\n\n# Function to fix data types:\ndef fix_data_types(dataset):\n    \n    new_dataset = dataset.copy()\n    \n    vars_to_fix_data_type = ['CryoSleep', 'VIP', 'exp_1', 'exp_2']\n    \n    for var in vars_to_fix_data_type:\n        new_dataset[var] = new_dataset[var].astype(int)\n    \n    return new_dataset\n\n\n# Function to remove unused variables:\ndef remove_variables(dataset):\n    \n    new_dataset = dataset.copy()\n    \n    variables_to_keep = ['PassengerId', 'CryoSleep', 'VIP', \n                         'log_Age', 'log_RoomService', 'log_FoodCourt', 'log_ShoppingMall', 'log_Spa', 'log_VRDeck',\n                         'exp_1', 'exp_2', \n                         'Home_Earth', 'Home_Europa', 'Home_Mars', \n                         'Dest_55 Cancri e', 'Dest_PSO J318.5-22', 'Dest_TRAPPIST-1e',\n                         'Transported']\n    \n    try:\n        new_dataset = new_dataset[variables_to_keep]\n    \n    except:\n        new_dataset = new_dataset[variables_to_keep[:-1]]\n    \n    return new_dataset\n\n\n# Function to apply all the feature engineering to a dataset:\ndef apply_fe(dataset):\n    \n    new_dataset = dataset.copy()\n    \n    new_dataset = fill_nan(dataset)\n    new_dataset = log_transform(new_dataset)\n    new_dataset = new_features(new_dataset)\n    new_dataset = one_hot_encode(new_dataset)\n    new_dataset = fix_data_types(new_dataset)\n    new_dataset = remove_variables(new_dataset)\n    \n    return new_dataset","metadata":{"execution":{"iopub.status.busy":"2022-08-09T17:17:54.218681Z","iopub.execute_input":"2022-08-09T17:17:54.219310Z","iopub.status.idle":"2022-08-09T17:17:54.239198Z","shell.execute_reply.started":"2022-08-09T17:17:54.219265Z","shell.execute_reply":"2022-08-09T17:17:54.237997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Things to improve:\n\n1. Missing values here are a pain in the smallest toe to say something nice about them. Filling them with zeros, means, medians, modes and everything might be a source of bias.\n\n<img width=\"300\" src=\"https://dss.fosterwebmarketing.com/upload/1094/pinky%20toe.png\">\n\nThat will be reviewed later, for now let's think about what is generating the bias:","metadata":{}},{"cell_type":"code","source":"train_dataset_fe = apply_fe(train_dataset)\ntrain_dataset_fe.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T17:17:54.241473Z","iopub.execute_input":"2022-08-09T17:17:54.241818Z","iopub.status.idle":"2022-08-09T17:17:54.312704Z","shell.execute_reply.started":"2022-08-09T17:17:54.241786Z","shell.execute_reply":"2022-08-09T17:17:54.311581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\nX = train_dataset_fe[train_dataset_fe.columns[:-1]]\ny = train_dataset_fe['Transported']\n\nX_train, X_dev, y_train, y_dev = train_test_split(X, y, test_size=0.20, random_state=1)\nlen(X_dev), len(X_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T17:17:54.314282Z","iopub.execute_input":"2022-08-09T17:17:54.314613Z","iopub.status.idle":"2022-08-09T17:17:54.897365Z","shell.execute_reply.started":"2022-08-09T17:17:54.314584Z","shell.execute_reply":"2022-08-09T17:17:54.896157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_dataset = pd.read_csv('../input/spaceship-titanic/test.csv')\ntest_dataset.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T17:17:54.899815Z","iopub.execute_input":"2022-08-09T17:17:54.900197Z","iopub.status.idle":"2022-08-09T17:17:54.938965Z","shell.execute_reply.started":"2022-08-09T17:17:54.900163Z","shell.execute_reply":"2022-08-09T17:17:54.937772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test = apply_fe(test_dataset)\nX_test.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T17:17:54.940626Z","iopub.execute_input":"2022-08-09T17:17:54.941113Z","iopub.status.idle":"2022-08-09T17:17:54.998575Z","shell.execute_reply.started":"2022-08-09T17:17:54.941066Z","shell.execute_reply":"2022-08-09T17:17:54.997401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Function to add interactions to dataset:\ndef interaction_exp(dataset):\n    \n    var_to_interact = ['CryoSleep', 'VIP', 'exp_1', 'exp_2']\n    num_vars = ['log_RoomService', 'log_FoodCourt','log_ShoppingMall', 'log_Spa', 'log_VRDeck']\n\n    dataset_exp = dataset.copy()\n\n    for i in var_to_interact:\n        for j in num_vars:\n            new_label = 'int_' + i + '_' + j[4:7]\n            dataset_exp[new_label] = dataset_exp[i] * dataset_exp[j]\n            \n    return dataset_exp","metadata":{"execution":{"iopub.status.busy":"2022-08-09T17:17:54.999928Z","iopub.execute_input":"2022-08-09T17:17:55.000292Z","iopub.status.idle":"2022-08-09T17:17:55.007382Z","shell.execute_reply.started":"2022-08-09T17:17:55.000258Z","shell.execute_reply":"2022-08-09T17:17:55.006284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_exp = interaction_exp(X_train)\nX_dev_exp = interaction_exp(X_dev)\nX_test_exp = interaction_exp(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T17:17:55.008907Z","iopub.execute_input":"2022-08-09T17:17:55.009500Z","iopub.status.idle":"2022-08-09T17:17:55.057445Z","shell.execute_reply.started":"2022-08-09T17:17:55.009466Z","shell.execute_reply":"2022-08-09T17:17:55.056297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import xgboost as xgb\nfrom sklearn.metrics import accuracy_score\n\nmodel = xgb.XGBRegressor(objective ='reg:squarederror', colsample_bytree = 0.77, learning_rate = 0.0015,\n                max_depth = 30, alpha = 4, n_estimators = 300)\nmodel.fit(X_train_exp.iloc[:,1:],y_train)\n\ny_pred_train = model.predict(X_train_exp.iloc[:,1:])\npredictions_train = [round(value) for value in y_pred_train]\ny_pred_dev = model.predict(X_dev_exp.iloc[:,1:])\npredictions_dev = [round(value) for value in y_pred_dev]\n\naccuracy_train = accuracy_score(y_train, predictions_train)\naccuracy_dev = accuracy_score(y_dev, predictions_dev)\n\nprint(\"Score on dev: %.2f%%\" % (accuracy_dev * 100.0))\nprint(\"Score on train: %.2f%%\" % (accuracy_train * 100.0))","metadata":{"execution":{"iopub.status.busy":"2022-08-09T17:18:47.529445Z","iopub.execute_input":"2022-08-09T17:18:47.529867Z","iopub.status.idle":"2022-08-09T17:18:54.210689Z","shell.execute_reply.started":"2022-08-09T17:18:47.529832Z","shell.execute_reply":"2022-08-09T17:18:54.209789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"bias of 17.3% and variance of 2.6%... Ooops...\n\n<img width=\"300\" src=\"https://memegenerator.net/img/instances/15072019.jpg\">","metadata":{}},{"cell_type":"code","source":"y_pred_test = model.predict(X_test_exp.iloc[:,1:])\npredictions_test = [round(value) for value in y_pred_test]\n\ntest_dataset['Transported'] = predictions_test\ntest_dataset['Transported'] = test_dataset['Transported'].astype(bool)\ntest_dataset.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T17:19:31.917899Z","iopub.execute_input":"2022-08-09T17:19:31.919190Z","iopub.status.idle":"2022-08-09T17:19:31.994207Z","shell.execute_reply.started":"2022-08-09T17:19:31.919132Z","shell.execute_reply":"2022-08-09T17:19:31.993054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = test_dataset[['PassengerId', 'Transported']]\nsubmission.to_csv('submission.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T17:19:48.296751Z","iopub.execute_input":"2022-08-09T17:19:48.297171Z","iopub.status.idle":"2022-08-09T17:19:48.312097Z","shell.execute_reply.started":"2022-08-09T17:19:48.297137Z","shell.execute_reply":"2022-08-09T17:19:48.311007Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Previous submission scored around 79.21% on the test dataset. Very close to the dev_dataset performance which is something desirable. \n\n<img width=\"300\" src=\"https://ep01.epimg.net/verne/imagenes/2017/01/10/articulo/1484060797_716917_1484061655_sumario_normal.jpg\">\n\nNow let's figure out what the frog happened.","metadata":{}},{"cell_type":"markdown","source":"### X_train_exp and predictions:","metadata":{}},{"cell_type":"code","source":"# Count of cases right and wrong classified:\nX_train_exp['Transported'] = y_train\nX_train_exp['Prediction'] = predictions_train\nX_train_exp['Prediction'] = X_train_exp['Prediction'].astype(bool)\npd.DataFrame(X_train_exp[['Transported', 'Prediction']].value_counts())","metadata":{"execution":{"iopub.status.busy":"2022-08-09T17:27:44.501157Z","iopub.execute_input":"2022-08-09T17:27:44.501530Z","iopub.status.idle":"2022-08-09T17:27:44.520373Z","shell.execute_reply.started":"2022-08-09T17:27:44.501497Z","shell.execute_reply":"2022-08-09T17:27:44.519357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"5748 right cases vs 1206 wrong classifications. \n\nLet's identify them to find the pattern biasing the model. yeah!!!!!!\n\n\n<img width=\"300\" src=\"https://scontent.fbog3-2.fna.fbcdn.net/v/t1.18169-9/23621715_139198490175223_5550282881484511641_n.jpg?_nc_cat=103&ccb=1-7&_nc_sid=09cbfe&_nc_eui2=AeEATYbugj6ZyBjeTwX7OkK9ybDPvTvP6FrJsM-9O8_oWr7lHtKRuGmUyEXdhYoprRw&_nc_ohc=lWnt-vJONs4AX9eIb99&_nc_ht=scontent.fbog3-2.fna&oh=00_AT9Ni-2bBhQ6NbEm7SDXuaUJPxTX2ZCU9iFjIlEl89nAGA&oe=6316F290\">","metadata":{}},{"cell_type":"code","source":"X_train_exp['predict_ok'] = np.where(X_train_exp['Transported'] == X_train_exp['Prediction'], 'ok', 'not_ok')\nX_train_exp.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:15:18.887159Z","iopub.execute_input":"2022-08-09T18:15:18.887688Z","iopub.status.idle":"2022-08-09T18:15:18.925538Z","shell.execute_reply.started":"2022-08-09T18:15:18.887642Z","shell.execute_reply":"2022-08-09T18:15:18.924590Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Calling old EDA again, but much cooler this time:\n\n<img width=\"300\" src=\"https://st1.uvnimg.com/dims4/default/6b1e00e/2147483647/thumbnail/1024x576%3E/quality/75/?url=https%3A%2F%2Fuvn-brightspot.s3.amazonaws.com%2Fassets%2Fvixes%2Fa%2Fabuelita-viendo-su-computadora-meme.jpg\">","metadata":{}}]}