{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# SS Titanic Feature Engineering\n\nby David Zambrano\n\n<img width=\"300\" src=\"https://cdn.drawception.com/drawings/KO3D1hhjOk.png\">\n\ncredits to the author of the image.\n\n## **Competition context:**\n\nWelcome to the year 2912, where your data science skills are needed to solve a cosmic mystery. We've received a transmission from four lightyears away and things aren't looking good.\n\nThe Spaceship Titanic was an interstellar passenger liner launched a month ago. With almost 13,000 passengers on board, the vessel set out on its maiden voyage transporting emigrants from our solar system to three newly habitable exoplanets orbiting nearby stars.\n\nWhile rounding Alpha Centauri en route to its first destination—the torrid 55 Cancri E—the unwary Spaceship Titanic collided with a spacetime anomaly hidden within a dust cloud. Sadly, it met a similar fate as its namesake from 1000 years before. Though the ship stayed intact, almost half of the passengers were transported to an alternate dimension!\n\nTo help rescue crews and retrieve the lost passengers, you are challenged to predict which passengers were transported by the anomaly using records recovered from the spaceship’s damaged computer system.\n\nFor more information about this competition click [here](https://www.kaggle.com/competitions/spaceship-titanic).","metadata":{"execution":{"iopub.status.busy":"2022-08-05T15:08:39.195801Z","iopub.execute_input":"2022-08-05T15:08:39.196706Z","iopub.status.idle":"2022-08-05T15:08:39.239536Z","shell.execute_reply.started":"2022-08-05T15:08:39.196588Z","shell.execute_reply":"2022-08-05T15:08:39.238081Z"}}},{"cell_type":"markdown","source":"## Data Description:\n\n**File name:** train.csv - Contains personal records for about two-thirds (~8700) of the passengers, to be used as training data.\n\n\n> **PassengerId** - A unique Id for each passenger. Each Id takes the form gggg_pp where gggg indicates a group the passenger is travelling with and pp is their number within the group. People in a group are often family members, but not always.\n\n> **HomePlanet** - The planet the passenger departed from, typically their planet of permanent residence.\n\n> **CryoSleep** - Indicates whether the passenger elected to be put into suspended animation for the duration of the voyage. Passengers in cryosleep are confined to their cabins.\n\n> **Cabin** - The cabin number where the passenger is staying. Takes the form deck/num/side, where side can be either P for Port or S for Starboard.\n\n> **Destination** - The planet the passenger will be debarking to.\n\n> **Age** - The age of the passenger.\n\n> **VIP** - Whether the passenger has paid for special VIP service during the voyage.\n\n> **RoomService, FoodCourt, ShoppingMall, Spa, VRDeck** - Amount the passenger has billed at each of the Spaceship Titanic's many luxury amenities.\n\n> **Name** - The first and last names of the passenger.\n\n> **Transported** - Whether the passenger was transported to another dimension. This is the target, the column you are trying to predict.","metadata":{}},{"cell_type":"markdown","source":"To review **Exploratory Data Analysis** click [here](https://www.kaggle.com/code/davidzambrano87/ss-titanic-eda-by-dz).","metadata":{}},{"cell_type":"markdown","source":"## Train Dataset:","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\ntrain_dataset = pd.read_csv('../input/spaceship-titanic/train.csv')\ntrain_dataset.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T19:46:14.592315Z","iopub.execute_input":"2022-08-09T19:46:14.593313Z","iopub.status.idle":"2022-08-09T19:46:14.641863Z","shell.execute_reply.started":"2022-08-09T19:46:14.593271Z","shell.execute_reply":"2022-08-09T19:46:14.640624Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Missing Values:","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport missingno as msnum \n\nmsnum.matrix(train_dataset)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T19:46:14.665275Z","iopub.execute_input":"2022-08-09T19:46:14.665720Z","iopub.status.idle":"2022-08-09T19:46:15.264279Z","shell.execute_reply.started":"2022-08-09T19:46:14.665683Z","shell.execute_reply":"2022-08-09T19:46:15.263406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Filling Missing Values:","metadata":{}},{"cell_type":"code","source":"# Function to fill nan values\ndef fill_nan(dataset):\n    \n    fill_with_mean = ['Age','RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck',]\n    fill_with_mode = ['HomePlanet', 'CryoSleep', 'Cabin', 'Destination', 'VIP']\n        \n    new_dataset = dataset\n    \n    for var in fill_with_mean:\n        new_dataset[var] = new_dataset[var].fillna(new_dataset[var].mean())   \n    \n    for var in fill_with_mode:\n        new_dataset[var] = new_dataset[var].fillna(new_dataset[var].mode().iloc[0])       \n    \n    return new_dataset\n\ntrain_dataset_fe = fill_nan(train_dataset)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T19:46:15.265939Z","iopub.execute_input":"2022-08-09T19:46:15.266509Z","iopub.status.idle":"2022-08-09T19:46:15.291154Z","shell.execute_reply.started":"2022-08-09T19:46:15.266475Z","shell.execute_reply":"2022-08-09T19:46:15.289902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Cuantitative Variables Transformation:","metadata":{}},{"cell_type":"code","source":"import numpy as np\n\n# Function to estimate log values\ndef log_transform(dataset):\n\n    vars_to_transform = ['Age', 'RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']\n\n    new_dataset = dataset.copy()\n    \n    for var in vars_to_transform:\n        new_label = 'log_' + var\n        new_dataset[new_label] = np.log(new_dataset[var] + 1)\n        new_dataset[new_label] = new_dataset[new_label] / new_dataset[new_label].max()\n        \n    return new_dataset\n\ntrain_dataset_fe = log_transform(train_dataset)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T19:46:15.292742Z","iopub.execute_input":"2022-08-09T19:46:15.293383Z","iopub.status.idle":"2022-08-09T19:46:15.310671Z","shell.execute_reply.started":"2022-08-09T19:46:15.293332Z","shell.execute_reply":"2022-08-09T19:46:15.309711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns\n\ndef density_comparison_fe(var1):\n    fig, axes = plt.subplots(1, 2, sharex=False, sharey=False, figsize=(15,4))\n    title = var1 + ' original vs log_transformed distribution'\n    fig.suptitle(title)\n    sns.kdeplot(train_dataset_fe[var1], shade=True, color=\"r\", ax = axes[0])\n    sns.kdeplot(train_dataset_fe['log_' + var1], shade=True, color=\"g\", ax = axes[1])\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T19:46:15.314219Z","iopub.execute_input":"2022-08-09T19:46:15.315097Z","iopub.status.idle":"2022-08-09T19:46:15.323331Z","shell.execute_reply.started":"2022-08-09T19:46:15.315016Z","shell.execute_reply":"2022-08-09T19:46:15.322353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"vars_to_transform = ['Age', 'RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']\n\nfor var in vars_to_transform:\n    density_comparison_fe(var)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T19:46:15.324962Z","iopub.execute_input":"2022-08-09T19:46:15.325514Z","iopub.status.idle":"2022-08-09T19:46:17.797671Z","shell.execute_reply.started":"2022-08-09T19:46:15.325481Z","shell.execute_reply":"2022-08-09T19:46:17.796306Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"log_columns = ['log_Age', 'log_RoomService', 'log_FoodCourt', 'log_ShoppingMall', \n                'log_Spa', 'log_VRDeck', 'Transported']\n\nfig, axes = plt.subplots(1, 2, sharex=False, sharey=False, figsize=(15,6))\nfig.suptitle('original vs log_transformed correlations')\nsns.heatmap(train_dataset.corr(), annot = True, cmap = 'coolwarm',  ax = axes[0])\nsns.heatmap(train_dataset_fe[log_columns].corr(), annot = True,cmap = 'coolwarm',ax = axes[1])\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T19:46:17.798989Z","iopub.execute_input":"2022-08-09T19:46:17.799422Z","iopub.status.idle":"2022-08-09T19:46:19.224886Z","shell.execute_reply.started":"2022-08-09T19:46:17.799391Z","shell.execute_reply":"2022-08-09T19:46:19.223655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Log transformation increased the absolute values of the correlations of every feature with the target, indicating that this should be useful to improve model's performance.","metadata":{}},{"cell_type":"code","source":"sns.pairplot(train_dataset_fe[log_columns], hue = \"Transported\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T19:46:19.226325Z","iopub.execute_input":"2022-08-09T19:46:19.226797Z","iopub.status.idle":"2022-08-09T19:46:41.559802Z","shell.execute_reply.started":"2022-08-09T19:46:19.226750Z","shell.execute_reply":"2022-08-09T19:46:41.558725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"New hints about being transported from the expenses in different amenities.","metadata":{}},{"cell_type":"markdown","source":"### Amenities Expenditure:","metadata":{}},{"cell_type":"markdown","source":"From the EDA it was pointed out that in some way, the distribution of these varaibles were related to the probability of being transported. Here it is explored in depth to identify a new variable that helps the model figure it out.","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 4, sharex=True, sharey=True, figsize=(18,3))\n\ntemp_df = train_dataset[[\"Transported\", \"RoomService\", \"VRDeck\", \"Spa\"]].copy()\n\ntemp_df[\"dummy_expenses_TT\"] = (temp_df[\"RoomService\"] > temp_df[\"VRDeck\"]) | (temp_df[\"RoomService\"] > temp_df[\"Spa\"])\nsns.heatmap(pd.crosstab(temp_df[\"dummy_expenses_TT\"], temp_df[\"Transported\"]), \n            annot = True, fmt = \".0f\", cmap = \"coolwarm\",ax=axes[0])\n\ntemp_df[\"dummy_expenses_TF\"] = (temp_df[\"RoomService\"] > temp_df[\"VRDeck\"]) | ~(temp_df[\"RoomService\"] > temp_df[\"Spa\"])\nsns.heatmap(pd.crosstab(temp_df[\"dummy_expenses_TF\"], temp_df[\"Transported\"]), \n            annot = True, fmt = \".0f\", cmap = \"coolwarm\",ax=axes[1])\n\ntemp_df[\"dummy_expenses_FT\"] = ~(temp_df[\"RoomService\"] > temp_df[\"VRDeck\"]) | (temp_df[\"RoomService\"] > temp_df[\"Spa\"])\nsns.heatmap(pd.crosstab(temp_df[\"dummy_expenses_FT\"], temp_df[\"Transported\"]), \n            annot = True, fmt = \".0f\", cmap = \"coolwarm\", ax=axes[2])\n\ntemp_df[\"dummy_expenses_FF\"] = ~(temp_df[\"RoomService\"] > temp_df[\"VRDeck\"]) | ~(temp_df[\"RoomService\"] > temp_df[\"Spa\"])\nsns.heatmap(pd.crosstab(temp_df[\"dummy_expenses_FF\"], temp_df[\"Transported\"]), \n            annot = True, fmt = \".0f\", cmap = \"coolwarm\",ax=axes[3])\nplt.show()\n\nsns.heatmap(temp_df.corr(), annot = True,cmap = 'coolwarm')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T19:46:41.561197Z","iopub.execute_input":"2022-08-09T19:46:41.561542Z","iopub.status.idle":"2022-08-09T19:46:43.165760Z","shell.execute_reply.started":"2022-08-09T19:46:41.561511Z","shell.execute_reply":"2022-08-09T19:46:43.164664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The findings indicate that dummy variables dummy_expenses_TT reached the highest correlation with the target variable, and in second place it is the dummy_expenses_FF but this one is also related with RoomService which is not desirable at all.","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 4, sharex=True, sharey=True, figsize=(18,3))\n\ntemp_df = train_dataset[[\"Transported\", \"FoodCourt\", \"VRDeck\", \"Spa\"]].copy()\n\ntemp_df[\"dummy_expenses_TT\"] = (temp_df[\"FoodCourt\"] > temp_df[\"VRDeck\"]) | (temp_df[\"FoodCourt\"] > temp_df[\"Spa\"])\nsns.heatmap(pd.crosstab(temp_df[\"dummy_expenses_TT\"], temp_df[\"Transported\"]), \n            annot = True, fmt = \".0f\", cmap = \"coolwarm\", ax=axes[0])\n\ntemp_df[\"dummy_expenses_TF\"] = (temp_df[\"FoodCourt\"] > temp_df[\"VRDeck\"]) | ~(temp_df[\"FoodCourt\"] > temp_df[\"Spa\"])\nsns.heatmap(pd.crosstab(temp_df[\"dummy_expenses_TF\"], temp_df[\"Transported\"]), \n            annot = True, fmt = \".0f\", cmap = \"coolwarm\", ax=axes[1])\n\ntemp_df[\"dummy_expenses_FT\"] = ~(temp_df[\"FoodCourt\"] > temp_df[\"VRDeck\"]) | (temp_df[\"FoodCourt\"] > temp_df[\"Spa\"])\nsns.heatmap(pd.crosstab(temp_df[\"dummy_expenses_FT\"], temp_df[\"Transported\"]), \n            annot = True, fmt = \".0f\", cmap = \"coolwarm\", ax=axes[2])\n\ntemp_df[\"dummy_expenses_FF\"] = ~(temp_df[\"FoodCourt\"] > temp_df[\"VRDeck\"]) | ~(temp_df[\"FoodCourt\"] > temp_df[\"Spa\"])\nsns.heatmap(pd.crosstab(temp_df[\"dummy_expenses_FF\"], temp_df[\"Transported\"]), \n            annot = True, fmt = \".0f\", cmap = \"coolwarm\", ax=axes[3])\nplt.show()\n\nsns.heatmap(temp_df.corr(), annot = True,cmap = 'coolwarm')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T20:05:13.788656Z","iopub.execute_input":"2022-08-09T20:05:13.789134Z","iopub.status.idle":"2022-08-09T20:05:15.065291Z","shell.execute_reply.started":"2022-08-09T20:05:13.789099Z","shell.execute_reply":"2022-08-09T20:05:15.064439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this second excercise is interesting to notice that the variable dummy_expenses_FT reaches the highes correlation with the target.","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 4, sharex=True, sharey=True, figsize=(18,3))\n\ntemp_df = train_dataset[[\"Transported\", \"FoodCourt\", \"RoomService\", \"VRDeck\", \"Spa\"]].copy()\n\ntemp_df[\"dummy_expenses_TT\"] = (temp_df[\"FoodCourt\"] > temp_df[\"VRDeck\"]) | ~(temp_df[\"RoomService\"] > temp_df[\"VRDeck\"])\nsns.heatmap(pd.crosstab(temp_df[\"dummy_expenses_TT\"], temp_df[\"Transported\"]), \n            annot = True, fmt = \".0f\", cmap = \"coolwarm\",ax=axes[0])\n\ntemp_df[\"dummy_expenses_TF\"] = (temp_df[\"FoodCourt\"] > temp_df[\"Spa\"]) | ~(temp_df[\"RoomService\"] > temp_df[\"Spa\"])\nsns.heatmap(pd.crosstab(temp_df[\"dummy_expenses_TF\"], temp_df[\"Transported\"]), \n            annot = True, fmt = \".0f\", cmap = \"coolwarm\",ax=axes[1])\n\ntemp_df[\"dummy_expenses_FT\"] = ~(temp_df[\"FoodCourt\"] > temp_df[\"VRDeck\"]) & ~(temp_df[\"RoomService\"] < temp_df[\"VRDeck\"])\nsns.heatmap(pd.crosstab(temp_df[\"dummy_expenses_FT\"], temp_df[\"Transported\"]), \n            annot = True, fmt = \".0f\", cmap = \"coolwarm\", ax=axes[2])\n\ntemp_df[\"dummy_expenses_FF\"] = ~(temp_df[\"FoodCourt\"] > temp_df[\"Spa\"]) & ~(temp_df[\"RoomService\"] < temp_df[\"Spa\"])\nsns.heatmap(pd.crosstab(temp_df[\"dummy_expenses_FF\"], temp_df[\"Transported\"]), \n            annot = True, fmt = \".0f\", cmap = \"coolwarm\",ax=axes[3])\nplt.show()\n\nsns.heatmap(temp_df.corr(), annot = True,cmap = 'coolwarm')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T20:25:34.980873Z","iopub.execute_input":"2022-08-09T20:25:34.981288Z","iopub.status.idle":"2022-08-09T20:25:36.450175Z","shell.execute_reply.started":"2022-08-09T20:25:34.981255Z","shell.execute_reply":"2022-08-09T20:25:36.449169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"ok, to sum up:","metadata":{}},{"cell_type":"code","source":"train_dataset_fe[\"exp_1\"] = (train_dataset_fe[\"RoomService\"] > train_dataset_fe[\"VRDeck\"]) | (train_dataset_fe[\"RoomService\"] > train_dataset_fe[\"Spa\"])\ntrain_dataset_fe[\"exp_2\"] = (train_dataset_fe[\"RoomService\"] > train_dataset_fe[\"VRDeck\"]) | ~(train_dataset_fe[\"RoomService\"] > train_dataset_fe[\"Spa\"])\ntrain_dataset_fe[\"exp_3\"] = ~(train_dataset_fe[\"RoomService\"] > train_dataset_fe[\"VRDeck\"]) | (train_dataset_fe[\"RoomService\"] > train_dataset_fe[\"Spa\"])\ntrain_dataset_fe[\"exp_4\"] = ~(train_dataset_fe[\"RoomService\"] > train_dataset_fe[\"VRDeck\"]) | ~(train_dataset_fe[\"RoomService\"] > train_dataset_fe[\"Spa\"])\ntrain_dataset_fe[\"exp_5\"] = (train_dataset_fe[\"FoodCourt\"] > train_dataset_fe[\"VRDeck\"]) | (train_dataset_fe[\"FoodCourt\"] > train_dataset_fe[\"Spa\"])\ntrain_dataset_fe[\"exp_6\"] = (train_dataset_fe[\"FoodCourt\"] > train_dataset_fe[\"VRDeck\"]) | ~(train_dataset_fe[\"FoodCourt\"] > train_dataset_fe[\"Spa\"])\ntrain_dataset_fe[\"exp_7\"] = ~(train_dataset_fe[\"FoodCourt\"] > train_dataset_fe[\"VRDeck\"]) | (train_dataset_fe[\"FoodCourt\"] > train_dataset_fe[\"Spa\"])\ntrain_dataset_fe[\"exp_8\"] = ~(train_dataset_fe[\"FoodCourt\"] > train_dataset_fe[\"VRDeck\"]) | ~(train_dataset_fe[\"FoodCourt\"] > train_dataset_fe[\"Spa\"])\ntrain_dataset_fe[\"exp_9\"] = (train_dataset_fe[\"FoodCourt\"] > train_dataset_fe[\"VRDeck\"]) | ~(train_dataset_fe[\"RoomService\"] > train_dataset_fe[\"VRDeck\"])\ntrain_dataset_fe[\"exp_10\"] = (train_dataset_fe[\"FoodCourt\"] > train_dataset_fe[\"Spa\"]) | ~(train_dataset_fe[\"RoomService\"] > train_dataset_fe[\"Spa\"])\ntrain_dataset_fe[\"exp_11\"] = ~(train_dataset_fe[\"FoodCourt\"] > train_dataset_fe[\"VRDeck\"]) & ~(train_dataset_fe[\"RoomService\"] < train_dataset_fe[\"VRDeck\"])\ntrain_dataset_fe[\"exp_12\"] = ~(train_dataset_fe[\"FoodCourt\"] > train_dataset_fe[\"Spa\"]) & ~(train_dataset_fe[\"RoomService\"] > train_dataset_fe[\"Spa\"])\n\nfe_vars = ['Transported', 'CryoSleep', 'VIP', 'log_Age', 'log_RoomService', 'log_FoodCourt', 'log_ShoppingMall', \n           'log_Spa', 'log_VRDeck', 'exp_1', 'exp_2', 'exp_3', 'exp_4', 'exp_5', 'exp_6', 'exp_7', 'exp_8',\n           'exp_9', 'exp_10', 'exp_11', 'exp_12']\n\nplt.subplots(figsize = (15,9))\nsns.heatmap(train_dataset_fe[fe_vars].corr(), annot = True,cmap = 'coolwarm', fmt = '.2f')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T20:26:36.778337Z","iopub.execute_input":"2022-08-09T20:26:36.779246Z","iopub.status.idle":"2022-08-09T20:26:38.824585Z","shell.execute_reply.started":"2022-08-09T20:26:36.779208Z","shell.execute_reply":"2022-08-09T20:26:38.823426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### One Hot Encoding:","metadata":{}},{"cell_type":"code","source":"home_dummy = pd.get_dummies(train_dataset_fe['HomePlanet'], prefix='Home')\ntrain_dataset_fe = pd.merge(left = train_dataset_fe, right = home_dummy, left_index = True, right_index = True)\n\ndestination_dummy = pd.get_dummies(train_dataset_fe['Destination'], prefix='Dest')\ntrain_dataset_fe = pd.merge(left = train_dataset_fe, right = destination_dummy, left_index = True,right_index = True)\n\ntemp_df = train_dataset_fe[['Transported', 'Home_Earth', 'Home_Europa', 'Home_Mars', 'Dest_55 Cancri e', \n                            'Dest_PSO J318.5-22', 'Dest_TRAPPIST-1e',]]\n\nplt.subplots(figsize = (12,9))\nsns.heatmap(temp_df.corr(), annot = True, cmap = 'coolwarm', fmt = \".2f\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T19:46:45.262103Z","iopub.execute_input":"2022-08-09T19:46:45.262666Z","iopub.status.idle":"2022-08-09T19:46:45.746448Z","shell.execute_reply.started":"2022-08-09T19:46:45.262629Z","shell.execute_reply":"2022-08-09T19:46:45.745226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To review **Exploratory Data Analysis** click [here](https://www.kaggle.com/code/davidzambrano87/ss-titanic-eda-by-dz).\n\n## Next step:\n\n> [Models and Error Analysis.](https://www.kaggle.com/code/davidzambrano87/ss-titanic-modeling-by-dz)","metadata":{}}]}