{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Spaceship Titanic with Gradient Boosting Classifier","metadata":{}},{"cell_type":"markdown","source":"**The purpose of this notebook is to show that it is possible to get a decent score with minimal steps.** This allows to have a good start in order to go further.\n\nI won't go deeper into the dataset (EDA, feature engineering, cleaning, preprocessing etc.) nor the classifier used here (Gradient Boosting). More information can be found in the links given in the second part of the notebook. Many high-level notebooks and discussion posts already have comprehensive guides to all models and tools.\n\nThis notebook is divided into two parts. The second part can be skipped.\n1. [Short way to get >0.80](#Minimal)\n2. [Further reading](#Further)\n\n","metadata":{}},{"cell_type":"markdown","source":"<a id='Minimal'></a>\n# Short way to get >0.80","metadata":{}},{"cell_type":"markdown","source":"Let's import only what is necessary for the first part.","metadata":{}},{"cell_type":"code","source":"import os\nimport pandas as pd\nfrom sklearn import preprocessing\nfrom sklearn.ensemble import GradientBoostingClassifier","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:31:59.151887Z","iopub.execute_input":"2022-08-07T15:31:59.152262Z","iopub.status.idle":"2022-08-07T15:31:59.613683Z","shell.execute_reply.started":"2022-08-07T15:31:59.152233Z","shell.execute_reply":"2022-08-07T15:31:59.612861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Some data modifications","metadata":{}},{"cell_type":"markdown","source":"In this section, few modifications to the datasets will be made to get slightly better datasets for the classifier.\n\nThe modifications are the following:\n- Step 1: From the `Cabin` feature in the form of `deck/num/side`, get `desk` and `side`.\n- Step 2: Drop the colums `PassengerId`, `Cabin` and `Name`.\n- Step 3: Convert categorical variables into indicator variables with `pandas.get_dummies` function.\n- Step 4: Fill NaN values (for train and test sets) with most frequent values **from the train set**.\n- Step 5: Standardize float values of the train and test sets with a scaler fitted to the train set.","metadata":{}},{"cell_type":"code","source":"df_train = pd.read_csv(\"../input/spaceship-titanic/train.csv\")\ndf_test = pd.read_csv(\"../input/spaceship-titanic/test.csv\")\ndf_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T11:28:38.326788Z","iopub.execute_input":"2022-08-06T11:28:38.328004Z","iopub.status.idle":"2022-08-06T11:28:38.441673Z","shell.execute_reply.started":"2022-08-06T11:28:38.327961Z","shell.execute_reply":"2022-08-06T11:28:38.440476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def data_edit(df):\n    # Step 1: Get Deck and Side from Cabin column\n    df[\"Desk\"] = [cabin[0] if (type(cabin)==str) else cabin for cabin in df['Cabin']]\n    df[\"Side\"] = [cabin[-1] if (type(cabin)==str) else cabin for cabin in df['Cabin']]\n    # Step 2: Drop columns\n    df = df.drop(columns=[\"PassengerId\",\"Cabin\",\"Name\"])\n    # Step 3: Get indicator variables\n    df = pd.get_dummies(df, drop_first=True)\n    return df","metadata":{"execution":{"iopub.status.busy":"2022-08-06T11:28:38.968840Z","iopub.execute_input":"2022-08-06T11:28:38.969744Z","iopub.status.idle":"2022-08-06T11:28:38.977755Z","shell.execute_reply.started":"2022-08-06T11:28:38.969700Z","shell.execute_reply":"2022-08-06T11:28:38.976365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def data_scaler(df_train,df_test) :\n    # Step 4: Fill NaN values with most frequent values from the train set\n    most_frequent = df_train.mode().loc[0,:]\n    df_train = df_train.fillna(value=most_frequent)\n    df_test = df_test.fillna(value=most_frequent)\n    # Step 5: Standardize float columns\n    num_cols = [\"Age\",\"RoomService\",\"FoodCourt\",\"ShoppingMall\",\"Spa\",\"VRDeck\"]\n    std = preprocessing.StandardScaler()\n    df_train[num_cols] = std.fit_transform(df_train[num_cols])\n    df_test[num_cols] = std.transform(df_test[num_cols])\n    return df_train, df_test","metadata":{"execution":{"iopub.status.busy":"2022-08-06T11:28:39.456552Z","iopub.execute_input":"2022-08-06T11:28:39.456984Z","iopub.status.idle":"2022-08-06T11:28:39.464479Z","shell.execute_reply.started":"2022-08-06T11:28:39.456949Z","shell.execute_reply":"2022-08-06T11:28:39.463210Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train = data_edit(df_train)\ndf_test = data_edit(df_test)\ndf_train, df_test = data_scaler(df_train,df_test)\ndf_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T11:28:39.952860Z","iopub.execute_input":"2022-08-06T11:28:39.953857Z","iopub.status.idle":"2022-08-06T11:28:40.054700Z","shell.execute_reply.started":"2022-08-06T11:28:39.953816Z","shell.execute_reply":"2022-08-06T11:28:40.053388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model creation and training","metadata":{}},{"cell_type":"markdown","source":"In this section, I will use **Gradient Boosting Classifier** and simply fit it to the training set. Then it predicts the outcome (transported or not) on the test set.","metadata":{}},{"cell_type":"code","source":"X_train = df_train.drop(\"Transported\",axis=1)\ny_train = df_train[\"Transported\"]\nX_test = df_test","metadata":{"execution":{"iopub.status.busy":"2022-08-06T00:28:03.334338Z","iopub.execute_input":"2022-08-06T00:28:03.334764Z","iopub.status.idle":"2022-08-06T00:28:03.342542Z","shell.execute_reply.started":"2022-08-06T00:28:03.334729Z","shell.execute_reply":"2022-08-06T00:28:03.341322Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"clf = GradientBoostingClassifier()\nclf.fit(X_train,y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T00:55:26.028630Z","iopub.execute_input":"2022-08-06T00:55:26.029075Z","iopub.status.idle":"2022-08-06T00:55:27.023969Z","shell.execute_reply.started":"2022-08-06T00:55:26.029031Z","shell.execute_reply":"2022-08-06T00:55:27.022774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_submission = pd.read_csv(\"../input/spaceship-titanic/sample_submission.csv\")\ndf_submission[\"Transported\"] = clf.predict(X_test)\ndf_submission.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T00:28:14.547590Z","iopub.execute_input":"2022-08-06T00:28:14.548005Z","iopub.status.idle":"2022-08-06T00:28:14.565318Z","shell.execute_reply.started":"2022-08-06T00:28:14.547967Z","shell.execute_reply":"2022-08-06T00:28:14.564221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**That's it! The .csv submission file gives a >0.80 score on the Kaggle leaderboard.** 🥳","metadata":{}},{"cell_type":"markdown","source":"<a id='Further'></a>\n# Further reading","metadata":{}},{"cell_type":"markdown","source":"In this second part, I will list several ideas to improve the dataset (such as feature engineering, better strategies for filling NaNs, etc.) in order to increase the score. There are infinite ways to do that, and I will cover a fraction of it.\n\nThese two top (and long) notebooks have a very deep investigation on the dataset:\n- https://www.kaggle.com/code/odins0n/spaceship-titanic-eda-27-different-models\n- https://www.kaggle.com/code/samuelcortinhas/spaceship-titanic-a-complete-guide","metadata":{}},{"cell_type":"markdown","source":"## Comparison with other classifiers","metadata":{}},{"cell_type":"markdown","source":"Scores of different classifiers on Kaggle leaderboard (replace the second last code cell above and beware of the type of the prediction - XGB yields \"0\" and \"1\" in `int`, CatBoost yields \"False\" and \"True\" in `str`):\n- GradientBoostingClassifier : 0.80313\n- AdaBoostClassifier : 0.79167\n- XGBClassifier : 0.79167\n- CatBoostClassifier : 0.80173\n\nMore information on the post : https://www.kaggle.com/competitions/spaceship-titanic/discussion/309535 \"Spaceship Titanic Compare BaseModels EveryBody Get Insight!\"","metadata":{}},{"cell_type":"markdown","source":"## Data engineering and preprocessing","metadata":{}},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\nimport numpy as np\nfrom ipywidgets import interact","metadata":{"execution":{"iopub.status.busy":"2022-08-07T16:09:58.912140Z","iopub.execute_input":"2022-08-07T16:09:58.912464Z","iopub.status.idle":"2022-08-07T16:09:58.918079Z","shell.execute_reply.started":"2022-08-07T16:09:58.912424Z","shell.execute_reply":"2022-08-07T16:09:58.917169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Recall the modifications on the datasets made previously:\n- Step 1: From the `Cabin` feature in the form of `deck/num/side`, get `desk` and `side`.\n- Step 2: Drop the colums `PassengerId`, `Cabin` and `Name`.\n- Step 3: Convert categorical variables into indicator variables with `pandas.get_dummies` function.\n- Step 4: Fill NaN values (for train and test sets) with most frequent values from the train set.\n- Step 5: Standardize float values of the train and test sets with a scaler fitted to the train set.\n\nAs one can see, this is obviously suboptimal. Here are some modifications to these 5 steps to leverage data.","metadata":{}},{"cell_type":"markdown","source":"- Step 1: The `Cabin` feature contains many information with its `deck/num/side` form, where `side` can be either `P` for Port or `S` for Starboard.\n\nAs shown below, the `T` deck is an outlier that can be replaced by another letter. The `num` can be gathered in four chunks.","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv(\"../input/spaceship-titanic/train.csv\")\ndf_temp = df[\"Cabin\"].str.split(\"/\", expand=True).rename(columns = {0:\"Deck\",1:\"Num\",2:\"Side\"})\ndf_temp[\"Transported\"] = df[\"Transported\"]\ndf_temp.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T16:09:59.436657Z","iopub.execute_input":"2022-08-07T16:09:59.436982Z","iopub.status.idle":"2022-08-07T16:09:59.483866Z","shell.execute_reply.started":"2022-08-07T16:09:59.436949Z","shell.execute_reply":"2022-08-07T16:09:59.483057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(16,4))\nplt.subplot(1,2,1)\nsns.countplot(data=df_temp.sort_values(by=\"Deck\"), x='Deck', hue='Transported') \nplt.title('Deck distribution')\nplt.subplot(1,2,2)\nsns.countplot(data=df_temp, x='Side', hue='Transported') \nplt.title('Side distribution')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T16:09:59.574985Z","iopub.execute_input":"2022-08-07T16:09:59.575555Z","iopub.status.idle":"2022-08-07T16:09:59.875378Z","shell.execute_reply.started":"2022-08-07T16:09:59.575520Z","shell.execute_reply":"2022-08-07T16:09:59.873885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_temp[\"Num\"] = df_temp[\"Num\"].astype(float)\nmaximum = max(df_temp[\"Num\"])\n\ndef plot_feature(bin_number):\n    divider = int(maximum/bin_number)\n    df_temp2 = df_temp.copy()\n    df_temp2[\"Num\"] = (df_temp2[\"Num\"]//divider)\n    plt.figure(figsize=(16,4))\n    sns.histplot(data=df_temp2, x='Num', hue='Transported', binwidth=1)\n    #sns.countplot(data=df_temp2, x='Num', hue='Transported') \n    plt.title('Num distribution')\n    plt.xticks(np.linspace(0,maximum//divider,41),np.linspace(0,maximum,41).astype(int),rotation=45)\n    plt.show()\n\ninteract(plot_feature, \n         bin_number = (0,80,5),\n        );","metadata":{"execution":{"iopub.status.busy":"2022-08-07T16:09:59.877021Z","iopub.execute_input":"2022-08-07T16:09:59.877290Z","iopub.status.idle":"2022-08-07T16:10:00.363959Z","shell.execute_reply.started":"2022-08-07T16:09:59.877265Z","shell.execute_reply":"2022-08-07T16:10:00.362548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The `Num` column can be transformed into four groups with for example the function\n\n`pd.cut(x = df_temp['Num'], bins = [0,290,700,1180,2000], labels = ['01', '02', '03', '04'])`","metadata":{}},{"cell_type":"markdown","source":"- Step 2: The columns`PassengerId` can be useful. I don't really know if `Name` is useful or not, the gender seems to be useless.\n\nEach Id takes the form `gggg_pp` where `gggg` indicates a group the passenger is travelling with and `pp` is their number within the group.\n\nAgain, as seen below, there a difference if the passenger travels alone or not, and the `gggg` can be gathered into groups.","metadata":{}},{"cell_type":"code","source":"df_temp = df[\"PassengerId\"].str.split(\"_\", expand=True).rename(columns = {0:\"GroupId\",1:\"GroupSize\"})\ndf_temp[\"Transported\"] = df[\"Transported\"]\n\ngroupsize = {}\nfor i in df_temp.index: # Works because PassengerId is already sorted\n    groupsize[df_temp.loc[i,\"GroupId\"]] = df_temp.loc[i,\"GroupSize\"]\nfor i in df_temp.index:\n    df_temp.loc[i,\"GroupSize\"] = groupsize[df_temp.loc[i,\"GroupId\"]]\ndf_temp = df_temp.assign(GroupId = df_temp[\"GroupId\"].astype(int))\ndf_temp.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T16:10:00.495886Z","iopub.execute_input":"2022-08-07T16:10:00.496315Z","iopub.status.idle":"2022-08-07T16:10:04.071218Z","shell.execute_reply.started":"2022-08-07T16:10:00.496288Z","shell.execute_reply":"2022-08-07T16:10:04.070114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(8,4))\nsns.histplot(data=df_temp.sort_values(by=\"GroupSize\"), x='GroupSize', hue='Transported', binwidth=1)\nplt.title('GroupSize distribution')","metadata":{"execution":{"iopub.status.busy":"2022-08-07T16:10:04.075068Z","iopub.execute_input":"2022-08-07T16:10:04.075386Z","iopub.status.idle":"2022-08-07T16:10:04.304432Z","shell.execute_reply.started":"2022-08-07T16:10:04.075363Z","shell.execute_reply":"2022-08-07T16:10:04.303506Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"maximum = max(df_temp[\"GroupId\"])\n\ndef plot_feature(bin_number):\n    divider = int(maximum/bin_number)\n    df_temp2 = df_temp.copy()\n    df_temp2[\"GroupId\"] = (df_temp2[\"GroupId\"]//divider)\n    plt.figure(figsize=(16,4))\n    sns.histplot(data=df_temp2, x='GroupId', hue='Transported', binwidth=1)\n    #sns.countplot(data=df_temp2, x='GroupId', hue='Transported') \n    plt.title('gggg distribution')\n    plt.xticks(np.linspace(0,maximum//divider,bin_number+1),\n               np.linspace(0,maximum,bin_number+1).astype(int),rotation=45)\n    plt.show()\n\ninteract(plot_feature, \n         bin_number = (1,40,1),\n        );","metadata":{"execution":{"iopub.status.busy":"2022-08-07T16:10:04.306236Z","iopub.execute_input":"2022-08-07T16:10:04.306521Z","iopub.status.idle":"2022-08-07T16:10:04.678386Z","shell.execute_reply.started":"2022-08-07T16:10:04.306497Z","shell.execute_reply":"2022-08-07T16:10:04.677195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Similar results can be found for the `Age` column, adjust the slider to find the best bins to group the values.","metadata":{}},{"cell_type":"markdown","source":"- Step ~~3~~ 6: Convert categorical variables into indicator variables with `pandas.get_dummies` function **only at the end** after filling the NaNs, tweaking the data and columns.","metadata":{}},{"cell_type":"markdown","source":"- Step 4: Fill NaN values with better strategies. Try to find patterns and your own rules. Remember that these rules must come from the train set to avoid data leakage.\n\nSome strategies can be found in this discussion post: https://www.kaggle.com/competitions/spaceship-titanic/discussion/315987 \"Some rules to fill NaNs\"\n\nI will discuss the `HomePlanet` column for example. People in the same group (in `PassengerId`) are from the same planet, so the HomePlanet can be filled if one knows that of the other passengers. Otherwise, if they travel alone, by looking at the repartition below, one can assign the HomePlanet using the Deck with this rule:\n```F, G : earth\nA, B, C, T : europa\nD, E : mars```","metadata":{}},{"cell_type":"code","source":"df_temp = df.copy()\ndf_temp[[\"Deck\",\"Num\",\"Side\"]] = df_temp[\"Cabin\"].str.split(\"/\", expand=True)\njoint=df_temp.groupby(['Deck','HomePlanet'])['HomePlanet'].size().unstack().fillna(0)\n\n# Heatmap of missing values\nplt.figure(figsize=(16,5))\nsns.heatmap(joint.T, annot=True, fmt='g')","metadata":{"execution":{"iopub.status.busy":"2022-08-07T16:15:19.459072Z","iopub.execute_input":"2022-08-07T16:15:19.459668Z","iopub.status.idle":"2022-08-07T16:15:19.690624Z","shell.execute_reply.started":"2022-08-07T16:15:19.459642Z","shell.execute_reply":"2022-08-07T16:15:19.689319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Step 5: **Logscale** then standardize float values of the train and test sets with a scaler fitted to the train set.\n\nI'll use some of the code from the notebook \"🚀 Spaceship Titanic: A complete guide 🏆\" mentioned earlier. Again, you should check it if you have some time!","metadata":{}},{"cell_type":"code","source":"df_normal = df.copy()\ndf_log = df.copy()\n\n# Expenditure features\nexp_feats=['RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']\n\n# Logarithmic scale\ndef logarithm(row,col):\n        row[col] = np.log(row[col]+1)\n        return row\nfor col in exp_feats:\n    df_log = df_log.apply(lambda x: logarithm(x,col),axis='columns')\n\n# Plot expenditure features\nfig=plt.figure(figsize=(16,20))\nfor i, var_name in enumerate(exp_feats):\n    # Left plot normal scale\n    ax=fig.add_subplot(5,4,4*i+1)\n    sns.histplot(data=df_normal, x=var_name, axes=ax, bins=30, kde=False, hue='Transported')\n    ax.set_title(var_name)\n    \n    # 2nd left plot normal scale (truncated)\n    ax=fig.add_subplot(5,4,4*i+2)\n    sns.histplot(data=df_normal, x=var_name, axes=ax, bins=30, kde=False, hue='Transported')\n    plt.ylim([0,200])\n    ax.set_title(var_name)\n    \n    # 2nd right plot after passing to the logarithm\n    ax=fig.add_subplot(5,4,4*i+3)\n    sns.histplot(data=df_log, x=var_name, axes=ax, bins=30, kde=True, hue='Transported')\n    ax.set_title(var_name)\n    \n    # Right plot after passing to the logarithm (truncated)\n    ax=fig.add_subplot(5,4,4*i+4)\n    sns.histplot(data=df_log, x=var_name, axes=ax, bins=30, kde=True, hue='Transported')\n    plt.ylim([0,200])\n    ax.set_title(var_name)\nfig.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T16:36:00.276739Z","iopub.execute_input":"2022-08-07T16:36:00.277086Z","iopub.status.idle":"2022-08-07T16:36:10.299755Z","shell.execute_reply.started":"2022-08-07T16:36:00.277060Z","shell.execute_reply":"2022-08-07T16:36:10.298713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Switching to logarithmic scale (3rd and 4th column) gives a better view and feel of the data: people who don't spend are more likely to be transported, people who spend the most in `FoodCourt` and `ShoppingMall` as well, but the rest are more likely to stay in the spaceship.\n\nIt's also a good idea to create a column `No_spending` to see if a passenger have spent or not. `No_spending` will be surely positively correlated to the `Transported` column.","metadata":{}},{"cell_type":"markdown","source":"I will stop here, my initial goal was to show that it is possible to get a very quick first feel of the data without having to read an overwhelming notebook. I hope this has helped you.\n\nThanks for reading to the end! Good luck for getting a 0.80 score and even higher!","metadata":{}}]}