{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Starting Guide to Spaceship Titanic EDA and Random Forest Model with Evaluation","metadata":{"id":"v7nXbwV_McUV"}},{"cell_type":"markdown","source":"## Notebook Description","metadata":{"id":"0Voqsl8MQM6w"}},{"cell_type":"markdown","source":"This notebook provides:\n* in-depth, feature-by-feature analysis of the Spaceship Titanic data;\n* data cleaning, preprocessing, and feature engineering;\n* training a Random Forest model;\n* evaluating model outputs.\n\n\nCredit to numerous other notebooks that provided inspiration and ideas incorporated here:\n* [Spaceship Titanic: A Complete Guide](https://www.kaggle.com/code/samuelcortinhas/spaceship-titanic-a-complete-guide)\n* [Spaceship Titanic EDA, XGBoost 80%+](https://www.kaggle.com/code/eisgandar/spaceship-titanic-eda-xgboost-80)\n\n\n**Please provide feedback in the comments which will help me continue to improve notebooks as I publish them!**","metadata":{"id":"s82IcyxQQQRi"}},{"cell_type":"markdown","source":"## Import Libraries","metadata":{"id":"HEeBOntHQR8T"}},{"cell_type":"code","source":"# data analysis \nimport numpy as np\nimport pandas as pd\npd.set_option('display.max_columns', None)\npd.set_option('display.max_rows', None) \n\n# data visualization \nimport matplotlib.pyplot as plt  \nimport seaborn as sns\n\n# general utilities \n# !pip install kaggle\nimport os\n\n# sklearn tools and models\nfrom sklearn.model_selection import GridSearchCV, RandomizedSearchCV\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.metrics import accuracy_score, confusion_matrix, classification_report\n\n","metadata":{"id":"yUFCry1ZMiR0","execution":{"iopub.status.busy":"2022-08-10T11:56:35.263133Z","iopub.execute_input":"2022-08-10T11:56:35.264728Z","iopub.status.idle":"2022-08-10T11:56:35.276951Z","shell.execute_reply.started":"2022-08-10T11:56:35.264674Z","shell.execute_reply":"2022-08-10T11:56:35.275618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Custom Functions\n\n*Functions written for use as part of the analysis of this data.*","metadata":{"id":"ovi6nMbZHrnX"}},{"cell_type":"code","source":"def missing_pct(data_lst, col):\n  '''Calculate the percentage of missing values from a feature in\n      both the train and test dataframes\n\n  Args:\n    data_lst: a list containing train and test dataframes, in that order \n    col: the specific feature (column) to calculate the missing values\n  \n  Returns: Text containing the rounded percentage of missing \n          features in the train and test data \n  '''\n\n  print(f'Percentage of {col} Missing Values in Training Data:', round(data_lst[0][col].isna().sum() / data_lst[0].shape[0], 3))\n  print(f'Percentage of {col} Missing Values in Test Data:', round(data_lst[1][col].isna().sum() / data_lst[1].shape[0], 3))\n","metadata":{"id":"8-aNI6J7L-p7","execution":{"iopub.status.busy":"2022-08-10T11:33:59.142964Z","iopub.execute_input":"2022-08-10T11:33:59.143399Z","iopub.status.idle":"2022-08-10T11:33:59.150812Z","shell.execute_reply.started":"2022-08-10T11:33:59.143364Z","shell.execute_reply":"2022-08-10T11:33:59.149640Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def transported_NAs(df, col):\n  '''Given a feature, determine how many passengers \n      with NAs were transported\n      \n  Args:\n    df: training df that contains 'Transported' column\n    col: specific column to test for missing data \n  \n  Returns: percentage of passengers with NAs in the col who were transported\n  '''\n  print(f'Analyzing {col}:')\n  print(\n      round(\n        df[df[col].isna()][['Transported']]\n        .mean()\n        .rename({'Transported':'Percentage of NAs Transported:'}), 3\n  ))","metadata":{"id":"ZnGMhPRVFPAP","execution":{"iopub.status.busy":"2022-08-10T11:33:59.742718Z","iopub.execute_input":"2022-08-10T11:33:59.743108Z","iopub.status.idle":"2022-08-10T11:33:59.749831Z","shell.execute_reply.started":"2022-08-10T11:33:59.743077Z","shell.execute_reply":"2022-08-10T11:33:59.748284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_categorical_barh(df: pd.DataFrame, \n                     col: str, \n                     title: str,\n                     figsize: tuple = (10,5),\n                     color: list = ['b']) -> None:\n  '''Generate hbar plot of one categorical dataframe feature\n\n  Args:\n    df: pandas DataFrame containing the data \n    col: categorical column to produce the barh chart \n    title: title of the barh chart \n    figsize: tuple containing chart dimensions; defaults to (10,5)\n    color: chart color; defaults to blue \n\n  Returns:\n    no object returned; plots the barh chart \n  '''\n  (\n    df[col]\n    .value_counts(normalize=True)\n    .sort_values(ascending=True)\n    .plot(kind='barh',\n          figsize=figsize,\n          title=title,\n          color=color)\n  )\n\n  plt.show()","metadata":{"id":"vwQ-xbWVQ5de","execution":{"iopub.status.busy":"2022-08-10T11:34:00.341193Z","iopub.execute_input":"2022-08-10T11:34:00.342222Z","iopub.status.idle":"2022-08-10T11:34:00.348482Z","shell.execute_reply.started":"2022-08-10T11:34:00.342183Z","shell.execute_reply":"2022-08-10T11:34:00.347552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_transported_pct(df: pd.DataFrame, \n                         col: str, \n                         title: str,\n                         figsize: tuple = (10,5),\n                         color: str = 'r') -> None:\n  '''Generate a horizontal bar chart of transported pct for a category\n\n  Args:\n    df: pandas dataframe containing the data \n    col: column to analyze transported percentage \n\n  Returns:\n    None; plots a chart of the pct of passengers transported\n  '''\n  (\n    df[[col, 'Transported']]\n    .groupby(col)\n    .agg('mean')\n    .sort_values('Transported')\n    .plot(kind='barh',\n          figsize=figsize,\n          color = color,\n          title=title)\n  )\n  plt.show()","metadata":{"id":"RNngKul900R_","execution":{"iopub.status.busy":"2022-08-10T11:34:00.981540Z","iopub.execute_input":"2022-08-10T11:34:00.982210Z","iopub.status.idle":"2022-08-10T11:34:00.990408Z","shell.execute_reply.started":"2022-08-10T11:34:00.982153Z","shell.execute_reply":"2022-08-10T11:34:00.988973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Import Data ","metadata":{"id":"siatd-GhPpQR"}},{"cell_type":"code","source":"train_df = pd.read_csv(\"../input/spaceship-titanic/train.csv\")\ntest_df = pd.read_csv(\"../input/spaceship-titanic/test.csv\")\n\nprint('train dataframe dimensions:', train_df.shape)\nprint('test dataframe dimensions:', test_df.shape)","metadata":{"id":"NceohZoKRhDo","outputId":"d96bc8d5-aa10-4fa2-c081-f7f017a8af3d","execution":{"iopub.status.busy":"2022-08-10T11:29:02.608587Z","iopub.execute_input":"2022-08-10T11:29:02.608978Z","iopub.status.idle":"2022-08-10T11:29:02.690078Z","shell.execute_reply.started":"2022-08-10T11:29:02.608947Z","shell.execute_reply":"2022-08-10T11:29:02.689197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# combine into a single list\nfull_data = [train_df, test_df]\nlen(full_data)","metadata":{"id":"TZ1Xa52pYwjs","outputId":"187332ea-b9bd-4688-ad6f-7e8a941d687a","execution":{"iopub.status.busy":"2022-08-10T11:31:09.653801Z","iopub.execute_input":"2022-08-10T11:31:09.654315Z","iopub.status.idle":"2022-08-10T11:31:09.664454Z","shell.execute_reply.started":"2022-08-10T11:31:09.654255Z","shell.execute_reply":"2022-08-10T11:31:09.663153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##  Exploratory Data Analysis","metadata":{"id":"BHzURwehQILI"}},{"cell_type":"markdown","source":"### High-Level Data Overview","metadata":{"id":"QZUWas_tVtUv"}},{"cell_type":"markdown","source":"**Train Data**","metadata":{"id":"cMcRuFN1ZMic"}},{"cell_type":"code","source":"train_df.info()","metadata":{"id":"wU6x626yVhTx","outputId":"bc3aa42d-5f9d-48b8-9ba2-6b1504633214","execution":{"iopub.status.busy":"2022-08-10T11:31:12.203899Z","iopub.execute_input":"2022-08-10T11:31:12.204390Z","iopub.status.idle":"2022-08-10T11:31:12.242670Z","shell.execute_reply.started":"2022-08-10T11:31:12.204352Z","shell.execute_reply":"2022-08-10T11:31:12.241654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"id":"RxcHyJlRVhPZ","outputId":"752a73e7-ec68-4722-f135-3c105065978c","execution":{"iopub.status.busy":"2022-08-10T11:31:12.958153Z","iopub.execute_input":"2022-08-10T11:31:12.958997Z","iopub.status.idle":"2022-08-10T11:31:12.986217Z","shell.execute_reply.started":"2022-08-10T11:31:12.958954Z","shell.execute_reply":"2022-08-10T11:31:12.985420Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Test Data**","metadata":{"id":"vGw196ICVlmi"}},{"cell_type":"code","source":"test_df.info()","metadata":{"id":"iEbMTBpQVoGj","outputId":"666aed37-61df-4f63-b199-96072748ab0e","execution":{"iopub.status.busy":"2022-08-10T11:31:14.424878Z","iopub.execute_input":"2022-08-10T11:31:14.425958Z","iopub.status.idle":"2022-08-10T11:31:14.442004Z","shell.execute_reply.started":"2022-08-10T11:31:14.425917Z","shell.execute_reply":"2022-08-10T11:31:14.441093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head()","metadata":{"id":"kpvMhawyVn_R","outputId":"ffaf6c6f-2776-4bed-c2bf-7e451b483611","execution":{"iopub.status.busy":"2022-08-10T11:31:15.046223Z","iopub.execute_input":"2022-08-10T11:31:15.046643Z","iopub.status.idle":"2022-08-10T11:31:15.067102Z","shell.execute_reply.started":"2022-08-10T11:31:15.046609Z","shell.execute_reply":"2022-08-10T11:31:15.066194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Analysis**\n* The training data is approximately double the size of the test data, providing an overall breakdown of routhly 2/3 training and 1/3 test.\n* With such a large test data set, we may need to check if there are any unknown features (i.e., categorical dummy variables) and make sure they are incorporated into the test data as features to prevent modeling errors. \n* Features in both the train and test data appear to be missing a small percentage of values; no features seem to have substantial missing data.","metadata":{"id":"xwbYL5w7BT-p"}},{"cell_type":"markdown","source":"### Descriptive Statistics","metadata":{"id":"VVfvF6HlW9mi"}},{"cell_type":"markdown","source":"**Descriptive Statistics -- Quantitative Features**","metadata":{"id":"nOb5Bg4MYha_"}},{"cell_type":"code","source":"train_df.describe().round().T","metadata":{"id":"wc_Cdd7FW9sS","outputId":"2b626c05-fdc3-4836-a64e-622cb6754249","execution":{"iopub.status.busy":"2022-08-10T11:31:17.336939Z","iopub.execute_input":"2022-08-10T11:31:17.337582Z","iopub.status.idle":"2022-08-10T11:31:17.381320Z","shell.execute_reply.started":"2022-08-10T11:31:17.337543Z","shell.execute_reply":"2022-08-10T11:31:17.380151Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.describe().round().T","metadata":{"id":"8Kp_eJV8Hsu4","outputId":"5ac19b30-e003-4469-fb53-832718bd43f5","execution":{"iopub.status.busy":"2022-08-10T11:31:17.982409Z","iopub.execute_input":"2022-08-10T11:31:17.983556Z","iopub.status.idle":"2022-08-10T11:31:18.023123Z","shell.execute_reply.started":"2022-08-10T11:31:17.983521Z","shell.execute_reply":"2022-08-10T11:31:18.021799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*Analysis of Training Data:*\n* Age: The youngest passenger is 0 years old, which could be an error or represent a newborn. The oldest is 79 with the mean being 29 and the median being 27 (indicating a slight skew). The train and test data sets are similar in this regard.\n* Spend Categories: Based on the median values being zero, most passengers don't seem to spend in the various locations. The data also appear to be skewed due to passengers who spend relatively large amounts in each of the categories.\n  * The test data spend amounts are generally lower than the train data across the board.\n  * In the test data the max ShoppingMall amount is only 8,292, which is far lower than the max of 23,492 in the train data. ","metadata":{"id":"1mbfKVOTW90p"}},{"cell_type":"markdown","source":"**Descriptive Statistics -- Categorical Data**","metadata":{"id":"c5lfse2jZ1Cx"}},{"cell_type":"code","source":"train_df.describe(include='object').round().T","metadata":{"id":"yDxcxCqUZ07v","outputId":"4b6bc88a-dbde-4c06-d6ad-b4b50316af35","execution":{"iopub.status.busy":"2022-08-10T11:31:21.120486Z","iopub.execute_input":"2022-08-10T11:31:21.121524Z","iopub.status.idle":"2022-08-10T11:31:21.165903Z","shell.execute_reply.started":"2022-08-10T11:31:21.121481Z","shell.execute_reply":"2022-08-10T11:31:21.165059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.describe(include='object').round().T","metadata":{"id":"yUZzhKKkMJrG","outputId":"de9b476a-b45e-401f-dbae-8a7cb843a068","execution":{"iopub.status.busy":"2022-08-10T11:31:21.695587Z","iopub.execute_input":"2022-08-10T11:31:21.696272Z","iopub.status.idle":"2022-08-10T11:31:21.731894Z","shell.execute_reply.started":"2022-08-10T11:31:21.696225Z","shell.execute_reply":"2022-08-10T11:31:21.731041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*Analysis:*\n* Train and Test data appears to be generally similar:\n  * HomePlanet: Three unique values with Earth as the most common, comprising slightly more than 50%.\n  * CryoSleep: Boolean with False comprising almost 65%.\n  * Destination: Three unique values with TRAPPIST-1e comprising 70%.\n  * VIP: Boolean with False comprising 99%. \n\nPassengerId, Cabin, and Name all comprise mostly unique values, although there appear to be passengers with duplicate names which we will need to investigate.","metadata":{"id":"65tcykggZ02o"}},{"cell_type":"markdown","source":"### Overview of Missing Values","metadata":{"id":"Unm6zoe2aF4-"}},{"cell_type":"code","source":"print('Percentage of missing data in train and test data sets:\\n')\nfor df in full_data:\n  (\n      df.isna().mean()\n      .plot(kind='barh',\n            figsize=(10,5))\n  )\n  plt.show()\n  print('')","metadata":{"id":"ofz47eameQu2","outputId":"e0a711ce-585e-4cfd-994b-af1f7cbdf773","execution":{"iopub.status.busy":"2022-08-10T11:31:23.369432Z","iopub.execute_input":"2022-08-10T11:31:23.370066Z","iopub.status.idle":"2022-08-10T11:31:23.909866Z","shell.execute_reply.started":"2022-08-10T11:31:23.370032Z","shell.execute_reply":"2022-08-10T11:31:23.908622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Analysis**\n* PassengerId contains zero missing values in the train and test data sets \n* In the train data set, Transported also contains zero missing values \n* All other features in both the train and test data sets contain roughly 0.02 to 0.025 percent missing values\n* For the train data the 0.02 correlates to approx. 180-200 values per feature ","metadata":{"id":"lq46KjdfcxOv"}},{"cell_type":"markdown","source":"## Summary of Data Preprocessing","metadata":{"id":"r7WKb7T7TTK3"}},{"cell_type":"markdown","source":"### Summary Steps\n* **PassengerId:** parse into Passenger_Group, Passenger_Num, and Group_Size\n* **HomePlanet:** impute most frequent and convert to categorical\n* **CryoSleep:** impute most frequent; convert to boolean\n* **Cabin:** Fill missing values with 'Z/99999/Z' and parse into new Cabin_Deck, Cabin_Num, and Cabin_Side features\n* **Destination:** impute most frequent and convert to categorical\n* **Age:** Impute median value\n* **VIP:** Impute most frequent and convert to boolean\n* **RoomService:** Impute median value\n* **FoodCourt:** Impute median value\n* **ShoppingMall:** Impute median value\n* **Spa:** Impute median value\n* **VRDeck:** Impute median value\n* **Name:** impute 'NoFirstName NoLastName'\n* **Transported:** converted to Boolean; train data only\n\n","metadata":{"id":"qqDfJAVHTS4A"}},{"cell_type":"markdown","source":"### Preprocess Data","metadata":{"id":"9mf_DXK3TSun"}},{"cell_type":"code","source":"# convert Transported to boolean; only in train_df \ntrain_df['Transported'] = train_df['Transported'].astype(bool)\n\n# format bulk of data sets \nfor df in full_data:\n\n  # IMPUTE MISSING VALUES \n  df['Name'].fillna('NoFirstName NoSurname', inplace=True)\n  df['HomePlanet'].fillna(df['HomePlanet'].value_counts().index[0], inplace=True)\n  df['CryoSleep'].fillna(df['CryoSleep'].value_counts().index[0], inplace=True)\n  df['Cabin'].fillna('Z/99999/Z', inplace=True)\n  df['Destination'].fillna(df['Destination'].value_counts().index[0], inplace=True)\n  df['VIP'].fillna(df['VIP'].value_counts().index[0], inplace=True)\n  df['Age'].fillna(df['Age'].median(), inplace=True)\n  df['RoomService'].fillna(df['RoomService'].median(), inplace=True)\n  df['FoodCourt'].fillna(df['FoodCourt'].median(), inplace=True)\n  df['ShoppingMall'].fillna(df['ShoppingMall'].median(), inplace=True)\n  df['Spa'].fillna(df['Spa'].median(), inplace=True)\n  df['VRDeck'].fillna(df['VRDeck'].median(), inplace=True)\n\n  # CREATE FEATURES \n  df['Passenger_Group'] = df['PassengerId'].apply(lambda x: x.split('_')[0]).astype(str)\n  df['Passenger_Num'] = df['PassengerId'].apply(lambda x: x.split('_')[1]).astype(str)\n  df['Cabin_Deck'] = df['Cabin'].apply(lambda x: x.split('/')[0]).astype(str)\n  df['Cabin_Num'] = df['Cabin'].apply(lambda x: x.split('/')[1]).astype(str)\n  df['Cabin_Side'] = df['Cabin'].apply(lambda x: x.split('/')[2]).astype(str) \n\n  # FORMAT DATA TYPES \n  df['HomePlanet'] = df['HomePlanet'].astype('category')\n  df['CryoSleep'] = df['CryoSleep'].astype('bool')\n  df['Cabin_Deck'] = df['Cabin_Deck'].astype('category')\n  df['Cabin_Side'] = df['Cabin_Side'].astype('category')\n  df['Destination'] = df['Destination'].astype('category')\n\n\n# use completed features to generate new features\nfor df in full_data:\n  df['Group_Size'] = df['Passenger_Group'].map(lambda x: pd.concat([train_df['Passenger_Group'], \n                                                                    test_df['Passenger_Group']]).value_counts()[x])\n\n\n","metadata":{"id":"UmX4sMctZZdp","execution":{"iopub.status.busy":"2022-08-10T11:31:29.515970Z","iopub.execute_input":"2022-08-10T11:31:29.516408Z","iopub.status.idle":"2022-08-10T11:32:35.229702Z","shell.execute_reply.started":"2022-08-10T11:31:29.516372Z","shell.execute_reply":"2022-08-10T11:32:35.228380Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.info()","metadata":{"id":"2SlDFGoeQ3f3","outputId":"787c277b-5a94-41de-b3cc-15f118a62b9a","execution":{"iopub.status.busy":"2022-08-10T11:32:35.231491Z","iopub.execute_input":"2022-08-10T11:32:35.231812Z","iopub.status.idle":"2022-08-10T11:32:35.253585Z","shell.execute_reply.started":"2022-08-10T11:32:35.231783Z","shell.execute_reply":"2022-08-10T11:32:35.252378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.info()","metadata":{"id":"U4tHA6U0Q3oy","outputId":"ede00d69-a2b7-4f1a-8319-5c5c16c3349a","execution":{"iopub.status.busy":"2022-08-10T11:32:35.255109Z","iopub.execute_input":"2022-08-10T11:32:35.255551Z","iopub.status.idle":"2022-08-10T11:32:35.274081Z","shell.execute_reply.started":"2022-08-10T11:32:35.255516Z","shell.execute_reply":"2022-08-10T11:32:35.272572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Categorical Feature Analysis\n\n*Note: This analysis will generally focus on the training dataset*","metadata":{"id":"Gt2HlTRffTZs"}},{"cell_type":"markdown","source":"### Transported\n*Whether the passenger was transported to another dimension. This is the target, the column we are trying to predict.*","metadata":{"id":"tCTdYLfN7OFJ"}},{"cell_type":"code","source":"plot_categorical_barh(df=train_df, col='Transported', title='Percentage of Passengers Transported')","metadata":{"id":"DDBayUgp7OLI","outputId":"82254d67-64fd-4967-8531-4db529a67dc3","execution":{"iopub.status.busy":"2022-08-10T11:33:04.147137Z","iopub.execute_input":"2022-08-10T11:33:04.147553Z","iopub.status.idle":"2022-08-10T11:33:04.340577Z","shell.execute_reply.started":"2022-08-10T11:33:04.147517Z","shell.execute_reply":"2022-08-10T11:33:04.339453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Analysis of Transported:**\n* Approximately half of the passengers were transported, with slightly more being transported than not. \n* At this point we don't have a way to confirm whether the test data has the same distribution.","metadata":{"id":"JgR4R9939NH3"}},{"cell_type":"markdown","source":"### PassengerId\n*A unique Id for each passenger. Each Id takes the form gggg_pp where gggg indicates a group the passenger is travelling with and pp is their number within the group. People in a group are often family members, but not always.*\n\nWe used PassengerId to create additional features: \n* *Passenger_Group* for each group \n* *Passenger_Num* for each group member \n* *Group_Size* to count the number of members in each group","metadata":{"id":"jXmIU4eiR7mb"}},{"cell_type":"markdown","source":"### Passenger Group Size\nNumber of members traveling as part of a specific passenger group. This feature was created from PassengerId.","metadata":{"id":"47_SBWqzBS-n"}},{"cell_type":"code","source":"plot_categorical_barh(df=train_df, col='Group_Size', \n                 title='Distribution of Passenger Group Sizes -- Train Data')","metadata":{"id":"7FSVYCLRbMqD","outputId":"b1a124e8-1164-4617-aa7f-167eb6e7abf9","execution":{"iopub.status.busy":"2022-08-10T11:33:07.024092Z","iopub.execute_input":"2022-08-10T11:33:07.024552Z","iopub.status.idle":"2022-08-10T11:33:07.244946Z","shell.execute_reply.started":"2022-08-10T11:33:07.024517Z","shell.execute_reply":"2022-08-10T11:33:07.243703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_categorical_barh(df=train_df, col='Group_Size', \n                 title='Distribution of Passenger Group Sizes -- Test Data',\n                 color=['g'])","metadata":{"id":"dbIEsXsAlkYg","outputId":"1c4f179b-cbb1-40f2-b255-28e8e954e7c3","execution":{"iopub.status.busy":"2022-08-10T11:33:07.595986Z","iopub.execute_input":"2022-08-10T11:33:07.596425Z","iopub.status.idle":"2022-08-10T11:33:07.816069Z","shell.execute_reply.started":"2022-08-10T11:33:07.596384Z","shell.execute_reply":"2022-08-10T11:33:07.814927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_transported_pct(df=train_df, \n                     col='Group_Size',\n                     title='Percentage Transported by Group Size -- Train Data')","metadata":{"id":"ejOxmTFfnqz2","outputId":"a03d5c01-26a6-444c-cd6a-524496d0f0c1","execution":{"iopub.status.busy":"2022-08-10T11:33:08.317929Z","iopub.execute_input":"2022-08-10T11:33:08.318948Z","iopub.status.idle":"2022-08-10T11:33:08.561751Z","shell.execute_reply.started":"2022-08-10T11:33:08.318904Z","shell.execute_reply":"2022-08-10T11:33:08.560650Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Analysis of Group Size**\n* Among all passengers, approximately half were transported to another dimension. \n* When analyzed by group size, it appears that passengers traveling in groups with two to seven members were between 4% and 14% more likely to be transported.\n* Passengers traveling alone were 5% less likely to be transported than the average.","metadata":{"id":"UEGgdq4h5w5T"}},{"cell_type":"markdown","source":"### Passenger Name\n*The first and last names of the passenger.*","metadata":{"id":"lV0hVjmSTv4w"}},{"cell_type":"markdown","source":"**Passengers with Duplicate Names**","metadata":{"id":"k4PF1I9dRlhp"}},{"cell_type":"code","source":"duplicate_names = train_df['Name'].value_counts()[train_df['Name'].value_counts() == 2].index\nprint('Number of passengers with same name:', len(duplicate_names))","metadata":{"id":"OQ-6C822cESS","outputId":"7afcb9cd-8946-4add-d70f-f1e73b275296","execution":{"iopub.status.busy":"2022-08-10T11:34:14.468915Z","iopub.execute_input":"2022-08-10T11:34:14.469443Z","iopub.status.idle":"2022-08-10T11:34:14.484949Z","shell.execute_reply.started":"2022-08-10T11:34:14.469399Z","shell.execute_reply":"2022-08-10T11:34:14.483904Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['Name'].value_counts()[train_df['Name'].value_counts() == 2]","metadata":{"id":"-9vJn8gMZBoe","outputId":"6d34d14e-dbba-4563-d845-04a28a9503a6","execution":{"iopub.status.busy":"2022-08-10T11:34:15.026336Z","iopub.execute_input":"2022-08-10T11:34:15.027062Z","iopub.status.idle":"2022-08-10T11:34:15.044657Z","shell.execute_reply.started":"2022-08-10T11:34:15.027024Z","shell.execute_reply":"2022-08-10T11:34:15.043182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# dataframe of just passengers with duplicate names\nduplicate_names_df = train_df[train_df['Name'].isin(list(duplicate_names))].copy()\nduplicate_names_df.shape","metadata":{"id":"P0dUzdaocEOL","outputId":"c0cce1ff-60e8-49dd-93bd-b151d6dd961d","execution":{"iopub.status.busy":"2022-08-10T11:34:19.078959Z","iopub.execute_input":"2022-08-10T11:34:19.080365Z","iopub.status.idle":"2022-08-10T11:34:19.094639Z","shell.execute_reply.started":"2022-08-10T11:34:19.080308Z","shell.execute_reply":"2022-08-10T11:34:19.093363Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"duplicate_names_df.sort_values(by='Name', ascending=True).head(6)","metadata":{"id":"iKocZQflpzkJ","execution":{"iopub.status.busy":"2022-08-10T11:34:23.336169Z","iopub.execute_input":"2022-08-10T11:34:23.336551Z","iopub.status.idle":"2022-08-10T11:34:23.368465Z","shell.execute_reply.started":"2022-08-10T11:34:23.336520Z","shell.execute_reply":"2022-08-10T11:34:23.367342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# likelihood of being transported with duplicate name \nduplicate_names_df['Transported'].value_counts(normalize=True)","metadata":{"id":"Dw4PzvMWVK_l","outputId":"f64bf0d1-cfcd-4aaa-ca55-523b7de7840f","execution":{"iopub.status.busy":"2022-08-10T11:34:24.737060Z","iopub.execute_input":"2022-08-10T11:34:24.737465Z","iopub.status.idle":"2022-08-10T11:34:24.746953Z","shell.execute_reply.started":"2022-08-10T11:34:24.737432Z","shell.execute_reply":"2022-08-10T11:34:24.745635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# transported_NAs(df=train_df, col='Name')","metadata":{"id":"QBWRoAQHVz6F","outputId":"27b43a62-da74-4db2-bc44-4fded99885d8","execution":{"iopub.status.busy":"2022-08-10T11:34:42.257290Z","iopub.execute_input":"2022-08-10T11:34:42.257676Z","iopub.status.idle":"2022-08-10T11:34:42.262561Z","shell.execute_reply.started":"2022-08-10T11:34:42.257645Z","shell.execute_reply":"2022-08-10T11:34:42.261379Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Analysis of Passenger Names**\n* We filled the missing values with 'NoFirstName NoSurname'\n* There are 20 passengers with the same name, but overall these are different entries.\n* Passengers with duplicate names appear to have a slightly lower likelihood of being transported, but it is a small data set.","metadata":{"id":"rFeU_xQaTA-W"}},{"cell_type":"markdown","source":"### Home Planet\n*Definition: The planet the passenger departed from, typically their planet of permanent residence.*","metadata":{"id":"-vZ_wxIDB5Di"}},{"cell_type":"code","source":"plot_categorical_barh(df=train_df, \n                      col='HomePlanet',\n                      color=['y', 'b', 'g'],\n                      title='Distribution of Passenger Home Planets -- Train Data')","metadata":{"id":"XVKu__eQqcEJ","outputId":"b2be77b5-de7a-4305-a4f8-74e25ff0c764","execution":{"iopub.status.busy":"2022-08-10T11:34:45.045580Z","iopub.execute_input":"2022-08-10T11:34:45.045983Z","iopub.status.idle":"2022-08-10T11:34:45.245667Z","shell.execute_reply.started":"2022-08-10T11:34:45.045950Z","shell.execute_reply":"2022-08-10T11:34:45.244388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_transported_pct(df=train_df,\n                     col='HomePlanet',\n                     title='Percentage Transported by Home Planet -- Train Data')","metadata":{"id":"NgtcIM_vrOHY","outputId":"bdd8b62f-37d0-44fe-b7d1-f9f0cf32427c","execution":{"iopub.status.busy":"2022-08-10T11:34:45.707351Z","iopub.execute_input":"2022-08-10T11:34:45.707857Z","iopub.status.idle":"2022-08-10T11:34:45.924776Z","shell.execute_reply.started":"2022-08-10T11:34:45.707812Z","shell.execute_reply":"2022-08-10T11:34:45.923808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Analysis of HomePlanet**\n* A majority of passengers are from Earth.\n* However, passengers from Earth appear to be almost 10% less likely to be transported, while passengers from Europa are 10%+ more likely to be transported. ","metadata":{"id":"wIEJDSRseGyg"}},{"cell_type":"markdown","source":"### CryoSleep\n*Indicates whether the passenger elected to be put into suspended animation for the duration of the voyage. Passengers in cryosleep are confined to their cabins.*","metadata":{"id":"15xnTyKperX6"}},{"cell_type":"code","source":"plot_categorical_barh(df=train_df, \n                      col='CryoSleep',\n                      title='Distribution of Passengers in CryoSleep -- Train Data')","metadata":{"id":"A3zDNASMykUc","outputId":"713747e1-5712-4293-cb9d-f6b792ad843f","execution":{"iopub.status.busy":"2022-08-10T11:34:50.685041Z","iopub.execute_input":"2022-08-10T11:34:50.686344Z","iopub.status.idle":"2022-08-10T11:34:50.878797Z","shell.execute_reply.started":"2022-08-10T11:34:50.686301Z","shell.execute_reply":"2022-08-10T11:34:50.877885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_transported_pct(df=train_df,\n                     col='CryoSleep',\n                     title='Passengers in CryoSleep who were Transported -- Train Data')","metadata":{"id":"VW3B8zNFzAfe","outputId":"1b7a12d5-423c-471a-be3a-7002b2839d53","execution":{"iopub.status.busy":"2022-08-10T11:34:51.521047Z","iopub.execute_input":"2022-08-10T11:34:51.521487Z","iopub.status.idle":"2022-08-10T11:34:51.742781Z","shell.execute_reply.started":"2022-08-10T11:34:51.521451Z","shell.execute_reply":"2022-08-10T11:34:51.741704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Analysis of CryoSleep**\n* Approximately 35% of passengers are in CryoSleep. \n* Passengers who are in cryo sleep appear to have an almost 80 percent chance of being transported, compared to approximately 35 percent for those who are not in cryo sleep. ","metadata":{"id":"AzCL5wfZQt40"}},{"cell_type":"markdown","source":"### Cabin\n*The cabin number where the passenger is staying. Takes the form deck/num/side, where side can be either P for Port or S for Starboard. Parsed cabin into Cabin_Deck, Cabin_Num, and Cabin_Side features.*","metadata":{"id":"hnzoWm8HLWsE"}},{"cell_type":"code","source":"plot_categorical_barh(df=train_df,\n                      col='Cabin_Deck',\n                      title='Passenger Distribution by Cabin Deck -- Train Data')","metadata":{"id":"Vs0I7FIcWYKk","outputId":"0e901cc6-8e04-4385-b958-69e6c8bc8da9","execution":{"iopub.status.busy":"2022-08-10T11:34:53.638839Z","iopub.execute_input":"2022-08-10T11:34:53.639637Z","iopub.status.idle":"2022-08-10T11:34:53.862090Z","shell.execute_reply.started":"2022-08-10T11:34:53.639585Z","shell.execute_reply":"2022-08-10T11:34:53.860974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_transported_pct(df=train_df,\n                     col='Cabin_Deck',\n                     title='Passengers Transported by Cabin Deck -- Train Data')","metadata":{"id":"C6sSa1cgVpfu","outputId":"60b56bde-a137-4da9-a1f4-4efd77600197","execution":{"iopub.status.busy":"2022-08-10T11:34:54.582056Z","iopub.execute_input":"2022-08-10T11:34:54.582470Z","iopub.status.idle":"2022-08-10T11:34:54.835231Z","shell.execute_reply.started":"2022-08-10T11:34:54.582439Z","shell.execute_reply":"2022-08-10T11:34:54.833909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_categorical_barh(df=train_df,\n                      col='Cabin_Side',\n                      title='Passenger Distribution by Cabin Side -- Train Data')","metadata":{"id":"6F5Q1PNXXcnR","outputId":"add1e7ce-b36d-458f-ea3b-bb6318219cec","execution":{"iopub.status.busy":"2022-08-10T11:34:55.503419Z","iopub.execute_input":"2022-08-10T11:34:55.504224Z","iopub.status.idle":"2022-08-10T11:34:55.836921Z","shell.execute_reply.started":"2022-08-10T11:34:55.504189Z","shell.execute_reply":"2022-08-10T11:34:55.835770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_transported_pct(df=train_df,\n                     col='Cabin_Side',\n                     title='Passengers Transported by Cabin Side -- Train Data')","metadata":{"id":"MGmGnwPJWB9E","outputId":"27ef2c24-2021-44cc-951c-f4eb4bb28310","execution":{"iopub.status.busy":"2022-08-10T11:34:56.551247Z","iopub.execute_input":"2022-08-10T11:34:56.552075Z","iopub.status.idle":"2022-08-10T11:34:56.754565Z","shell.execute_reply.started":"2022-08-10T11:34:56.552032Z","shell.execute_reply":"2022-08-10T11:34:56.753701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Analysis of Cabin and Related Features**\n* Cabin Decks G and F contain the most passengers, but those passengers on Decks B and C are the most likely to be transported.\n* Passengers are generally evenly distributed between the Port and Starboard sides with passengers on the Starboard side having a slightly higher chance of being transported.","metadata":{"id":"iWCKD5e0gwUB"}},{"cell_type":"markdown","source":"### Destination\n*The planet the passenger will be debarking to.*\n","metadata":{"id":"Erb6tw3TRmdj"}},{"cell_type":"code","source":"plot_categorical_barh(df=df,\n                      col='Destination',\n                      color=['r', 'b', 'g'],\n                      title='Distribution of Passenger Destinations -- Train Data')","metadata":{"id":"zukkcOL3WRKF","outputId":"a2675ff2-964d-4665-f60c-3729b3aa025b","execution":{"iopub.status.busy":"2022-08-10T11:34:59.290108Z","iopub.execute_input":"2022-08-10T11:34:59.291173Z","iopub.status.idle":"2022-08-10T11:34:59.492136Z","shell.execute_reply.started":"2022-08-10T11:34:59.291126Z","shell.execute_reply":"2022-08-10T11:34:59.491021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_transported_pct(df=train_df,\n                     col='Destination',\n                     title='Distribution of Transported Passengers by Destination -- Train Data')","metadata":{"id":"ccR7AyuzWg36","outputId":"9de5a52e-41a3-45fa-e42d-343d187011fe","execution":{"iopub.status.busy":"2022-08-10T11:35:08.006064Z","iopub.execute_input":"2022-08-10T11:35:08.006468Z","iopub.status.idle":"2022-08-10T11:35:08.224966Z","shell.execute_reply.started":"2022-08-10T11:35:08.006435Z","shell.execute_reply":"2022-08-10T11:35:08.223835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Analysis of Destination:**\n* TRAPPIST-1e is the most common destination and guests traveling there have roughly the same chance of being transported as the full passenger population.\n* Passengers traveling to the destination '55 Cancri e' appaer to have a slightly elevated probability of being transported.","metadata":{"id":"d8HEFwwpVw2E"}},{"cell_type":"markdown","source":"### VIP Status\n*Whether the passenger has paid for special VIP service during the voyage.*","metadata":{"id":"oi8cESEDV3Vu"}},{"cell_type":"code","source":"plot_categorical_barh(df=train_df,\n                      col='VIP',\n                      title='Distribution of Passenger VIP Status -- Train Data')","metadata":{"id":"zb2c_RV8X2rF","outputId":"340a4a39-60b6-4a26-8936-1bc490e66ae2","execution":{"iopub.status.busy":"2022-08-10T11:35:12.755780Z","iopub.execute_input":"2022-08-10T11:35:12.756219Z","iopub.status.idle":"2022-08-10T11:35:12.946140Z","shell.execute_reply.started":"2022-08-10T11:35:12.756181Z","shell.execute_reply":"2022-08-10T11:35:12.945165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_transported_pct(df=train_df,\n                     col='VIP',\n                     title='VIP Passengers who were Transported -- Train Data')","metadata":{"id":"vwh6LZUZYp1r","outputId":"18dd3b99-5b2c-43df-e0c8-371eea1c0107","execution":{"iopub.status.busy":"2022-08-10T11:35:13.191613Z","iopub.execute_input":"2022-08-10T11:35:13.192025Z","iopub.status.idle":"2022-08-10T11:35:13.395021Z","shell.execute_reply.started":"2022-08-10T11:35:13.191993Z","shell.execute_reply":"2022-08-10T11:35:13.393974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Analysis of VIP Status**\n* Although relatively few passengers paid for VIP status, those tho did appear to have a slightly lower rate of transportation than the overall passenger population.","metadata":{"id":"C80jwVByUzUx"}},{"cell_type":"markdown","source":"## Quantitative Feature Analysis","metadata":{"id":"3bS_jQ9Gc8Zb"}},{"cell_type":"markdown","source":"### Age\n*The age of the passenger.*","metadata":{"id":"6Bvm2PSSy5dM"}},{"cell_type":"code","source":"round(train_df['Age'].describe(), 2)","metadata":{"id":"2EfDdo8cRUDS","outputId":"0c732a69-2bb3-4703-c67f-e27aba54629e","execution":{"iopub.status.busy":"2022-08-10T11:35:17.685765Z","iopub.execute_input":"2022-08-10T11:35:17.686177Z","iopub.status.idle":"2022-08-10T11:35:17.699693Z","shell.execute_reply.started":"2022-08-10T11:35:17.686144Z","shell.execute_reply":"2022-08-10T11:35:17.698766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['Age'].plot(kind='hist',\n                     title='Distribution of Passenger Ages -- Train Data',\n                     figsize = (10, 6))\nplt.show()","metadata":{"id":"KHEqips9y5TG","outputId":"a18888bf-64f8-4ae6-f157-513c5caa1878","execution":{"iopub.status.busy":"2022-08-10T11:35:18.298444Z","iopub.execute_input":"2022-08-10T11:35:18.298841Z","iopub.status.idle":"2022-08-10T11:35:18.530860Z","shell.execute_reply.started":"2022-08-10T11:35:18.298804Z","shell.execute_reply":"2022-08-10T11:35:18.529644Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Analysis of Passenger Age**\n* The mean and median age are roughly 27 and 28 years old, indicating that the data is only lightly skewed.\n* The age range is 0 to 79, with the fewest passengers at the upper end of the range.","metadata":{"id":"gtZbp3qNksFQ"}},{"cell_type":"markdown","source":"### Room Service\n*Amount passenger billed for room service.*\n","metadata":{"id":"waUdpEJYgyxl"}},{"cell_type":"code","source":"round(train_df['RoomService'].describe(), 2)","metadata":{"id":"06bvtqTng92T","outputId":"3ff1c1bb-9a72-4bf3-b08a-17e1f3185f0a","execution":{"iopub.status.busy":"2022-08-10T11:35:21.574397Z","iopub.execute_input":"2022-08-10T11:35:21.575529Z","iopub.status.idle":"2022-08-10T11:35:21.588011Z","shell.execute_reply.started":"2022-08-10T11:35:21.575488Z","shell.execute_reply":"2022-08-10T11:35:21.587065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['RoomService'].plot(kind='hist',\n                             log=True,\n                             bins=20,\n                             title='Distribution of Passenger Room Service Spend -- Train Data, Log Scale',\n                             figsize=(10, 6))\nplt.show()","metadata":{"id":"0R9rfi9ug9w1","outputId":"60226d4f-0a13-443c-9c92-40d1343b57c1","execution":{"iopub.status.busy":"2022-08-10T11:35:22.127013Z","iopub.execute_input":"2022-08-10T11:35:22.127703Z","iopub.status.idle":"2022-08-10T11:35:22.826795Z","shell.execute_reply.started":"2022-08-10T11:35:22.127666Z","shell.execute_reply":"2022-08-10T11:35:22.825560Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Percentage of passengers spending any money on room service:', \n  round(train_df[train_df['RoomService'] > 0].shape[0] / train_df.shape[0], 2))","metadata":{"id":"eoeBRovqhYq6","outputId":"1efeb106-b29f-463e-f11d-8344a0905d86","execution":{"iopub.status.busy":"2022-08-10T11:35:22.874512Z","iopub.execute_input":"2022-08-10T11:35:22.875480Z","iopub.status.idle":"2022-08-10T11:35:22.884339Z","shell.execute_reply.started":"2022-08-10T11:35:22.875438Z","shell.execute_reply":"2022-08-10T11:35:22.882891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(\n    train_df[train_df['RoomService'] > 0][['RoomService', 'Transported']]\n    .groupby('Transported')\n    .agg('mean')\n    .plot(kind='barh',\n        figsize=(10,5),\n        color=['r'],\n        title='Passengers who Spent Any Money on Room Service  -- Train Data')\n)\nplt.show()","metadata":{"id":"1DqIFYNnPdRn","outputId":"42a48a6f-ac52-4dd6-9338-88293a482bef","execution":{"iopub.status.busy":"2022-08-10T11:35:23.880971Z","iopub.execute_input":"2022-08-10T11:35:23.882140Z","iopub.status.idle":"2022-08-10T11:35:24.116978Z","shell.execute_reply.started":"2022-08-10T11:35:23.882081Z","shell.execute_reply":"2022-08-10T11:35:24.115684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Analysis of Room Service Spend**\n* Only 34% of passengers spent money on Room Service.\n* Of those who did, the data is heavily skewed by a few passengers who spent large amounts.\n* Passengers who spent any money on room service appear to have a lower likelihood of being transported. ","metadata":{"id":"1_uhFgsklqfN"}},{"cell_type":"markdown","source":"### Food Court\n*Amount passenger spent at the food court*","metadata":{"id":"VQBwYTEGlm7h"}},{"cell_type":"code","source":"round(train_df['FoodCourt'].describe(), 2)","metadata":{"id":"XxvToAf3lnHP","outputId":"5926f65e-5610-44f7-e768-3f997389d0ad","execution":{"iopub.status.busy":"2022-08-10T11:35:26.845497Z","iopub.execute_input":"2022-08-10T11:35:26.845926Z","iopub.status.idle":"2022-08-10T11:35:26.858913Z","shell.execute_reply.started":"2022-08-10T11:35:26.845889Z","shell.execute_reply":"2022-08-10T11:35:26.857520Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['FoodCourt'].plot(kind='hist',\n                           log=True,\n                           title='Distribution of Passenger Food Court Spend -- Train Data, Log Scale',\n                           bins=20,\n                           figsize=(10,6))\nplt.show()","metadata":{"id":"ZYSx7asInHBt","outputId":"5fad9fa5-a7a7-4406-c4d6-4ce4af0c56c5","execution":{"iopub.status.busy":"2022-08-10T11:35:27.503739Z","iopub.execute_input":"2022-08-10T11:35:27.504448Z","iopub.status.idle":"2022-08-10T11:35:28.040370Z","shell.execute_reply.started":"2022-08-10T11:35:27.504402Z","shell.execute_reply":"2022-08-10T11:35:28.039026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Percentage of passengers spending any money at the Food Court:', \n  round(train_df[train_df['FoodCourt'] > 0].shape[0] / train_df.shape[0], 2))","metadata":{"id":"QrOwmkTDlnLW","outputId":"052593bb-31d0-4331-9326-de02f6a5cf7b","execution":{"iopub.status.busy":"2022-08-10T11:35:28.848898Z","iopub.execute_input":"2022-08-10T11:35:28.849308Z","iopub.status.idle":"2022-08-10T11:35:28.860032Z","shell.execute_reply.started":"2022-08-10T11:35:28.849276Z","shell.execute_reply":"2022-08-10T11:35:28.858618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(\n    train_df[train_df['FoodCourt'] > 0][['FoodCourt', 'Transported']]\n    .groupby('Transported')\n    .agg('mean')\n    .plot(kind='barh',\n        figsize=(10,5),\n        color=['r'],\n        title='Passengers who Spent Any Money at the Food Court  -- Train Data')\n)\nplt.show()","metadata":{"id":"XtDv3wwamUzV","outputId":"ca62808f-8b3c-4487-f1cd-dd2a3e5afb53","execution":{"iopub.status.busy":"2022-08-10T11:35:30.183911Z","iopub.execute_input":"2022-08-10T11:35:30.184410Z","iopub.status.idle":"2022-08-10T11:35:30.392141Z","shell.execute_reply.started":"2022-08-10T11:35:30.184367Z","shell.execute_reply":"2022-08-10T11:35:30.391319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Analysis of Food Court Spend**\n* Similar to other categories, roughly one-third of passengers spent money in the food court.\n* The spend was heavily skewed by a few passengers spending large amounts.\n* Customers who spent money in the food court appear to be disproportionally transported compared to the overall passenger population. ","metadata":{"id":"9ZuVlbcVnrjm"}},{"cell_type":"markdown","source":"### Shopping Mall\n*Amount passenger spent at shopping mall.*","metadata":{"id":"FImMdfn9PY8x"}},{"cell_type":"code","source":"round(train_df['ShoppingMall'].describe(), 2)","metadata":{"id":"Mre708r2PY1M","outputId":"dd802b57-8fd0-4ae2-f89a-b4c244bd47cf","execution":{"iopub.status.busy":"2022-08-10T11:35:32.236175Z","iopub.execute_input":"2022-08-10T11:35:32.236555Z","iopub.status.idle":"2022-08-10T11:35:32.249557Z","shell.execute_reply.started":"2022-08-10T11:35:32.236524Z","shell.execute_reply":"2022-08-10T11:35:32.248441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['ShoppingMall'].plot(kind='hist',\n                              log=True,\n                              bins=20,\n                              title='Distribution of Passenger Shopping Mall Spend -- Train Data, Log Scale',\n                              figsize=(10,6))\nplt.show()","metadata":{"id":"3BBEMLKAQoBQ","outputId":"304cf68e-df4f-4224-b01c-eea821b43ff0","execution":{"iopub.status.busy":"2022-08-10T11:35:32.862249Z","iopub.execute_input":"2022-08-10T11:35:32.862869Z","iopub.status.idle":"2022-08-10T11:35:33.549503Z","shell.execute_reply.started":"2022-08-10T11:35:32.862835Z","shell.execute_reply":"2022-08-10T11:35:33.548199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Percentage of passengers spending any money at the shopping mall:', \n  round(train_df[train_df['ShoppingMall'] > 0].shape[0] / train_df.shape[0], 2))","metadata":{"id":"VFSnWZrdPYwt","outputId":"42a83112-1939-4bf8-af77-39f4a1dfe37e","execution":{"iopub.status.busy":"2022-08-10T11:35:33.551586Z","iopub.execute_input":"2022-08-10T11:35:33.551938Z","iopub.status.idle":"2022-08-10T11:35:33.560245Z","shell.execute_reply.started":"2022-08-10T11:35:33.551909Z","shell.execute_reply":"2022-08-10T11:35:33.558687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(\n    train_df[train_df['ShoppingMall'] > 0][['ShoppingMall', 'Transported']]\n    .groupby('Transported')\n    .agg('median')\n    .plot(kind='barh',\n        figsize=(10,5),\n        color=['r', 'b'],\n        title='Passengers who Spent Any Money at the shopping mall  -- Train Data')\n)\nplt.show()","metadata":{"id":"Rrl5QHOORAxN","outputId":"24df0a20-e0d9-428d-a4cf-522d5eb4457d","execution":{"iopub.status.busy":"2022-08-10T11:35:34.278235Z","iopub.execute_input":"2022-08-10T11:35:34.278655Z","iopub.status.idle":"2022-08-10T11:35:34.495755Z","shell.execute_reply.started":"2022-08-10T11:35:34.278623Z","shell.execute_reply":"2022-08-10T11:35:34.494578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Analysis of Shopping Mall Spend**\n* Similar to other categories, roughly one-third of passengers spent money in the Shopping Mall.\n* This spend is heavily skewed by a small number of passengers who spent large amounts.\n* Passengers who spent money in the Shopping Mall appear to be much more likely to be transported.","metadata":{"id":"fbjkmWMJSKsg"}},{"cell_type":"markdown","source":"### Spa\n*Amount passenger spent at the spa*","metadata":{"id":"acCfMTK0RA3H"}},{"cell_type":"code","source":"round(train_df['Spa'].describe(), 2)","metadata":{"id":"HAjzMJF3SoRA","outputId":"14edce55-15bb-4f0f-bcfb-08da7e14f23c","execution":{"iopub.status.busy":"2022-08-10T11:35:36.011181Z","iopub.execute_input":"2022-08-10T11:35:36.011843Z","iopub.status.idle":"2022-08-10T11:35:36.025231Z","shell.execute_reply.started":"2022-08-10T11:35:36.011811Z","shell.execute_reply":"2022-08-10T11:35:36.023892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['Spa'].plot(kind='hist',\n                     log=True,\n                     bins=20,\n                     title='Distribution of Passenger Spa Spend -- Train Data, Log Scale',\n                     figsize=(10,6))\nplt.show()","metadata":{"id":"uTkFJTRgSoVD","outputId":"c516aebc-afc9-4a5e-e260-37bdc0a0c9cf","execution":{"iopub.status.busy":"2022-08-10T11:35:36.595447Z","iopub.execute_input":"2022-08-10T11:35:36.596166Z","iopub.status.idle":"2022-08-10T11:35:37.146397Z","shell.execute_reply.started":"2022-08-10T11:35:36.596129Z","shell.execute_reply":"2022-08-10T11:35:37.145170Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Percentage of passengers spending any money at the spa:', \n  round(train_df[train_df['Spa'] > 0].shape[0] / train_df.shape[0], 2))","metadata":{"id":"xFUREhJkTOzh","outputId":"a3e50894-3d0f-45f1-84e0-fd2f3fa1933e","execution":{"iopub.status.busy":"2022-08-10T11:35:37.786776Z","iopub.execute_input":"2022-08-10T11:35:37.787583Z","iopub.status.idle":"2022-08-10T11:35:37.796934Z","shell.execute_reply.started":"2022-08-10T11:35:37.787542Z","shell.execute_reply":"2022-08-10T11:35:37.795856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(\n    train_df[train_df['Spa'] > 0][['Spa', 'Transported']]\n    .groupby('Transported')\n    .agg('median')\n    .plot(kind='barh',\n        figsize=(10,5),\n        color=['r', 'b'],\n        title='Passengers who Spent Any Money at the Spa  -- Train Data')\n)\nplt.show()","metadata":{"id":"Ojdvhk4OTO5y","outputId":"d8113680-22f0-4b9d-c806-49b9f45deb60","execution":{"iopub.status.busy":"2022-08-10T11:35:38.287499Z","iopub.execute_input":"2022-08-10T11:35:38.287907Z","iopub.status.idle":"2022-08-10T11:35:38.505672Z","shell.execute_reply.started":"2022-08-10T11:35:38.287873Z","shell.execute_reply":"2022-08-10T11:35:38.504504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Analysis of Customer Spa Spend**\n* Slightly more than one-third of passengers spent money at the spa.\n* As with the other spend categories, Spa spend is heavily skewed by a few high-spending passengers.\n* Passengers who spent money at the Spa appear much less likely to be transported than the overall population. ","metadata":{"id":"I9oGfFVZTwfU"}},{"cell_type":"markdown","source":"### VR Deck","metadata":{"id":"QUQIqv8gTvz2"}},{"cell_type":"code","source":"round(train_df['VRDeck'].describe(), 2)","metadata":{"id":"5PlcaTGXT7Ns","outputId":"1ac1f470-3eab-45e6-91a7-51ca263f9f9e","execution":{"iopub.status.busy":"2022-08-10T11:35:40.869068Z","iopub.execute_input":"2022-08-10T11:35:40.869905Z","iopub.status.idle":"2022-08-10T11:35:40.882980Z","shell.execute_reply.started":"2022-08-10T11:35:40.869863Z","shell.execute_reply":"2022-08-10T11:35:40.881871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['VRDeck'].plot(kind='hist',\n                     log=True,\n                     bins=20,\n                     title='Distribution of Passenger VRDeck Spend -- Train Data, Log Scale',\n                     figsize=(10,6))\nplt.show()","metadata":{"id":"ehMDv6NEIb3k","outputId":"6c8c0bc9-1704-464e-817c-a0e7c5554750","execution":{"iopub.status.busy":"2022-08-10T11:35:41.440321Z","iopub.execute_input":"2022-08-10T11:35:41.441078Z","iopub.status.idle":"2022-08-10T11:35:41.976885Z","shell.execute_reply.started":"2022-08-10T11:35:41.441035Z","shell.execute_reply":"2022-08-10T11:35:41.975769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Percentage of passengers spending any money at the VR Deck:', \n  round(train_df[train_df['VRDeck'] > 0].shape[0] / train_df.shape[0], 2))","metadata":{"outputId":"c7bcccaf-62b4-4235-8009-f151698c8373","id":"yOPHz3aZUx1y","execution":{"iopub.status.busy":"2022-08-10T11:35:42.058222Z","iopub.execute_input":"2022-08-10T11:35:42.058615Z","iopub.status.idle":"2022-08-10T11:35:42.068103Z","shell.execute_reply.started":"2022-08-10T11:35:42.058583Z","shell.execute_reply":"2022-08-10T11:35:42.067268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(\n    train_df[train_df['VRDeck'] > 0][['VRDeck', 'Transported']]\n    .groupby('Transported')\n    .agg('median')\n    .plot(kind='barh',\n        figsize=(10,5),\n        color=['r', 'b'],\n        title='Passengers who Spent Any Money at the VR Deck  -- Train Data')\n)\nplt.show()","metadata":{"outputId":"d37631ee-2ba0-42ec-fc7a-bf51f845f5bc","id":"8k9qkZtdUx11","execution":{"iopub.status.busy":"2022-08-10T11:35:42.756432Z","iopub.execute_input":"2022-08-10T11:35:42.757538Z","iopub.status.idle":"2022-08-10T11:35:42.978473Z","shell.execute_reply.started":"2022-08-10T11:35:42.757490Z","shell.execute_reply":"2022-08-10T11:35:42.977577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Analysis of Customer VR Deck Spend**\n* Slightly more than one-third of passengers spent money at the VR Deck.\n* As with the other spend categories, VR Deck spend is heavily skewed by a few high-spending passengers.\n* Passengers who spent money at the VR Deck appear much less likely to be transported than the overall population. ","metadata":{"id":"ghEcEpWuUx13"}},{"cell_type":"markdown","source":"## Finalize Data for ML Models","metadata":{"id":"yOM1hExb-coP"}},{"cell_type":"markdown","source":"### One-Hot Encoding","metadata":{"id":"2j9IBQaoS777"}},{"cell_type":"code","source":"# create y as the target value \ny = train_df['Transported']\ny.shape","metadata":{"id":"8qdj6fDw-cg9","outputId":"38003c00-11be-4f10-80ba-b821d1a7d156","execution":{"iopub.status.busy":"2022-08-10T11:35:44.763055Z","iopub.execute_input":"2022-08-10T11:35:44.763807Z","iopub.status.idle":"2022-08-10T11:35:44.771311Z","shell.execute_reply.started":"2022-08-10T11:35:44.763766Z","shell.execute_reply":"2022-08-10T11:35:44.770153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# features to use for the model \ntrain_features = ['HomePlanet', 'CryoSleep', 'Destination', 'VIP',\n                  'Group_Size', 'Cabin_Deck', 'Cabin_Side',\n                  'Age', 'RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']\n\nlen(train_features)","metadata":{"id":"zSPI-FZJ-cdz","outputId":"dca59069-4b3d-49c5-8ec0-08c7c522db96","execution":{"iopub.status.busy":"2022-08-10T11:35:45.763091Z","iopub.execute_input":"2022-08-10T11:35:45.764076Z","iopub.status.idle":"2022-08-10T11:35:45.771080Z","shell.execute_reply.started":"2022-08-10T11:35:45.764038Z","shell.execute_reply":"2022-08-10T11:35:45.769933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# generate dummy variables from categoricals \nX = pd.get_dummies(train_df[train_features], drop_first=True)\nX.shape\n","metadata":{"id":"Chef8BZG-cZ-","outputId":"c1993cfc-b098-4a69-fc1b-e9642a0d984e","execution":{"iopub.status.busy":"2022-08-10T11:35:46.458309Z","iopub.execute_input":"2022-08-10T11:35:46.459130Z","iopub.status.idle":"2022-08-10T11:35:46.476839Z","shell.execute_reply.started":"2022-08-10T11:35:46.459074Z","shell.execute_reply":"2022-08-10T11:35:46.475369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X.head()","metadata":{"id":"VZO2asfq-cVd","outputId":"424e2994-5d43-4430-93d8-8021d539d04f","execution":{"iopub.status.busy":"2022-08-10T11:35:47.862754Z","iopub.execute_input":"2022-08-10T11:35:47.863179Z","iopub.status.idle":"2022-08-10T11:35:47.890213Z","shell.execute_reply.started":"2022-08-10T11:35:47.863132Z","shell.execute_reply":"2022-08-10T11:35:47.889040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test = pd.get_dummies(test_df[train_features], drop_first=True)\nX_test.shape","metadata":{"id":"1fscCPShAgG4","outputId":"f54861fa-7414-4f37-f6d4-28f19627d90a","execution":{"iopub.status.busy":"2022-08-10T11:35:48.279106Z","iopub.execute_input":"2022-08-10T11:35:48.280069Z","iopub.status.idle":"2022-08-10T11:35:48.298872Z","shell.execute_reply.started":"2022-08-10T11:35:48.280028Z","shell.execute_reply":"2022-08-10T11:35:48.297607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# confirm same columns; identify any discrepancies\n[col for col in list(X.columns) if col not in list(X_test.columns)]\n","metadata":{"id":"r3Ul9WumAgcL","outputId":"541060ca-23fd-495d-94fd-f5c90fe0121f","execution":{"iopub.status.busy":"2022-08-10T11:35:49.212183Z","iopub.execute_input":"2022-08-10T11:35:49.213294Z","iopub.status.idle":"2022-08-10T11:35:49.223919Z","shell.execute_reply.started":"2022-08-10T11:35:49.213234Z","shell.execute_reply":"2022-08-10T11:35:49.222540Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Check Feature Correlation","metadata":{"id":"o5nnqOMBAggv"}},{"cell_type":"code","source":"# correlation map of the fatures \ncolormap = plt.cm.RdBu\nplt.figure(figsize=(18,14))\nplt.title('Pearson Correlation of Features', y=1.05, size=15)\nsns.heatmap(X.astype(float).corr(),\n            linewidths=0.1,\n            vmax=1.0, \n            square=True, \n            cmap=colormap, \n            linecolor='white', \n            annot=True)\n\nplt.show()","metadata":{"id":"QeQCUpdzD4Hl","outputId":"68942511-e377-4ed1-b63d-24fee40ce8a7","execution":{"iopub.status.busy":"2022-08-10T11:35:53.789792Z","iopub.execute_input":"2022-08-10T11:35:53.790209Z","iopub.status.idle":"2022-08-10T11:35:56.486964Z","shell.execute_reply.started":"2022-08-10T11:35:53.790176Z","shell.execute_reply":"2022-08-10T11:35:56.485769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train Random Forest Model","metadata":{"id":"B5Lxz8eVf8b-"}},{"cell_type":"markdown","source":"### Identify Optimal RF Parameters","metadata":{"id":"GpMwbnMaGnKb"}},{"cell_type":"markdown","source":"**Establish RandomizedSearchCV**","metadata":{"id":"8Bf2_ZHMFrMX"}},{"cell_type":"code","source":"# rf_model_cv = RandomForestClassifier(oob_score=True, random_state=1, n_jobs=-1)\n\n# param_grid_rf = {'criterion' : [\"gini\", \"entropy\"], \n#                  'max_features': ['log2', 'sqrt', None],\n#                  'max_depth': [2, 4, 8, 16, 32, 64],\n#                  'min_samples_leaf' : [1, 5, 10, 20], \n#                  'min_samples_split' : [2, 4, 10, 14, 18], \n#                  'n_estimators': [100, 200, 300, 400, 500]}\n\n# gs_rf = RandomizedSearchCV(estimator=rf_model_cv, \n#                            param_distributions=param_grid_rf, \n#                            scoring='accuracy', \n#                            n_iter = 100,\n#                            cv=5, \n#                            n_jobs=-1)\n\n# gs_rf.fit(X.values, y)","metadata":{"id":"o6SIP8Ka7Fnh","outputId":"f3e206f0-47bf-48ff-f839-7c12f77bc2b3","execution":{"iopub.status.busy":"2022-08-10T11:36:08.095498Z","iopub.execute_input":"2022-08-10T11:36:08.095924Z","iopub.status.idle":"2022-08-10T11:46:06.904205Z","shell.execute_reply.started":"2022-08-10T11:36:08.095885Z","shell.execute_reply":"2022-08-10T11:46:06.901953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Identify Best Estimator and Features**","metadata":{"id":"Zmq4I7dEFwGA"}},{"cell_type":"code","source":"# print(gs_rf.best_estimator_)","metadata":{"id":"cYzaEiNO8DpL","outputId":"aeb9a2a7-176e-4bb1-fa15-17eed6095a17","execution":{"iopub.status.busy":"2022-08-10T11:46:06.906727Z","iopub.execute_input":"2022-08-10T11:46:06.907129Z","iopub.status.idle":"2022-08-10T11:46:06.914257Z","shell.execute_reply.started":"2022-08-10T11:46:06.907079Z","shell.execute_reply":"2022-08-10T11:46:06.912914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# print(gs_rf.best_params_)","metadata":{"id":"8ZeCVIdbFmIe","outputId":"e5adf644-f361-49d3-cc9b-aaff0c898336","execution":{"iopub.status.busy":"2022-08-10T11:46:06.916048Z","iopub.execute_input":"2022-08-10T11:46:06.916526Z","iopub.status.idle":"2022-08-10T11:46:06.926722Z","shell.execute_reply.started":"2022-08-10T11:46:06.916484Z","shell.execute_reply":"2022-08-10T11:46:06.925509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# print(round(gs_rf.best_score_, 4))","metadata":{"id":"dO8JPrwXFotS","outputId":"0f243016-071b-4677-e6f5-83ca6cdfa9d3","execution":{"iopub.status.busy":"2022-08-10T11:46:06.929336Z","iopub.execute_input":"2022-08-10T11:46:06.929781Z","iopub.status.idle":"2022-08-10T11:46:06.938902Z","shell.execute_reply.started":"2022-08-10T11:46:06.929739Z","shell.execute_reply":"2022-08-10T11:46:06.937589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Model Parameters**\nRunning the above code produced the following parameters with a best score of 0.8003\n* 'n_estimators': 300, \n* 'min_samples_split': 18, \n* 'min_samples_leaf': 5, \n* 'max_features': 'sqrt', \n* 'max_depth': 16, \n* 'criterion': 'entropy'","metadata":{}},{"cell_type":"markdown","source":"### Train Model","metadata":{"id":"2OX2onp0F2nC"}},{"cell_type":"code","source":"# specify the model with optimal parameters\nrf_model = RandomForestClassifier(criterion='entropy', \n                                  n_estimators=300,\n                                  min_samples_split=18,\n                                  min_samples_leaf=5,\n                                  max_features='sqrt',\n                                  oob_score=True,\n                                  max_depth=16,\n                                  random_state=1,\n                                  n_jobs=-1)","metadata":{"id":"-BycnA7iGtPG","execution":{"iopub.status.busy":"2022-08-10T11:51:34.239657Z","iopub.execute_input":"2022-08-10T11:51:34.240037Z","iopub.status.idle":"2022-08-10T11:51:34.245906Z","shell.execute_reply.started":"2022-08-10T11:51:34.240008Z","shell.execute_reply":"2022-08-10T11:51:34.245003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# fit to training data \nrf_model.fit(X.values, y)\n\n# predict on test data \nrf_predictions = rf_model.predict(X_test.values)\n","metadata":{"id":"qpy2xFIxHdsS","execution":{"iopub.status.busy":"2022-08-10T11:51:35.237829Z","iopub.execute_input":"2022-08-10T11:51:35.238709Z","iopub.status.idle":"2022-08-10T11:51:37.336475Z","shell.execute_reply.started":"2022-08-10T11:51:35.238667Z","shell.execute_reply":"2022-08-10T11:51:37.335197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Feature Importance","metadata":{"id":"x0pxDiG0Hr-c"}},{"cell_type":"code","source":"plt.figure(figsize = (12,10))\n\nplt.barh(X.columns, rf_model.feature_importances_)\nplt.title('Random Forest Model Feature Importance')\nplt.show()","metadata":{"id":"B3xmaln6HzpY","outputId":"e205baec-b030-4138-dcea-5fb07e992703","execution":{"iopub.status.busy":"2022-08-10T11:58:06.277106Z","iopub.execute_input":"2022-08-10T11:58:06.277721Z","iopub.status.idle":"2022-08-10T11:58:06.749321Z","shell.execute_reply.started":"2022-08-10T11:58:06.277676Z","shell.execute_reply":"2022-08-10T11:58:06.747883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Analysis of Model and Feature Importance\n\nPlotting RF feature importance confirms much of the intuition we gained during the feature analysis, with additional wrinkles: \n* Whether the passenger was in CryoSleep or not stands out as an extremely important feature for this model. While only one-third of passengers were in CryoSleep, 80% of those passengers were transported.\n* Customer spend at various locations (Room Service, Spa, VR Deck, Food Court, and to a lesser extent the Shopping Mall) were also strong features, particularly the Spa. Our EDA showed major swings among customers who spent at these locations so it is not surprising they were important. \n* Other features contributed, but to lesser extent than those mentioned above. Age, Home Planet, and Group Size are notable additional features. A few Passenger Decks (F, G) and the Starboard side stick out as well. ","metadata":{"id":"pTEyqpm9TiW_"}},{"cell_type":"markdown","source":"## Final Output for Submission","metadata":{"id":"BK8g1LQbH4F4"}},{"cell_type":"code","source":"output = pd.DataFrame({'PassengerId': test_df.PassengerId, 'Transported': rf_predictions})\noutput.head()","metadata":{"id":"F-PsqoiVIetI","outputId":"fc985a05-6b9e-445f-ae42-2e95566cbd1b","execution":{"iopub.status.busy":"2022-08-10T11:58:18.093302Z","iopub.execute_input":"2022-08-10T11:58:18.093710Z","iopub.status.idle":"2022-08-10T11:58:18.112669Z","shell.execute_reply.started":"2022-08-10T11:58:18.093680Z","shell.execute_reply":"2022-08-10T11:58:18.111654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"output.to_csv('submission.csv', index=False)\nprint(\"Your submission was successfully saved!\")","metadata":{"id":"6xYK_jfCI1gk","execution":{"iopub.status.busy":"2022-08-10T11:58:19.382006Z","iopub.execute_input":"2022-08-10T11:58:19.383016Z","iopub.status.idle":"2022-08-10T11:58:19.399575Z","shell.execute_reply.started":"2022-08-10T11:58:19.382976Z","shell.execute_reply":"2022-08-10T11:58:19.398593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}