{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-08-07T13:54:28.598818Z","iopub.execute_input":"2022-08-07T13:54:28.599283Z","iopub.status.idle":"2022-08-07T13:54:28.606791Z","shell.execute_reply.started":"2022-08-07T13:54:28.599243Z","shell.execute_reply":"2022-08-07T13:54:28.605867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Spaceship Titanic Project\nEveryone who dips their toes in the waters of data science knows about the Titanic dataset. If only our community of analytics experts had been around before that fateful voyage! In this project, we have the opportunity to look at data from the *Spaceship Titanic*, which will collide with a spacetime anomaly in the year 2912! Almost half of the nearly 13,000 passengers will be transported to an alternate dimension! From the damaged spaceship computer system, our mission is to predict which passengers were transported by the anomaly. ","metadata":{}},{"cell_type":"markdown","source":"## Data Exploration\nLet's dig in to the dataset to get a sense of our data structure and look for potential predictors. ","metadata":{}},{"cell_type":"code","source":"# import data into pandas df\ntrain = pd.read_csv('../input/spaceship-titanic/train.csv')\n\n# examine first 5 rows\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:28.662877Z","iopub.execute_input":"2022-08-07T13:54:28.663792Z","iopub.status.idle":"2022-08-07T13:54:28.700973Z","shell.execute_reply.started":"2022-08-07T13:54:28.663759Z","shell.execute_reply":"2022-08-07T13:54:28.700200Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# get info on variables\ntrain.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:28.723495Z","iopub.execute_input":"2022-08-07T13:54:28.724393Z","iopub.status.idle":"2022-08-07T13:54:28.741422Z","shell.execute_reply.started":"2022-08-07T13:54:28.724310Z","shell.execute_reply":"2022-08-07T13:54:28.740245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In our training set, we have 8693 passengers. For each passenger, we are provided 13 features with information about the background information as well as on-cruise spending habits of each passenger. Our target column is 'Transported', which indicates whether the passenger was transported to another dimension or not. ","metadata":{}},{"cell_type":"markdown","source":"### Checking for nulls","metadata":{}},{"cell_type":"code","source":"# Compute nulls in each column\ntrain.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:28.781203Z","iopub.execute_input":"2022-08-07T13:54:28.781793Z","iopub.status.idle":"2022-08-07T13:54:28.795174Z","shell.execute_reply.started":"2022-08-07T13:54:28.781763Z","shell.execute_reply":"2022-08-07T13:54:28.794163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have around 200 nulls in a number of columns. If we drop every row with a null, we would lose about 2000 data points. Let's remove nulls once we have a set list of potential features.","metadata":{}},{"cell_type":"markdown","source":"### Numeric Features\n\nLet's start by looking at the numeric variables for any correlations between them and the target. ","metadata":{}},{"cell_type":"code","source":"# extract numeric vars\nnumerics = ['int16', 'int32', 'int64', 'float16', 'float32', 'float64']\ntrain_numerics = train.select_dtypes(include=numerics)\n\n# map boolean target to integer value and append result to numeric table\ntrain_numerics['Transported'] = train['Transported'].astype('int')\n\n# compute pearson correlations\ntrain_corrs = train_numerics.corr()\n\n# import plotting libraries and plot heatmap\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\nsns.heatmap(train_corrs)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:28.838943Z","iopub.execute_input":"2022-08-07T13:54:28.839534Z","iopub.status.idle":"2022-08-07T13:54:29.074945Z","shell.execute_reply.started":"2022-08-07T13:54:28.839498Z","shell.execute_reply":"2022-08-07T13:54:29.074040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_corrs","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:29.076593Z","iopub.execute_input":"2022-08-07T13:54:29.076897Z","iopub.status.idle":"2022-08-07T13:54:29.090879Z","shell.execute_reply.started":"2022-08-07T13:54:29.076870Z","shell.execute_reply":"2022-08-07T13:54:29.089649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We don't see any particularly strong correlations with the 'Transported' variable. The three most promising of the batch are VRDeck, Spa, and RoomService, which each correlate negatively with transportation (and thus positively with survival!). Perhaps the guests who could afford to spend more on these amenities were more likely to be safe when disaster struck. \nFrom our correlation values, it does not seem that any two numeric variables are too highly correlated. Consulting the documentation, none of these features should give us cause for concern of redundant information.\n\n### Categorical Features\nLet's explore our categorical variables.\n\n#### CryoSleep\nOne interesting feature that should jump out is 'CryoSleep'. Were passengers who were asleep more or less likely to survive? Let's quickly compare survival rates for these groups.","metadata":{}},{"cell_type":"code","source":"# group data by 'CryoSleep' and compute survival rates\ncryo_survival = train.groupby('CryoSleep')['Transported'].mean()\n\n# convert transport rate to survival rate\ncryo_survival = 1 - cryo_survival\n\n# bar plot\ncryo_survival.plot.bar(rot=0)\nplt.ylabel('Survival Rate')\nplt.title('Proportion of Survivors by Sleep Status')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:29.092484Z","iopub.execute_input":"2022-08-07T13:54:29.092861Z","iopub.status.idle":"2022-08-07T13:54:29.230661Z","shell.execute_reply.started":"2022-08-07T13:54:29.092823Z","shell.execute_reply":"2022-08-07T13:54:29.229897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cryo_survival","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:29.232501Z","iopub.execute_input":"2022-08-07T13:54:29.233261Z","iopub.status.idle":"2022-08-07T13:54:29.240815Z","shell.execute_reply.started":"2022-08-07T13:54:29.233229Z","shell.execute_reply":"2022-08-07T13:54:29.239987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"CryoSleep seems to be a promising predictor! 67% of those not in cryo sleep survived, while only 18% of those asleep did!\n\nCryoSleep should also set a minimum standard for our model performance. If the training set is representative of the test set, then guessing purely based on this variable should give us 67% accuracy. If any of the models we implement do not achieve at least this standard of accuracy then they are not optimal for our predictions and we can do a lot better. \n\n### Cabin\n\"Cabin\" is also a potential source of interesting information. Maybe the spacetime vortex struck specific locations on the ship! This variable takes the form deck/num/side, which we can separate into its component parts. Num is probably not as relevant as the deck or side, so let's extract those two variables and plot survival rates across them.","metadata":{}},{"cell_type":"code","source":"# Separate cabin column by '/'\ntrain[['cabin_deck', 'cabin_num', 'cabin_side']] = train['Cabin'].str.split('/', expand = True)\n\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:29.241868Z","iopub.execute_input":"2022-08-07T13:54:29.242532Z","iopub.status.idle":"2022-08-07T13:54:29.282515Z","shell.execute_reply.started":"2022-08-07T13:54:29.242498Z","shell.execute_reply":"2022-08-07T13:54:29.281602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# group data by cabin_deck and compute survival rates\ndeck_survival = 1 - train.groupby('cabin_deck')['Transported'].mean()\n\ndeck_survival","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:29.283604Z","iopub.execute_input":"2022-08-07T13:54:29.283882Z","iopub.status.idle":"2022-08-07T13:54:29.292698Z","shell.execute_reply.started":"2022-08-07T13:54:29.283856Z","shell.execute_reply":"2022-08-07T13:54:29.291768Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plot survival by deck\ndeck_survival.plot.bar(rot=0)\nplt.ylabel('Survival Rate')\nplt.title('Survival by Deck')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:29.294142Z","iopub.execute_input":"2022-08-07T13:54:29.294624Z","iopub.status.idle":"2022-08-07T13:54:29.457579Z","shell.execute_reply.started":"2022-08-07T13:54:29.294593Z","shell.execute_reply":"2022-08-07T13:54:29.456309Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"80% of deck T survived, while only 26% of those in cabin B survived. With a range of survival values, it seems that cabin deck is another promising feature. Let's check the distribution across cabin side, which is a binary variable. ","metadata":{}},{"cell_type":"code","source":"# group data by cabin_side and compute survival rates\nside_survival = 1 - train.groupby('cabin_side')['Transported'].mean()\n\nside_survival","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:29.460003Z","iopub.execute_input":"2022-08-07T13:54:29.460376Z","iopub.status.idle":"2022-08-07T13:54:29.469889Z","shell.execute_reply.started":"2022-08-07T13:54:29.460331Z","shell.execute_reply":"2022-08-07T13:54:29.469138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plot survival by cabin side\nside_survival.plot.bar(rot=0)\nplt.ylabel('Survival Rate')\nplt.xlabel('Cabin Side')\nplt.title('Survival by Cabin Side')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:29.471057Z","iopub.execute_input":"2022-08-07T13:54:29.472360Z","iopub.status.idle":"2022-08-07T13:54:29.579110Z","shell.execute_reply.started":"2022-08-07T13:54:29.472323Z","shell.execute_reply":"2022-08-07T13:54:29.578001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Those staying in the port side were slightly more likely to survive, but not by a large margin. Let's keep cabin side in our back pocket for now.","metadata":{}},{"cell_type":"markdown","source":"### Passenger ID\nPassenger ID contains which group a specific passenger belonged to. Let's get a sense of whether certain groups were more likely to survive than others. ","metadata":{}},{"cell_type":"code","source":"# separate passenger id by group and individual\ntrain[['group', 'number']] =  train['PassengerId'].str.split('_', expand=True)\n\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:29.580567Z","iopub.execute_input":"2022-08-07T13:54:29.581152Z","iopub.status.idle":"2022-08-07T13:54:29.618767Z","shell.execute_reply.started":"2022-08-07T13:54:29.581109Z","shell.execute_reply":"2022-08-07T13:54:29.617740Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# group by passenger group and compute survival\ngroup_survival = 1 - train.groupby('group')['Transported'].mean()\n\ngroup_survival","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:29.620284Z","iopub.execute_input":"2022-08-07T13:54:29.620908Z","iopub.status.idle":"2022-08-07T13:54:29.634515Z","shell.execute_reply.started":"2022-08-07T13:54:29.620863Z","shell.execute_reply":"2022-08-07T13:54:29.633380Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Unfortunately there are 6217 unique groups in this set! Since this is quite a large number of categories, and we don't know how well this information is representative of our test set, let's keep group out of our analysis for now.","metadata":{}},{"cell_type":"markdown","source":"### HomePlanet and Destination\nThere are a couple different ways the journey path could influence survival. Those from the same homeplanet or going to the same destination may be more likely to engage in certain behaviors or congregate in specific areas of the ship that may have saved them at the time of the accident. ","metadata":{}},{"cell_type":"code","source":"# group by homeplanet\nhomeplanet_survival = 1 - train.groupby('HomePlanet')['Transported'].mean()\n\nhomeplanet_survival","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:29.635618Z","iopub.execute_input":"2022-08-07T13:54:29.636360Z","iopub.status.idle":"2022-08-07T13:54:29.645514Z","shell.execute_reply.started":"2022-08-07T13:54:29.636320Z","shell.execute_reply":"2022-08-07T13:54:29.644601Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# group by destination\n# group by homeplanet\ndestination_survival = 1 - train.groupby('Destination')['Transported'].mean()\n\ndestination_survival","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:29.647699Z","iopub.execute_input":"2022-08-07T13:54:29.648756Z","iopub.status.idle":"2022-08-07T13:54:29.662643Z","shell.execute_reply.started":"2022-08-07T13:54:29.648729Z","shell.execute_reply":"2022-08-07T13:54:29.661546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Those from Europa and those transiting to 55 Cancri e both have around a relatively low change of survival, but for the other two respective destinations and home planets, the effect does not seem so strong. \n","metadata":{}},{"cell_type":"markdown","source":"## Summary of Features and Selection\nWe can rank the numerical features by absolute value of correlation with 'Transported'. From strongest to weakest:\n* RoomService\n* Spa\n* VRDeck\n* Age, ShoppingMall, FoodCourt (very weak)\n\nWe can rank the categorical features by the discrepancies we observed earlier. From stronges to weakest:\n* CryoSleep\n* cabin_deck\n* HomePlanet, Destination\n* cabin_side\n\nFor our model, let's start with the top numerical and categorical features. Then, we can selectively add features to observe their influence on model performance. Recall that we have not yet dealt with nulls in this set. Let's make a new dataframe with all our potential features. At the most, we are only going to use the top two numerical features and top two categorical features.","metadata":{}},{"cell_type":"code","source":"# list all potential features\nall_features = ['CryoSleep', 'RoomService', 'cabin_deck', 'Spa']\n\n# make new dataframe\ntrain_trim = train[all_features]\ntrain_trim['Transported'] = train['Transported']\n\n# remove all nulls\ntrain_trim = train_trim.dropna(axis=0)\n\n# reset index\ntrain_trim = train_trim.reset_index(drop=True)\n\ntrain_trim.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:29.664626Z","iopub.execute_input":"2022-08-07T13:54:29.665399Z","iopub.status.idle":"2022-08-07T13:54:29.688762Z","shell.execute_reply.started":"2022-08-07T13:54:29.665364Z","shell.execute_reply":"2022-08-07T13:54:29.687415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## K Nearest Neighbors Model\nWe will start with a KNN classifier using 'CryoSleep' to predict model performance, then selectively add features and observe the effect on performance. We will start with 2 nearest neighbors and optimize. We will use 4 fold cross validation and compute the mean accuracy of the folds to evaluate performance.","metadata":{}},{"cell_type":"markdown","source":"### 1 Feature: CryoSleep","metadata":{}},{"cell_type":"code","source":"# import KNN and Kfold\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.model_selection import KFold\n\n# train knn model on features and labels from training set with n neighbors\ndef train_knn(k, train_features, train_labels):\n    knn = KNeighborsClassifier(n_neighbors = k)\n    knn.fit(train_features, train_labels)\n    return knn\n\n# test knn\ndef test_knn(knn, test_features, test_labels):\n    tester_df = pd.DataFrame()\n    predictions = knn.predict(test_features)\n    tester_df['true'] = test_labels\n    tester_df['pred'] = predictions\n    return (tester_df['pred']==tester_df['true']).mean()\n\n# perform k-fold cross validation with 4 folds using train and test functions\ndef cross_validate_knn(data, labels, k):\n    kf = KFold(n_splits = 4, shuffle = True)\n    fold_accs = []\n    for train_index, test_index in kf.split(data):\n        train_features, test_features = data.loc[train_index], data.loc[test_index]\n        train_labels, test_labels = labels.loc[train_index], labels.loc[test_index]\n        knn = train_knn(k, train_features, train_labels)\n        fold_accs.append(test_knn(knn, test_features, test_labels))\n        \n    return fold_accs\n","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:29.691124Z","iopub.execute_input":"2022-08-07T13:54:29.691749Z","iopub.status.idle":"2022-08-07T13:54:29.701622Z","shell.execute_reply.started":"2022-08-07T13:54:29.691709Z","shell.execute_reply":"2022-08-07T13:54:29.700483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# convert boolean CryoSleep to int\ntrain_trim['CryoSleep'] = train_trim['CryoSleep'].astype(int)\n\n# feed data and labels to model\ndata = pd.DataFrame(train_trim['CryoSleep'])\nlabels = train_trim['Transported']\nknn_one_accuracies = cross_validate_knn(data, labels, 1)\nnp.mean(knn_one_accuracies)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:29.708132Z","iopub.execute_input":"2022-08-07T13:54:29.708544Z","iopub.status.idle":"2022-08-07T13:54:30.032972Z","shell.execute_reply.started":"2022-08-07T13:54:29.708514Z","shell.execute_reply":"2022-08-07T13:54:30.031782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Univariate Model: Finding most predictive feature and optimizing hyperparams\nLet's feed each of our potential features to the model one by one, then plot the corresponding accuracies of a univariate KNN model. We will also vary the number of nearest neighbors for each model.","metadata":{}},{"cell_type":"code","source":"# extract labels\nlabels = train_trim['Transported']\n\n# extract cryosleep feature\ncryo_sleep = pd.DataFrame(train_trim['CryoSleep'])\n\n# extract room service feature\nroom_service = pd.DataFrame(train_trim['RoomService'])\n\n# extract and dummy code cabin_deck feature\ncabin_deck = pd.get_dummies(train_trim['cabin_deck'])\n\n# extract spa feature\nspa = pd.DataFrame(train_trim['Spa'])\n\n# store feature vars in list to iterate\nfeature_list = [cryo_sleep, room_service, cabin_deck, spa]","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:30.034940Z","iopub.execute_input":"2022-08-07T13:54:30.035371Z","iopub.status.idle":"2022-08-07T13:54:30.045285Z","shell.execute_reply.started":"2022-08-07T13:54:30.035334Z","shell.execute_reply":"2022-08-07T13:54:30.044128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# set list of k-vals\nk_list = range(1, 11)\n\n# initialize dictionary to store mean performance for each feature and k val\nperformance_vals = {}\n\n# iterate features\nfor i in range(len(feature_list)):\n    # extract data\n    data = feature_list[i]\n    \n    # initialize list of performance vals \n    feature_performance_vals = []\n    \n    # iterate k vals\n    for k in k_list:\n        # run univariate model for feature and k val. store mean performance across folds\n        feature_performance_vals.append(np.mean(cross_validate_knn(data, labels, k)))\n\n    # add list of performance vals to dictionary\n    performance_vals[all_features[i]] = feature_performance_vals\n    \n# convert dictionary to dataframe\nperformance_vals = pd.DataFrame(performance_vals, index=k_list)\nperformance_vals.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:30.046944Z","iopub.execute_input":"2022-08-07T13:54:30.047354Z","iopub.status.idle":"2022-08-07T13:54:44.256453Z","shell.execute_reply.started":"2022-08-07T13:54:30.047314Z","shell.execute_reply":"2022-08-07T13:54:44.255400Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"These performance metrics are disappointing given that we could predict survival with 67% accuracy using CryoSleep alone (with no model). Let's turn to some other potential models to answer our question. ","metadata":{}},{"cell_type":"markdown","source":"## Neural Network Model","metadata":{}},{"cell_type":"code","source":"from sklearn.neural_network import MLPClassifier\nfrom sklearn.model_selection import KFold\n\n# trains neural network with one hidden layer of n_a architecture\ndef train_nn(n_a, train_features, train_labels):\n    nn = MLPClassifier(hidden_layer_sizes = n_a)\n    nn.fit(train_features, train_labels)\n    return nn\n\n# tests neural network and returns overall accuracy\ndef test_nn(nn, test_features, test_labels):\n    predictions = nn.predict(test_features)\n    tester_df = pd.DataFrame()\n    tester_df['true'] = test_labels\n    tester_df['pred'] = predictions\n    return np.mean(tester_df['true'] == tester_df['pred'])\n\n# implements 4-fold cross validation, takes in architecture for neural network\ndef cross_validate_nn(data, labels, n_a):\n    kf = KFold(n_splits = 4, shuffle = True)\n    fold_accs = []\n    for train_index, test_index in kf.split(data):\n        train_features, test_features = data.loc[train_index], data.loc[test_index]\n        train_labels, test_labels = labels.loc[train_index], labels.loc[test_index]\n        nn = train_nn(n_a, train_features, train_labels)\n        fold_accs.append(test_nn(nn, test_features, test_labels))\n    return fold_accs","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:44.258798Z","iopub.execute_input":"2022-08-07T13:54:44.259118Z","iopub.status.idle":"2022-08-07T13:54:44.266865Z","shell.execute_reply.started":"2022-08-07T13:54:44.259072Z","shell.execute_reply":"2022-08-07T13:54:44.265873Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# set list of neuron layers\nn_list = range(1, 6)\n\n# initialize dictionary to store mean performance for each feature and k val\nperformance_vals = {}\n\n# iterate features\nfor i in range(len(feature_list)):\n    # extract data\n    data = feature_list[i]\n    \n    # initialize list of performance vals \n    feature_performance_vals = []\n    \n    # iterate k vals\n    for n in n_list:\n        # run univariate model for feature and k val. store mean performance across folds\n        feature_performance_vals.append(np.mean(cross_validate_nn(data, labels, n)))\n\n    # add list of performance vals to dictionary\n    performance_vals[all_features[i]] = feature_performance_vals\n    \n# convert dictionary to dataframe\nperformance_vals = pd.DataFrame(performance_vals, index=n_list)\nperformance_vals.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:54:44.268069Z","iopub.execute_input":"2022-08-07T13:54:44.268531Z","iopub.status.idle":"2022-08-07T13:55:29.079384Z","shell.execute_reply.started":"2022-08-07T13:54:44.268505Z","shell.execute_reply":"2022-08-07T13:55:29.078315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = pd.concat(feature_list, axis=1)\n\nnp.mean(cross_validate_nn(data,labels,3))","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:55:29.082744Z","iopub.execute_input":"2022-08-07T13:55:29.085143Z","iopub.status.idle":"2022-08-07T13:55:32.797993Z","shell.execute_reply.started":"2022-08-07T13:55:29.085107Z","shell.execute_reply":"2022-08-07T13:55:32.797127Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The neural network performs better, but just barely above our standard. ","metadata":{}},{"cell_type":"markdown","source":"## Gradient Boosting Classifier w/ CatBoost\nLet's take an alternative approach both to model building and feature selection. With gradient boosting, we don't have to worry so much about careful selection and engineering of our features. We can expand our dimensionality back to that of the original set (and add some dimensionality by vectorizing our categorical features). We are going to implement a CatBoost gradient boosting ensemble model for classification, and use the library's feature importance metrics to evaluate our feature selection step. Then, we can engage in an iterative process of model building.\n\n### Data Preparation\n#### Dealing with nulls\nSince we are going to be using all features for this implementation, we need to devise a more elegant strategy for dealing with the null values we discovered earlier. Let's use some simple imputation. For numerics, we will fill nulls with the mean. For categoricals, we will fill with the mode. However, this is a bit problematic for missing passenger names since they are unique to each person. It is unlikely that first names will give us much insight, but last names might hold some weight (maybe we should have done some EDA on this).\n","metadata":{}},{"cell_type":"code","source":"# for numeric values, replace nulls with mean\nnumerics = train.select_dtypes(include=np.number).columns.to_list()\ntrain[numerics] = train[numerics].fillna(train[numerics].mean())\n\n# for categoricals, replace nulls with mode. first, let's use name to generate lastname var and drop firstname\ntrain[['first_name', 'last_name']] = train['Name'].str.split(' ', expand=True)\ntrain = train.drop(columns=['Name', 'first_name'])\ncategoricals = train.select_dtypes(exclude=np.number).columns.to_list()\ntrain[categoricals] = train[categoricals].fillna(train[categoricals].mode().iloc[0])","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:55:32.799186Z","iopub.execute_input":"2022-08-07T13:55:32.799467Z","iopub.status.idle":"2022-08-07T13:55:32.875691Z","shell.execute_reply.started":"2022-08-07T13:55:32.799442Z","shell.execute_reply":"2022-08-07T13:55:32.874780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# drop target and name from categoricals list for catboost\ncategoricals.remove('Transported')\n\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:55:32.876894Z","iopub.execute_input":"2022-08-07T13:55:32.877179Z","iopub.status.idle":"2022-08-07T13:55:32.896743Z","shell.execute_reply.started":"2022-08-07T13:55:32.877154Z","shell.execute_reply":"2022-08-07T13:55:32.896053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:55:32.897950Z","iopub.execute_input":"2022-08-07T13:55:32.898229Z","iopub.status.idle":"2022-08-07T13:55:32.917481Z","shell.execute_reply.started":"2022-08-07T13:55:32.898205Z","shell.execute_reply":"2022-08-07T13:55:32.916557Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Model Building","metadata":{}},{"cell_type":"code","source":"# import catboost\nfrom catboost import CatBoostClassifier\nfrom sklearn.metrics import accuracy_score\n\n# train catbost classifier\ndef train_cat(iterations, learning_rate, X, y, cat_features):\n    # initialize catboost\n    clf = CatBoostClassifier(iterations=iterations, learning_rate=learning_rate)\n    \n    # return fitted model\n    clf.fit(X, y, cat_features=cat_features)\n    \n    # print params\n    print('CatBoost model is fitted: ' + str(clf.is_fitted()))\n    print('CatBoost model parameters:')\n    print(clf.get_params())\n    return clf\n\n# test catboost classifier\ndef test_cat(clf, test_features, test_labels):\n    predictions = clf.predict(data=test_features)\n    return accuracy_score(test_labels, predictions=='True')\n\n# implements four-fold cross-validation for catboost model\ndef cross_validate_cat(iterations, learning_rate, data, labels, cat_features):\n    kf = KFold(n_splits = 4, shuffle = True)\n    fold_accs = []\n    for train_index, test_index in kf.split(data):\n        train_features, test_features = data.loc[train_index], data.loc[test_index]\n        train_labels, test_labels = labels.loc[train_index], labels.loc[test_index]\n        clf = train_cat(iterations, learning_rate, train_features, train_labels, cat_features)\n        fold_accs.append(test_cat(clf, test_features, test_labels))\n    return fold_accs","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:55:32.921806Z","iopub.execute_input":"2022-08-07T13:55:32.922194Z","iopub.status.idle":"2022-08-07T13:55:32.932510Z","shell.execute_reply.started":"2022-08-07T13:55:32.922164Z","shell.execute_reply":"2022-08-07T13:55:32.931150Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fold_accs = cross_validate_cat(10, 0.1, train.drop(columns = ['Transported']), train['Transported'], categoricals)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:55:32.934365Z","iopub.execute_input":"2022-08-07T13:55:32.935162Z","iopub.status.idle":"2022-08-07T13:55:33.581977Z","shell.execute_reply.started":"2022-08-07T13:55:32.935123Z","shell.execute_reply":"2022-08-07T13:55:33.580864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fold_accs","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:55:33.583012Z","iopub.execute_input":"2022-08-07T13:55:33.583300Z","iopub.status.idle":"2022-08-07T13:55:33.591375Z","shell.execute_reply.started":"2022-08-07T13:55:33.583273Z","shell.execute_reply":"2022-08-07T13:55:33.590173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Already we are getting some better performance from CatBoost! Let's take a look at our importance metrics from our features.","metadata":{}},{"cell_type":"code","source":"clf = train_cat(10, 0.1, train.drop(columns = ['Transported']), train['Transported'], categoricals)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:55:33.594260Z","iopub.execute_input":"2022-08-07T13:55:33.594664Z","iopub.status.idle":"2022-08-07T13:55:33.770612Z","shell.execute_reply.started":"2022-08-07T13:55:33.594539Z","shell.execute_reply":"2022-08-07T13:55:33.769714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"importances = clf.get_feature_importance()\nfeatures = train.drop(columns=['Transported']).columns.to_list()\nimportances = sorted(importances)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:55:33.771860Z","iopub.execute_input":"2022-08-07T13:55:33.772252Z","iopub.status.idle":"2022-08-07T13:55:33.779770Z","shell.execute_reply.started":"2022-08-07T13:55:33.772217Z","shell.execute_reply":"2022-08-07T13:55:33.778873Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=(12, 6))\nplt.barh(features, importances, align='center')\nplt.title('Feature Importance')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:55:33.781021Z","iopub.execute_input":"2022-08-07T13:55:33.781381Z","iopub.status.idle":"2022-08-07T13:55:33.992785Z","shell.execute_reply.started":"2022-08-07T13:55:33.781347Z","shell.execute_reply":"2022-08-07T13:55:33.991583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is quite curious as it seems in opposition to our findings of EDA! It seems that CryoSleep in fact had no importance rating in this model. Perhaps there is a way to combine our insights from exploration to boost this model's accuracy. ","metadata":{}},{"cell_type":"markdown","source":"### Grid Search\nTo find the best set of parameters for our model, we will use CatBoost's inbuilt grid search function. We are going to optimize parameters in batches. For our first grid search, we will optimize border_count and 12_leaf_reg independently. Using the best results, we will then optimize iterations and learning_rate together. Then, we will search for the best depth.","metadata":{}},{"cell_type":"code","source":"# define ranges for parameters that we are searching\nparams = [\n    {'border_count':[32,5,10,20,50,100,200]},\n    {'l2_leaf_reg':[3,1,5,10,100]},\n    {'iterations':[250,100,500,1000],\n     'learning_rate':[0.03,0.001,0.01,0.1,0.2,0.3]},\n    {'depth': [3,1,2,6,4,5,7,8,9,10]}\n]","metadata":{"execution":{"iopub.status.busy":"2022-08-07T14:17:42.532152Z","iopub.execute_input":"2022-08-07T14:17:42.532562Z","iopub.status.idle":"2022-08-07T14:17:42.538796Z","shell.execute_reply.started":"2022-08-07T14:17:42.532529Z","shell.execute_reply":"2022-08-07T14:17:42.537987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We are having an issue where the catboost classifier outputs strings of 'True' or 'False' for predictions, but our original values are boolean. Let's map our original labels to strings for consistency.","metadata":{}},{"cell_type":"code","source":"train['Transported'] = train['Transported'].astype('str')","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:55:34.018938Z","iopub.execute_input":"2022-08-07T13:55:34.019291Z","iopub.status.idle":"2022-08-07T13:55:34.034834Z","shell.execute_reply.started":"2022-08-07T13:55:34.019259Z","shell.execute_reply":"2022-08-07T13:55:34.033662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T13:55:34.036616Z","iopub.execute_input":"2022-08-07T13:55:34.036963Z","iopub.status.idle":"2022-08-07T13:55:34.059228Z","shell.execute_reply.started":"2022-08-07T13:55:34.036931Z","shell.execute_reply":"2022-08-07T13:55:34.058233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV\nclf=CatBoostClassifier()\n\n# iterate params grids\nfor param_grid in params:\n    grid_clf = GridSearchCV(estimator=clf, param_grid=param_grid, cv=3, n_jobs=-1, error_score='raise')\n    grid_clf.fit(train.drop(columns=['Transported']), train['Transported'], cat_features=categoricals, verbose=0)\n    print(grid_clf.best_params_)\n    print(grid_clf.best_score_)\n    clf = grid_clf.best_estimator_","metadata":{"execution":{"iopub.status.busy":"2022-08-07T14:19:20.155910Z","iopub.execute_input":"2022-08-07T14:19:20.156359Z","iopub.status.idle":"2022-08-07T14:41:30.521610Z","shell.execute_reply.started":"2022-08-07T14:19:20.156293Z","shell.execute_reply":"2022-08-07T14:41:30.520132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Our optimized model has a performance just under 0.8 (average of four folds). Out best parameters are:\n- border_count: 200\n- l2_leaf_ref: 10\n- iterations: 500\n- learning rate: 0.1\n- depth: 6\n\nFor now, let's go ahead and submit our predictions with this best model\n\n### Fit best estimator","metadata":{}},{"cell_type":"code","source":"# fit best estimator to all training data\nclf.fit(train.drop(columns=['Transported']), train['Transported'], cat_features=categoricals, verbose=0)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Format test set and generate predictions","metadata":{}},{"cell_type":"code","source":"# format testing set to match \ntest = pd.read_csv('../input/spaceship-titanic/test.csv')\ntest[['first_name', 'last_name']] = test['Name'].str.split(' ', expand=True)\ntest = test.drop(columns=['first_name', 'Name'])\ntest[['cabin_deck', 'cabin_num', 'cabin_side']] = test['Cabin'].str.split('/', expand = True)\ntest[['group', 'number']] =  test['PassengerId'].str.split('_', expand=True)\nnumerics = test.select_dtypes(include=np.number).columns.to_list()\ntest[numerics] = test[numerics].fillna(test[numerics].mean())\ntest[categoricals] = test[categoricals].fillna(test[categoricals].mode().iloc[0])\n\n# generate predictions\npredictions = clf.predict(test) == 'True'","metadata":{"execution":{"iopub.status.busy":"2022-08-07T14:54:04.755525Z","iopub.execute_input":"2022-08-07T14:54:04.756349Z","iopub.status.idle":"2022-08-07T14:54:04.871935Z","shell.execute_reply.started":"2022-08-07T14:54:04.756311Z","shell.execute_reply":"2022-08-07T14:54:04.870561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions","metadata":{"execution":{"iopub.status.busy":"2022-08-07T14:54:10.883906Z","iopub.execute_input":"2022-08-07T14:54:10.884989Z","iopub.status.idle":"2022-08-07T14:54:10.891552Z","shell.execute_reply.started":"2022-08-07T14:54:10.884939Z","shell.execute_reply":"2022-08-07T14:54:10.890402Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = test.PassengerId.copy().to_frame()\nsubmission['Transported'] = predictions","metadata":{"execution":{"iopub.status.busy":"2022-08-07T14:56:41.130923Z","iopub.execute_input":"2022-08-07T14:56:41.131357Z","iopub.status.idle":"2022-08-07T14:56:41.137395Z","shell.execute_reply.started":"2022-08-07T14:56:41.131316Z","shell.execute_reply":"2022-08-07T14:56:41.136536Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T14:56:43.076949Z","iopub.execute_input":"2022-08-07T14:56:43.078356Z","iopub.status.idle":"2022-08-07T14:56:43.092290Z","shell.execute_reply.started":"2022-08-07T14:56:43.078305Z","shell.execute_reply":"2022-08-07T14:56:43.090213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('submission.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T14:57:31.295341Z","iopub.execute_input":"2022-08-07T14:57:31.295729Z","iopub.status.idle":"2022-08-07T14:57:31.308490Z","shell.execute_reply.started":"2022-08-07T14:57:31.295698Z","shell.execute_reply":"2022-08-07T14:57:31.307207Z"},"trusted":true},"execution_count":null,"outputs":[]}]}