{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Beginner's Guide to Spaceship Titanic: Scikit-Learn Pipeline, RF, XGB, and LGBM","metadata":{"id":"v7nXbwV_McUV"}},{"cell_type":"markdown","source":"## Notebook Description","metadata":{"id":"0Voqsl8MQM6w"}},{"cell_type":"markdown","source":"* This notebook references the in-depth EDA and feature analysis I performed [in my prior Spaceship Titanic Notebook](https://www.kaggle.com/danb91/spacshiptitanic-starter-eda-rf-model-eval-80). \n* This notebook also builds upon that by improving the pipeline for the EDA and using the data to train RF, XGB, and LBG models.\n* Please provide feedback in the comments which will help me continue to improve notebooks as I publish them!\n","metadata":{"id":"s82IcyxQQQRi"}},{"cell_type":"markdown","source":"Credit to the following tutorials and notebooks that were extremely helpful while developing this notebook:\n* [Beginner’s Guide to XGBoost for Classification Problems](https://towardsdatascience.com/beginners-guide-to-xgboost-for-classification-problems-50f75aac5390)\n* [Lasso regression with Pipelines (Tutorial)](https://www.kaggle.com/code/bextuychiev/lasso-regression-with-pipelines-tutorial/notebook)\n* [You Are Missing Out on LightGBM. It Crushes XGBoost in Every Aspect](https://www.kaggle.com/code/prashant111/lightgbm-classifier-in-python/notebook)\n* [LightGBM Classifier in Python](https://www.kaggle.com/code/prashant111/lightgbm-classifier-in-python/notebook)\n* [Kaggler’s Guide to LightGBM Hyperparameter Tuning with Optuna in 2021](https://towardsdatascience.com/kagglers-guide-to-lightgbm-hyperparameter-tuning-with-optuna-in-2021-ed048d9838b5)","metadata":{"id":"JGep0UYtux7q"}},{"cell_type":"markdown","source":"## Import Libraries","metadata":{"id":"HEeBOntHQR8T"}},{"cell_type":"code","source":"# data analysis \nimport numpy as np\nimport pandas as pd\npd.set_option('display.max_columns', None)\npd.set_option('display.max_rows', None) \n\n# data visualization \nimport matplotlib.pyplot as plt  \nimport seaborn as sns\n\n# general utilities \nimport os\n","metadata":{"id":"yUFCry1ZMiR0","execution":{"iopub.status.busy":"2022-08-14T12:27:58.516778Z","iopub.execute_input":"2022-08-14T12:27:58.517746Z","iopub.status.idle":"2022-08-14T12:27:59.722654Z","shell.execute_reply.started":"2022-08-14T12:27:58.517615Z","shell.execute_reply":"2022-08-14T12:27:59.721310Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# sklearn tools and models\nfrom sklearn.model_selection import GridSearchCV, RandomizedSearchCV\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.preprocessing import OneHotEncoder, StandardScaler\nfrom sklearn.impute import SimpleImputer\n\n# XGBoost\nimport xgboost as xgb\n\n# LGBM\nimport lightgbm as lgbm","metadata":{"id":"Xv013vFcvnHo","execution":{"iopub.status.busy":"2022-08-14T12:27:59.724979Z","iopub.execute_input":"2022-08-14T12:27:59.725609Z","iopub.status.idle":"2022-08-14T12:28:01.517935Z","shell.execute_reply.started":"2022-08-14T12:27:59.725562Z","shell.execute_reply":"2022-08-14T12:28:01.516950Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Import Data ","metadata":{"id":"siatd-GhPpQR"}},{"cell_type":"code","source":"train_df = pd.read_csv(\"../input/spaceship-titanic/train.csv\")\ntest_df = pd.read_csv(\"../input/spaceship-titanic/test.csv\")\n\nprint('train dataframe dimensions:', train_df.shape)\nprint('test dataframe dimensions:', test_df.shape)","metadata":{"id":"NceohZoKRhDo","outputId":"a1a1375c-afad-4852-ae3a-b592a9f10e5e","execution":{"iopub.status.busy":"2022-08-14T12:28:01.519399Z","iopub.execute_input":"2022-08-14T12:28:01.520028Z","iopub.status.idle":"2022-08-14T12:28:01.598379Z","shell.execute_reply.started":"2022-08-14T12:28:01.519993Z","shell.execute_reply":"2022-08-14T12:28:01.597466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# combine into a single list\nfull_data = [train_df, test_df]\nlen(full_data)","metadata":{"id":"TZ1Xa52pYwjs","outputId":"1440961a-32fe-4553-f04f-6e31458d09bd","execution":{"iopub.status.busy":"2022-08-14T12:28:01.600396Z","iopub.execute_input":"2022-08-14T12:28:01.601048Z","iopub.status.idle":"2022-08-14T12:28:01.607153Z","shell.execute_reply.started":"2022-08-14T12:28:01.601012Z","shell.execute_reply":"2022-08-14T12:28:01.606109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Summary of Data Preprocessing\nThe following preprocessing steps are based on the EDA [in my prior Spaceship Titanic Notebook](https://www.kaggle.com/danb91/spacshiptitanic-starter-eda-rf-model-eval-80). ","metadata":{"id":"r7WKb7T7TTK3"}},{"cell_type":"markdown","source":"### Summary Steps\n* **PassengerId:** parse into Passenger_Group, Passenger_Num, and Group_Size\n* **HomePlanet:** impute most frequent and convert to categorical\n* **CryoSleep:** impute most frequent; convert to boolean\n* **Cabin:** Fill missing values with 'Z/99999/Z' and parse into new Cabin_Deck, Cabin_Num, and Cabin_Side features\n* **Destination:** impute most frequent and convert to categorical\n* **Age:** Impute median value\n* **VIP:** Impute most frequent and convert to boolean\n* **RoomService:** Impute median value\n* **FoodCourt:** Impute median value\n* **ShoppingMall:** Impute median value\n* **Spa:** Impute median value\n* **VRDeck:** Impute median value\n* **Name:** impute 'NoFirstName NoLastName'\n* **Transported:** converted to Boolean; train data only\n\n","metadata":{"id":"qqDfJAVHTS4A"}},{"cell_type":"markdown","source":"### Manual Preprocessing Steps","metadata":{"id":"9mf_DXK3TSun"}},{"cell_type":"code","source":"# convert Transported to boolean; only in train_df \ntrain_df['Transported'] = train_df['Transported'].astype(bool)","metadata":{"id":"uYnCEYXzqrdp","execution":{"iopub.status.busy":"2022-08-14T12:31:47.290017Z","iopub.execute_input":"2022-08-14T12:31:47.290426Z","iopub.status.idle":"2022-08-14T12:31:47.302154Z","shell.execute_reply.started":"2022-08-14T12:31:47.290394Z","shell.execute_reply":"2022-08-14T12:31:47.301268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# custom missing data fill\nfor df in full_data:\n  df['Name'].fillna('NoFirstName NoSurname', inplace=True)\n  df['Cabin'].fillna('Z/99999/Z', inplace=True)","metadata":{"id":"JKIOYKnWs2jr","execution":{"iopub.status.busy":"2022-08-14T12:31:47.773103Z","iopub.execute_input":"2022-08-14T12:31:47.773801Z","iopub.status.idle":"2022-08-14T12:31:47.784118Z","shell.execute_reply.started":"2022-08-14T12:31:47.773761Z","shell.execute_reply":"2022-08-14T12:31:47.782937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# CREATE FEATURES and format \nfor df in full_data:\n  # parse PassengerId\n  df['Passenger_Group'] = df['PassengerId'].apply(lambda x: x.split('_')[0]).astype('str')\n  df['Passenger_Num'] = df['PassengerId'].apply(lambda x: x.split('_')[1]).astype(str)\n  # parse Cabin\n  df['Cabin_Deck'] = df['Cabin'].apply(lambda x: x.split('/')[0]).astype('category')\n  df['Cabin_Num'] = df['Cabin'].apply(lambda x: x.split('/')[1]).astype(str)\n  df['Cabin_Side'] = df['Cabin'].apply(lambda x: x.split('/')[2]).astype('category') \n  # confirm dtypes for later processing \n  df['PassengerId'] = df['PassengerId'].astype(str)\n  df['HomePlanet'] = df['HomePlanet'].astype('category')\n  df['CryoSleep'] = df['CryoSleep'].astype('category')\n  df['Cabin'] = df['Cabin'].astype(str)\n  df['Destination'] = df['Destination'].astype('category')\n  df['VIP'] = df['VIP'].astype('category')\n","metadata":{"id":"LNc2gjPPtS2F","execution":{"iopub.status.busy":"2022-08-14T12:31:51.438366Z","iopub.execute_input":"2022-08-14T12:31:51.439373Z","iopub.status.idle":"2022-08-14T12:31:51.508924Z","shell.execute_reply.started":"2022-08-14T12:31:51.439339Z","shell.execute_reply":"2022-08-14T12:31:51.508081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# use completed features to generate new features\nfor df in full_data:\n  # calculate group size \n  df['Group_Size'] = df['Passenger_Group'].map(lambda x: pd.concat([train_df['Passenger_Group'], \n                                                                    test_df['Passenger_Group']]).value_counts()[x])\n  \n  df['Group_Size'] = df['Group_Size'].astype(float)\n  \n  ","metadata":{"id":"lslcXIB5wjKT","execution":{"iopub.status.busy":"2022-08-14T12:31:52.241282Z","iopub.execute_input":"2022-08-14T12:31:52.242036Z","iopub.status.idle":"2022-08-14T12:32:54.232186Z","shell.execute_reply.started":"2022-08-14T12:31:52.242000Z","shell.execute_reply":"2022-08-14T12:32:54.231118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T12:32:54.235410Z","iopub.execute_input":"2022-08-14T12:32:54.237118Z","iopub.status.idle":"2022-08-14T12:32:54.268923Z","shell.execute_reply.started":"2022-08-14T12:32:54.237079Z","shell.execute_reply":"2022-08-14T12:32:54.267946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Manually Create Additional Features","metadata":{}},{"cell_type":"code","source":"# define categorical and quantitative features \ncat_cols = train_df.select_dtypes(include='category').columns.tolist()\nquant_cols = train_df.select_dtypes(include='number').columns.tolist()\n","metadata":{"id":"b0iTGHs71esr","execution":{"iopub.status.busy":"2022-08-14T12:32:54.272619Z","iopub.execute_input":"2022-08-14T12:32:54.273013Z","iopub.status.idle":"2022-08-14T12:32:54.286200Z","shell.execute_reply.started":"2022-08-14T12:32:54.272981Z","shell.execute_reply":"2022-08-14T12:32:54.284497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# manually fill_na() in quant columns to calculate the total \nfor df in full_data:\n    for col in quant_cols:\n        df[col].fillna(value=df[col].median(), inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T12:32:54.288793Z","iopub.execute_input":"2022-08-14T12:32:54.289188Z","iopub.status.idle":"2022-08-14T12:32:54.307454Z","shell.execute_reply.started":"2022-08-14T12:32:54.289155Z","shell.execute_reply":"2022-08-14T12:32:54.306256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Create Onboard Spend Features**","metadata":{}},{"cell_type":"code","source":"# create Total_Spend feature \nfor df in full_data:\n    df['Total_Spend'] = (  df['RoomService'] \n                         + df['FoodCourt'] \n                         + df['ShoppingMall']\n                         + df['Spa'] \n                         + df['VRDeck']  )","metadata":{"execution":{"iopub.status.busy":"2022-08-14T12:32:54.308686Z","iopub.execute_input":"2022-08-14T12:32:54.309380Z","iopub.status.idle":"2022-08-14T12:32:54.317548Z","shell.execute_reply.started":"2022-08-14T12:32:54.309347Z","shell.execute_reply":"2022-08-14T12:32:54.316683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# create Any_Spend feature \nfor df in full_data:\n    df['Any_Spend'] = np.where(df['Total_Spend'] > 0, True, False)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T12:32:54.318693Z","iopub.execute_input":"2022-08-14T12:32:54.319033Z","iopub.status.idle":"2022-08-14T12:32:54.329012Z","shell.execute_reply.started":"2022-08-14T12:32:54.319002Z","shell.execute_reply":"2022-08-14T12:32:54.328125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['Any_Spend'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T12:32:54.330194Z","iopub.execute_input":"2022-08-14T12:32:54.330688Z","iopub.status.idle":"2022-08-14T12:32:54.342832Z","shell.execute_reply.started":"2022-08-14T12:32:54.330644Z","shell.execute_reply":"2022-08-14T12:32:54.341948Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"round(train_df[train_df['Total_Spend'] > 0]['Total_Spend'].describe(), 2)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T12:32:54.344416Z","iopub.execute_input":"2022-08-14T12:32:54.346108Z","iopub.status.idle":"2022-08-14T12:32:54.364760Z","shell.execute_reply.started":"2022-08-14T12:32:54.346062Z","shell.execute_reply":"2022-08-14T12:32:54.363947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for df in full_data:\n    df['Spend_Category'] = ''\n    df.loc[df['Total_Spend'].between(0, 1, 'left'), 'Spend_Category'] = 'Zero_Spend'\n    df.loc[df['Total_Spend'].between(1, 800, 'both'), 'Spend_Category'] = 'Under_800'\n    df.loc[df['Total_Spend'].between(800, 1200, 'right'), 'Spend_Category'] = 'Median_1200'\n    df.loc[df['Total_Spend'].between(1200, 2700, 'right'), 'Spend_Category'] = 'Upper_2700'\n    df.loc[df['Total_Spend'].between(2700, 100000, 'right'), 'Spend_Category'] = 'Big_Spender'\n    df['Spend_Category'] = df['Spend_Category'].astype('category')","metadata":{"execution":{"iopub.status.busy":"2022-08-14T12:32:54.365765Z","iopub.execute_input":"2022-08-14T12:32:54.366178Z","iopub.status.idle":"2022-08-14T12:32:54.391370Z","shell.execute_reply.started":"2022-08-14T12:32:54.366146Z","shell.execute_reply":"2022-08-14T12:32:54.390388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Create Age Category**","metadata":{}},{"cell_type":"code","source":"for df in full_data:\n    df['Age_Category'] = ''\n    df.loc[df['Age'].between(0, 18, 'both'), 'Age_Category'] = 'Under_18'\n    df.loc[df['Age'].between(18, 40, 'right'), 'Age_Category'] = 'Adult'\n    df.loc[df['Age'].between(40, 60, 'right'), 'Age_Category'] = 'Middle_Age'\n    df.loc[df['Age'].between(60, 100, 'right'), 'Age_Category'] = 'Over_60'\n    df['Age_Category'] = df['Age_Category'].astype('category')","metadata":{"execution":{"iopub.status.busy":"2022-08-14T12:32:54.395684Z","iopub.execute_input":"2022-08-14T12:32:54.396445Z","iopub.status.idle":"2022-08-14T12:32:54.417239Z","shell.execute_reply.started":"2022-08-14T12:32:54.396406Z","shell.execute_reply":"2022-08-14T12:32:54.415970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Preprocessing Pipeline ","metadata":{"id":"p5WQHg6J4BLW"}},{"cell_type":"markdown","source":"**Create Pipelines**","metadata":{"id":"5bfSPAqBqC7J"}},{"cell_type":"code","source":"# update feature lists \ncat_cols = train_df.select_dtypes(include='category').columns.tolist()\nquant_cols = train_df.select_dtypes(include='number').columns.tolist()\n","metadata":{"execution":{"iopub.status.busy":"2022-08-14T12:32:54.419116Z","iopub.execute_input":"2022-08-14T12:32:54.419479Z","iopub.status.idle":"2022-08-14T12:32:54.428698Z","shell.execute_reply.started":"2022-08-14T12:32:54.419440Z","shell.execute_reply":"2022-08-14T12:32:54.427673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_cols","metadata":{"execution":{"iopub.status.busy":"2022-08-14T12:32:54.430420Z","iopub.execute_input":"2022-08-14T12:32:54.430762Z","iopub.status.idle":"2022-08-14T12:32:54.442252Z","shell.execute_reply.started":"2022-08-14T12:32:54.430731Z","shell.execute_reply":"2022-08-14T12:32:54.441155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"quant_cols","metadata":{"execution":{"iopub.status.busy":"2022-08-14T12:32:54.443711Z","iopub.execute_input":"2022-08-14T12:32:54.444083Z","iopub.status.idle":"2022-08-14T12:32:54.452677Z","shell.execute_reply.started":"2022-08-14T12:32:54.444051Z","shell.execute_reply":"2022-08-14T12:32:54.451909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_pipeline = Pipeline(steps=[\n    ('impute', SimpleImputer(strategy='most_frequent')),\n    ('one-hot', OneHotEncoder(drop='first', handle_unknown='ignore', sparse=False))\n])","metadata":{"id":"zb-14Fen8OaQ","execution":{"iopub.status.busy":"2022-08-14T12:32:54.454353Z","iopub.execute_input":"2022-08-14T12:32:54.455071Z","iopub.status.idle":"2022-08-14T12:32:54.463172Z","shell.execute_reply.started":"2022-08-14T12:32:54.455024Z","shell.execute_reply":"2022-08-14T12:32:54.461850Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"quant_pipeline = Pipeline(steps=[\n    ('impute', SimpleImputer(strategy='median')),\n    ('scale', StandardScaler())\n])","metadata":{"id":"kJwqX88o5TIb","execution":{"iopub.status.busy":"2022-08-14T12:32:54.464507Z","iopub.execute_input":"2022-08-14T12:32:54.465489Z","iopub.status.idle":"2022-08-14T12:32:54.473245Z","shell.execute_reply.started":"2022-08-14T12:32:54.465455Z","shell.execute_reply":"2022-08-14T12:32:54.472126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# combine categorical and quantitative pipelines \nfull_pipeline = ColumnTransformer(transformers=[\n    ('category', cat_pipeline, cat_cols),\n    ('quant', quant_pipeline, quant_cols)\n])","metadata":{"id":"ZzXBCzGaHg7H","execution":{"iopub.status.busy":"2022-08-14T12:32:54.474691Z","iopub.execute_input":"2022-08-14T12:32:54.475166Z","iopub.status.idle":"2022-08-14T12:32:54.483345Z","shell.execute_reply.started":"2022-08-14T12:32:54.475134Z","shell.execute_reply":"2022-08-14T12:32:54.482460Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Preprocess Train and Test Data**","metadata":{"id":"l6jIXzTRqFBq"}},{"cell_type":"code","source":"X_train = full_pipeline.fit_transform(train_df)\nX_train.shape","metadata":{"id":"FIJQEitGIA16","outputId":"13bae459-96b7-477c-c14f-a8947a7b2723","execution":{"iopub.status.busy":"2022-08-14T12:32:54.485321Z","iopub.execute_input":"2022-08-14T12:32:54.485755Z","iopub.status.idle":"2022-08-14T12:32:54.565105Z","shell.execute_reply.started":"2022-08-14T12:32:54.485714Z","shell.execute_reply":"2022-08-14T12:32:54.563668Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test = full_pipeline.fit_transform(test_df)\nX_test.shape","metadata":{"id":"as2YmfDQiSrj","outputId":"85421130-b0a3-4c14-bc2f-80df265afb8b","execution":{"iopub.status.busy":"2022-08-14T12:32:54.568996Z","iopub.execute_input":"2022-08-14T12:32:54.569582Z","iopub.status.idle":"2022-08-14T12:32:54.616046Z","shell.execute_reply.started":"2022-08-14T12:32:54.569536Z","shell.execute_reply":"2022-08-14T12:32:54.615188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# create y as the target value \ny = train_df['Transported']\ny.shape","metadata":{"id":"RIEs2OFGi3ax","outputId":"7f39d7d3-57ac-43e6-ec41-23307f82e6b3","execution":{"iopub.status.busy":"2022-08-14T12:32:54.617781Z","iopub.execute_input":"2022-08-14T12:32:54.618507Z","iopub.status.idle":"2022-08-14T12:32:54.625919Z","shell.execute_reply.started":"2022-08-14T12:32:54.618463Z","shell.execute_reply":"2022-08-14T12:32:54.625050Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Check Feature Correlation","metadata":{"id":"o5nnqOMBAggv"}},{"cell_type":"code","source":"# correlation map of the fatures \ncolormap = plt.cm.RdBu\nplt.figure(figsize=(18,14))\nplt.title('Pearson Correlation of Features', y=1.05, size=15)\nsns.heatmap(pd.DataFrame(X_train).astype(float).corr(),\n            linewidths=0.1,\n            vmax=1.0, \n            square=True, \n            cmap=colormap, \n            linecolor='white', \n            annot=True)\n\nplt.show()","metadata":{"id":"QeQCUpdzD4Hl","outputId":"352dd253-c26c-4f10-d057-f52c41e5bc79","execution":{"iopub.status.busy":"2022-08-14T12:32:54.627238Z","iopub.execute_input":"2022-08-14T12:32:54.628332Z","iopub.status.idle":"2022-08-14T12:32:58.677814Z","shell.execute_reply.started":"2022-08-14T12:32:54.628286Z","shell.execute_reply":"2022-08-14T12:32:58.676642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train Random Forest Model","metadata":{"id":"B5Lxz8eVf8b-"}},{"cell_type":"markdown","source":"Hyperparameters are centered around prior best models: \n* parameters: {'n_estimators': 360, 'min_samples_split': 24, 'min_samples_leaf': 1, 'max_features': 'log2', 'max_depth': 16, 'criterion': 'gini'}\n\n* score: 0.7985","metadata":{}},{"cell_type":"code","source":"rf_model_cv = RandomForestClassifier(oob_score=True, random_state=1)\n\nparam_grid_rf = {'criterion' : [\"gini\"], \n                 'max_features': ['log2'],\n                 'max_depth': np.arange(15, 20, 1),\n                 'min_samples_leaf' : np.arange(1, 5, 1), \n                 'min_samples_split' : np.arange(20, 30, 1), \n                 'n_estimators': np.arange(350, 400, 5)\n                }\n\ngs_rf = RandomizedSearchCV(estimator=rf_model_cv, \n                           param_distributions=param_grid_rf, \n                           scoring='accuracy', \n                           n_iter = 50,\n                           cv=5, \n                           n_jobs=-1)\n\ngs_rf.fit(X_train, y)","metadata":{"id":"o6SIP8Ka7Fnh","outputId":"f4638f56-02ce-4b7f-8bd8-afce3c9aee8a","execution":{"iopub.status.busy":"2022-08-14T14:46:16.750433Z","iopub.execute_input":"2022-08-14T14:46:16.750906Z","iopub.status.idle":"2022-08-14T14:50:50.167455Z","shell.execute_reply.started":"2022-08-14T14:46:16.750852Z","shell.execute_reply":"2022-08-14T14:50:50.166116Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(round(gs_rf.best_score_, 4))","metadata":{"id":"dO8JPrwXFotS","outputId":"345b1653-2898-467f-ed0c-afe25bb2c540","execution":{"iopub.status.busy":"2022-08-14T14:50:50.170270Z","iopub.execute_input":"2022-08-14T14:50:50.170800Z","iopub.status.idle":"2022-08-14T14:50:50.176423Z","shell.execute_reply.started":"2022-08-14T14:50:50.170740Z","shell.execute_reply":"2022-08-14T14:50:50.175546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(gs_rf.best_params_)","metadata":{"id":"8ZeCVIdbFmIe","outputId":"2f7061e5-d139-4265-9656-51923bbfba56","execution":{"iopub.status.busy":"2022-08-14T14:50:50.177709Z","iopub.execute_input":"2022-08-14T14:50:50.178909Z","iopub.status.idle":"2022-08-14T14:50:50.189945Z","shell.execute_reply.started":"2022-08-14T14:50:50.178832Z","shell.execute_reply":"2022-08-14T14:50:50.188603Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train XGBoost Classifier","metadata":{"id":"56q_HM8SSqG4"}},{"cell_type":"markdown","source":"Grid Search is centered around prior best model:\n* {'subsample': 0.80, 'reg_lambda': 0.55, 'reg_alpha': 0.25, 'max_depth': 4, 'learning_rate': 0.15, 'gamma': 0.45, 'colsample_bytree': 0.95}\n* 0.801","metadata":{}},{"cell_type":"code","source":"# instantiate model\nxgb_model_cv = xgb.XGBClassifier(objective = 'binary:logistic', seed=1)\n\nparam_grid_xgb = {'learning_rate': np.arange(0.12, 0.18, 0.01),\n                  'gamma': np.arange(0.4, 0.55, 0.01), \n                  'reg_alpha': np.arange(0.02, 0.35, .1), \n                  'reg_lambda': np.arange(0.45, 0.55, 0.1),\n                  'max_depth': np.arange(2, 5, 1),\n                  'subsample': np.arange(0.75, 0.85, 0.1),\n                  'colsample_bytree': np.arange(0.9, 1, .01)}\n\ngs_xgb = RandomizedSearchCV(estimator=xgb_model_cv, \n                            param_distributions=param_grid_xgb, \n                            scoring='accuracy', \n                            n_iter = 50,\n                            cv=5, \n                            n_jobs=-1)\n\ngs_xgb.fit(X_train, y)","metadata":{"id":"BOD1SsFiSpoi","outputId":"7ddceb56-f617-48e3-e040-f3bc59561199","execution":{"iopub.status.busy":"2022-08-14T14:51:32.539190Z","iopub.execute_input":"2022-08-14T14:51:32.540270Z","iopub.status.idle":"2022-08-14T14:53:27.588336Z","shell.execute_reply.started":"2022-08-14T14:51:32.540225Z","shell.execute_reply":"2022-08-14T14:53:27.587408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"round(gs_xgb.best_score_, 4)","metadata":{"id":"vAJ_or-drj7w","outputId":"0e121a15-6606-43ba-c1c5-36d52a4b35ea","execution":{"iopub.status.busy":"2022-08-14T14:53:27.590136Z","iopub.execute_input":"2022-08-14T14:53:27.590685Z","iopub.status.idle":"2022-08-14T14:53:27.597779Z","shell.execute_reply.started":"2022-08-14T14:53:27.590650Z","shell.execute_reply":"2022-08-14T14:53:27.596342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gs_xgb.best_params_","metadata":{"id":"l1u_zK3crkIL","outputId":"a9a6b856-935d-4876-c682-ee72ff1aec86","execution":{"iopub.status.busy":"2022-08-14T14:53:27.599207Z","iopub.execute_input":"2022-08-14T14:53:27.599535Z","iopub.status.idle":"2022-08-14T14:53:27.614442Z","shell.execute_reply.started":"2022-08-14T14:53:27.599506Z","shell.execute_reply":"2022-08-14T14:53:27.613613Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train LGBM Classifier","metadata":{"id":"AHc5OKcJL96I"}},{"cell_type":"markdown","source":"Parameter range based on previous best parameters:\n* {'reg_lambda': 0.86, 'reg_alpha': 0.46, 'num_leaves': 12, 'n_estimators': 490, 'min_child_samples': 45, 'max_depth': 7, 'learning_rate': 0.09, 'boosting_type': 'dart'}\n* Accuracy Score: 0.8071","metadata":{}},{"cell_type":"code","source":"lgbm_classifier = lgbm.LGBMClassifier(objective='binary') \nlgbm_classifier.fit(X=X_train, y=y)","metadata":{"id":"cdjZQSZnMEjH","outputId":"d7856e6f-32ed-4e5a-a6c6-b496011473a8","execution":{"iopub.status.busy":"2022-08-13T14:42:08.148636Z","iopub.execute_input":"2022-08-13T14:42:08.149828Z","iopub.status.idle":"2022-08-13T14:42:08.395040Z","shell.execute_reply.started":"2022-08-13T14:42:08.149782Z","shell.execute_reply":"2022-08-13T14:42:08.394148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# instantiate model\nlgbm_model_cv = lgbm.LGBMClassifier(objective = 'binary', random_state=1)\n\nparam_grid_lgbm = {'boosting_type': ['dart'],\n                   'num_leaves': np.arange(10, 15, 1),\n                   'max_depth': np.arange(6, 8, 1),\n                   'learning_rate': np.arange(0.07, 0.12, 0.01),\n                   'n_estimators': np.arange(480, 510, 2),\n                   'reg_alpha': np.arange(0.4, 0.6, 0.02), \n                   'min_child_samples': np.arange(40, 60, 2),\n                   'reg_lambda': np.arange(0.85, 1, 0.01)\n                  }\n\ngs_lgbm = RandomizedSearchCV(estimator=lgbm_model_cv, \n                             param_distributions=param_grid_lgbm, \n                             scoring='accuracy', \n                             n_iter = 250,\n                             cv=5, \n                             n_jobs=-1)\n\ngs_lgbm.fit(X_train, y)","metadata":{"id":"mXLHKz-gNskY","outputId":"58afc071-8163-4433-e040-7737e1b92792","execution":{"iopub.status.busy":"2022-08-14T14:55:58.484518Z","iopub.execute_input":"2022-08-14T14:55:58.485325Z","iopub.status.idle":"2022-08-14T15:34:16.996047Z","shell.execute_reply.started":"2022-08-14T14:55:58.485283Z","shell.execute_reply":"2022-08-14T15:34:16.994370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"round(gs_lgbm.best_score_, 4)","metadata":{"id":"X5XXanFpcmqK","outputId":"022bb8e0-58d0-4743-d0e5-34b258e352cf","execution":{"iopub.status.busy":"2022-08-14T16:34:08.223295Z","iopub.execute_input":"2022-08-14T16:34:08.223845Z","iopub.status.idle":"2022-08-14T16:34:08.245743Z","shell.execute_reply.started":"2022-08-14T16:34:08.223803Z","shell.execute_reply":"2022-08-14T16:34:08.244494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gs_lgbm.best_params_","metadata":{"id":"FiLagrkxNsh2","outputId":"155d80a9-bba5-4b98-c8c1-ae150271ff9f","execution":{"iopub.status.busy":"2022-08-14T15:34:17.008748Z","iopub.execute_input":"2022-08-14T15:34:17.009137Z","iopub.status.idle":"2022-08-14T15:34:17.019973Z","shell.execute_reply.started":"2022-08-14T15:34:17.009103Z","shell.execute_reply":"2022-08-14T15:34:17.019086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Final Model","metadata":{"id":"jOyUaRcAsEPu"}},{"cell_type":"code","source":"# based on best run\nlgbm_model = lgbm.LGBMClassifier(objective='binary',\n                                 boosting_type='dart',\n                                 learning_rate=0.11,\n                                 max_depth=7,\n                                 min_child_samples=50,\n                                 n_estimators=505,\n                                 num_leaves=11,\n                                 reg_alpha=0.48,\n                                 reg_lambda=0.95)","metadata":{"id":"2cbcWOjwsEKB","execution":{"iopub.status.busy":"2022-08-14T13:36:56.396320Z","iopub.execute_input":"2022-08-14T13:36:56.396718Z","iopub.status.idle":"2022-08-14T13:36:56.403181Z","shell.execute_reply.started":"2022-08-14T13:36:56.396687Z","shell.execute_reply":"2022-08-14T13:36:56.402211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# fit to training data \nlgbm_model.fit(X_train, y)\n\n# predict on test data \nlgbm_predictions = lgbm_model.predict(X_test)\n","metadata":{"id":"426sb75lNsd9","execution":{"iopub.status.busy":"2022-08-14T13:37:00.908249Z","iopub.execute_input":"2022-08-14T13:37:00.909472Z","iopub.status.idle":"2022-08-14T13:37:03.703784Z","shell.execute_reply.started":"2022-08-14T13:37:00.909428Z","shell.execute_reply":"2022-08-14T13:37:03.702656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Final Output for Submission","metadata":{"id":"BK8g1LQbH4F4"}},{"cell_type":"code","source":"output = pd.DataFrame({'PassengerId': test_df.PassengerId, 'Transported': lgbm_predictions})\noutput.shape","metadata":{"id":"F-PsqoiVIetI","outputId":"75671d61-cb1a-4e94-e1e4-b8072d14a65a","execution":{"iopub.status.busy":"2022-08-14T13:37:03.705662Z","iopub.execute_input":"2022-08-14T13:37:03.706016Z","iopub.status.idle":"2022-08-14T13:37:03.720184Z","shell.execute_reply.started":"2022-08-14T13:37:03.705983Z","shell.execute_reply":"2022-08-14T13:37:03.718994Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"output.head()","metadata":{"id":"chOWGVc4wVBg","outputId":"0a6c2a10-0c3b-4349-9e29-a119258bc4e6","execution":{"iopub.status.busy":"2022-08-14T13:37:06.820207Z","iopub.execute_input":"2022-08-14T13:37:06.820615Z","iopub.status.idle":"2022-08-14T13:37:06.832065Z","shell.execute_reply.started":"2022-08-14T13:37:06.820580Z","shell.execute_reply":"2022-08-14T13:37:06.831084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"output.to_csv('submission.csv', index=False)\nprint(\"Your submission was successfully saved!\")","metadata":{"id":"6xYK_jfCI1gk","execution":{"iopub.status.busy":"2022-08-14T13:37:07.781465Z","iopub.execute_input":"2022-08-14T13:37:07.782239Z","iopub.status.idle":"2022-08-14T13:37:07.797876Z","shell.execute_reply.started":"2022-08-14T13:37:07.782201Z","shell.execute_reply":"2022-08-14T13:37:07.796622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# provided by kaggle starter code \nfor dirname, _, filenames in os.walk('/kaggle/working'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"id":"vsL_SuEvJyMU","execution":{"iopub.status.busy":"2022-08-13T15:12:36.165727Z","iopub.execute_input":"2022-08-13T15:12:36.166944Z","iopub.status.idle":"2022-08-13T15:12:36.174044Z","shell.execute_reply.started":"2022-08-13T15:12:36.166902Z","shell.execute_reply":"2022-08-13T15:12:36.172846Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}