{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Volcanic feature importance using Boruta-SHAP\nIn this notebook we shall produce a selection of the most important features of the [INGV - Volcanic Eruption Prediction](https://www.kaggle.com/c/predict-volcanic-eruptions-ingv-oe) data using the [Boruta-SHAP](https://www.kaggle.com/carlmcbrideellis/feature-selection-using-the-borutashap-package) package.\nFor the input I use the `train.csv` produced by the excellent notebook [\"INGV Volcanic Eruption Prediction - LGBM Baseline\"](https://www.kaggle.com/ajcostarino/ingv-volcanic-eruption-prediction-lgbm-baseline) written by [Adam James](https://www.kaggle.com/ajcostarino). I shall write the results of the feature selection to the file `selected_features.csv`."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"import numpy   as np\nimport pandas  as pd\npd.set_option('display.max_columns', None)\n!pip install BorutaShap","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true,"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":false,"_kg_hide-output":true},"cell_type":"code","source":"train   = pd.read_csv('../input/the-volcano-and-the-regularized-greedy-forest/volcano_train.csv')\nX_train = train.drop([\"segment_id\",\"time_to_eruption\"],axis=1)\ny_train = train[\"time_to_eruption\"]\n\nfrom xgboost import XGBRegressor\nmodel = XGBRegressor()\n\nfrom BorutaShap import BorutaShap\nFeature_Selector = BorutaShap(model=model,importance_measure='shap', classification=False)\nFeature_Selector.fit(X=X_train, y=y_train, n_trials=35, random_state=0);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Produce a box-plot of the accepted features"},{"metadata":{"trusted":true},"cell_type":"code","source":"Feature_Selector.plot(which_features='accepted', figsize=(20,12))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Return a subset of the original data with the selected features"},{"metadata":{"trusted":true},"cell_type":"code","source":"selected_features = Feature_Selector.Subset()\nselected_features","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"write out a `selected_features.csv` file"},{"metadata":{"trusted":true},"cell_type":"code","source":"selected_features.to_csv('selected_features.csv',index=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Produce a `submission.csv` using the RGF\nFor completeness we shall produce a `submission.csv`, here using the [Regularized Greedy Forest](https://www.kaggle.com/carlmcbrideellis/introduction-to-the-regularized-greedy-forest) for the estimator:"},{"metadata":{"trusted":true},"cell_type":"code","source":"test   = pd.read_csv('../input/the-volcano-and-the-regularized-greedy-forest/volcano_test.csv')\nX_train          = selected_features\nselected_columns = selected_features.columns\nX_test           = test[selected_columns]\n\nfrom rgf.sklearn import RGFRegressor\nregressor = RGFRegressor(max_leaf=10000, \n                         algorithm=\"RGF_Sib\", \n                         test_interval=100, \n                         loss=\"LS\",\n                         verbose=False)\nregressor.fit(X_train, y_train)\npredictions = regressor.predict(X_test)\n\nsample = pd.read_csv('../input/predict-volcanic-eruptions-ingv-oe/sample_submission.csv')\nsample.iloc[:,1:] = predictions\nsample.to_csv('submission.csv',index=False)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}