{"cells":[{"metadata":{},"cell_type":"markdown","source":"# The Volcano and the Regularized Greedy Forest\nThis is a demonstration script using the ***Regularized Greedy Forest*** regressor (RGF)(see my notebook [\"Introduction to the Regularized Greedy Forest\"](https://www.kaggle.com/carlmcbrideellis/introduction-to-the-regularized-greedy-forest)) for the [INGV - Volcanic Eruption Prediction](https://www.kaggle.com/c/predict-volcanic-eruptions-ingv-oe) competition. The RGF performs as well as XGBoost, and is a very useful estimator to include when one is creating a [stacking ensemble](https://www.kaggle.com/carlmcbrideellis/stacking-ensemble-using-the-house-prices-data), which combines multiple estimators to produce one strong result. For the input I use the `train.csv` and `test.csv` produced by the excellent notebook [\"INGV Volcanic Eruption Prediction - LGBM Baseline\"](https://www.kaggle.com/ajcostarino/ingv-volcanic-eruption-prediction-lgbm-baseline) written by [Adam James](https://www.kaggle.com/ajcostarino). (For completeness I include these `train.csv` and `test.csv` files in the **Output** section of this notebook, as they take nearly three hours to produce). I have not undertaken any feature selection (for example using the [Boruta-SHAP](https://www.kaggle.com/carlmcbrideellis/feature-selection-using-the-borutashap-package) package), nor have I performed any cross validation, hyperparameter tuning, *etc.* so there is *plenty* of room for improvement.\n\nI hope you find the RGF technique useful, and good luck!"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import pandas  as pd\nimport numpy   as np","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true,"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":false},"cell_type":"code","source":"train  = pd.read_csv('../input/ingv-lgbm-baseline-the-train-test-csv-files/volcano_train.csv')\ntest   = pd.read_csv('../input/ingv-lgbm-baseline-the-train-test-csv-files/volcano_test.csv')\nsample = pd.read_csv('../input/predict-volcanic-eruptions-ingv-oe/sample_submission.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"X_train       = train.drop([\"segment_id\",\"time_to_eruption\"],axis=1)\ny_train       = train[\"time_to_eruption\"]\nX_test        = test.drop(\"segment_id\",axis=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from rgf.sklearn import RGFRegressor\n\nregressor = RGFRegressor(max_leaf=2000, \n                         algorithm=\"RGF_Sib\", \n                         test_interval=100, \n                         loss=\"LS\",\n                         verbose=False)\n\nregressor.fit(X_train, y_train)\npredictions = regressor.predict(X_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sample.iloc[:,1:] = predictions\nsample.to_csv('submission.csv',index=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Appendix\nWrite out a copy of the `train.csv` and `test.csv` files used in this work."},{"metadata":{"trusted":true},"cell_type":"code","source":"train.to_csv('volcano_train.csv')\ntest.to_csv('volcano_test.csv')","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}