{"cells":[{"metadata":{},"cell_type":"markdown","source":"Thanks for the dataset (generated by TSFresh library) from [Alexander Lyubchenko](https://www.kaggle.com/carpediemamigo):\n- https://www.kaggle.com/carpediemamigo/ingv-catboost-baseline-tsfresh/data\n- https://www.kaggle.com/carpediemamigo/ingv-tsfresh-7730"},{"metadata":{},"cell_type":"markdown","source":"Since this competition permits the use of automated machine learning tool(s) (“AMLT”), this notebook uses **h2o automl** without tuning hyperparameters, and keeps the same features as my previous notebook (features after resampling):\n- https://www.kaggle.com/patrick0302/ingv-volcanic-eruption-prediction-add-resampling\n\nAlso, notebooks below are some very useful references:\n- https://www.kaggle.com/ajcostarino/ingv-volcanic-eruption-prediction-lgbm-baseline\n- https://www.kaggle.com/tunguz/lanl-earthquake-with-h2o-automl"},{"metadata":{},"cell_type":"markdown","source":"Here are present results with different runtimes:\n\n**Note: Even with the same settings (max_models, seed, and max_runtime_secs), the prediction results of each run seem to be somehow different.**\n\n\n|Trial|   Runtime(mins) |   Public Score |   AutoML validation score|   |\n|---:|----------:|---------------:|--------------------:|   |\n|  0 |         1 |    ~7.65320e+06  |         ~5.18657e+06 |   |\n|  1 |        10 |    ~6.16750e+06  |         ~3.74239e+06 |   |\n|  2 |        30 |    ~6.01172e+06 |         ~3.36789e+06 |   |\n|  3 |       120 |    ~5.81476e+06 |         ~3.27152e+06 |  **Best result !!!** |\n|  4 |       360 |    ~5.95679e+06 |          ~3.11598e+06\t|  **Overfitting :(** |\n\n\n\n"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport lightgbm as lgb\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import h2o\nprint(h2o.__version__)\nfrom h2o.automl import H2OAutoML\n\nh2o.init(max_mem_size='16G')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/predict-volcanic-eruptions-ingv-oe/train.csv')\ntest = pd.read_csv('/kaggle/input/predict-volcanic-eruptions-ingv-oe/sample_submission.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"scaled_feature_df = pd.read_csv('../input/ingv-tsfresh-7730/train.csv', sep = ';', index_col=0)\nscaled_feature_df = scaled_feature_df.loc[train['segment_id']]\nscaled_test_df = pd.read_csv('../input/ingv-tsfresh-7730/test.csv', sep = ';', index_col=0)\nscaled_test_df = scaled_test_df.loc[test['segment_id']]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Use lightgbm to select important features"},{"metadata":{"trusted":true},"cell_type":"code","source":"from sklearn.feature_selection import SelectFromModel","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sfm = SelectFromModel(estimator=lgb.LGBMRegressor())\nX = scaled_feature_df.drop('time_to_eruption',axis=1).copy()\nX.columns = list(np.arange(len(X.columns)))\ny = scaled_feature_df['time_to_eruption']\nsfm.fit(X, y)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"selected_features = list(scaled_feature_df.drop('time_to_eruption',axis=1).columns[sfm.get_support()])\nselected_features","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Number of selected features: ' + str(len(selected_features)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"scaled_feature_df = scaled_feature_df[selected_features]\nscaled_test_df = scaled_test_df[selected_features]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Create model (h2o automl)"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_h2o = h2o.H2OFrame(scaled_feature_df)\ntrain_label_h2o = h2o.H2OFrame(train[['time_to_eruption']])\ntrain_h2o['time_to_eruption'] = train_label_h2o['time_to_eruption']\n\ntest_feature_h2o = h2o.H2OFrame(scaled_test_df)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(train_h2o.shape)\nprint(test_feature_h2o.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"x = test_feature_h2o.columns\ny = 'time_to_eruption'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"aml = H2OAutoML(max_models=1000, seed=121, stopping_metric='MAE',\n                max_runtime_secs=360*60) # set 360 minutes\naml.train(x=x, y=y, training_frame=train_h2o)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# View the AutoML Leaderboard\nlb = aml.leaderboard\nlb.head(rows=lb.nrows)  # Print all rows instead of default (10 rows)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# The leader model is stored here\naml.leader","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# If you need to generate predictions on a test set, you can make\n# predictions directly on the `\"H2OAutoML\"` object, or on the leader\n# model object directly\n\npreds = aml.predict(test_feature_h2o)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"submission = pd.DataFrame()\nsubmission['segment_id'] = test['segment_id']\nsubmission['time_to_eruption'] = preds.as_data_frame().values.flatten()\nsubmission.loc[submission['time_to_eruption']<0, 'time_to_eruption'] = 0 #make sure all prediction values are larger than 0\nsubmission.to_csv('submission_recent.csv', header=True, index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"submission","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}