{
  "id": 211120,
  "title": "🌋 18th Place Solution + processed dataset",
  "url": "/competitions/predict-volcanic-eruptions-ingv-oe/discussion/211120",
  "author_name": "Ekhtiar Syed",
  "post_date": "2021-01-13T17:11:16.327000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi Kaggle Community,</p>\n<p>I wanted to share the work I did for this competition for others so I can have some feedback on what I could have done better. I also wanted to do something for the newcomers in data science, so I cleaned up my competition  and the dataset, added some documentation, and made it available as public. </p>\n<p>I hope this also is helpful for the Volcanic Activity Research Community, and the machine learning research community in general. </p>\n<p>📖 Notebook - <a href=\"https://www.kaggle.com/ekhtiar/18th-place-predicting-eruption-full-tutorial\" target=\"_blank\">Predicting 🌋 Eruption Full Tutorial\n</a><br>\n📈 Processed dataset - <a href=\"https://www.kaggle.com/ekhtiar/ingv-parquet\" target=\"_blank\">INGV Processed Features With ts-fresh\n</a></p>\n<p><strong>Little About The Processing…</strong><br>\nFirst of all, I processed of all I processed a wide number of features for the original dataset using tsfresh. Tsfresh is an amazing package that automates feature generation of timeseries data. You can find the processed features dataset here.</p>\n<p>As some features are very computation heavy, I either needed to downsample the 60000 datapoints per segment or process in smaller batches. I choose to go with creating multiple batches per segment. So, for each segment I divided the datapoints into 6 pieces (10,000 datapoints per piece).</p>\n<p>Currently, we have 2854 features for our dataset. At first I processed all the possible features of the train dataset from the INGV competition using ts-fresh library (ComprehensiveFCParameters). This generated almost 8000 features. Then I removed highly correlated columns, and quasi-constant features. This brought our features down to 2854. I also applied a recursive feature elimination to take the top 501 features (500 seemed too goodie-to-shoe of a number). These columns are hard-coded in this notebook.</p>\n<p><strong>Little About Modelling</strong><br>\nFor making this prediction, I have used LGBMRegressor from the LightGBM (LGBM) framework. I am taking a two-fold approach, where I first use a single LGBM model for the entire dataset. This model is used on the test set to make an initial prediction. Then for multiple LGBM models is created for different ranges of time to eruption. Since these models concentrates on a specific range, they can be more specialized. Then finally, we have six output or prediction for each segment in our test dataset. We take the median of these outputs to get our final prediction.</p>",
  "messages": [
    {
      "id": 1151971,
      "postDate": "2021-01-13T17:11:16.327Z",
      "content": "<p>Hi Kaggle Community,</p>\n<p>I wanted to share the work I did for this competition for others so I can have some feedback on what I could have done better. I also wanted to do something for the newcomers in data science, so I cleaned up my competition  and the dataset, added some documentation, and made it available as public. </p>\n<p>I hope this also is helpful for the Volcanic Activity Research Community, and the machine learning research community in general. </p>\n<p>📖 Notebook - <a href=\"https://www.kaggle.com/ekhtiar/18th-place-predicting-eruption-full-tutorial\" target=\"_blank\">Predicting 🌋 Eruption Full Tutorial\n</a><br>\n📈 Processed dataset - <a href=\"https://www.kaggle.com/ekhtiar/ingv-parquet\" target=\"_blank\">INGV Processed Features With ts-fresh\n</a></p>\n<p><strong>Little About The Processing…</strong><br>\nFirst of all, I processed of all I processed a wide number of features for the original dataset using tsfresh. Tsfresh is an amazing package that automates feature generation of timeseries data. You can find the processed features dataset here.</p>\n<p>As some features are very computation heavy, I either needed to downsample the 60000 datapoints per segment or process in smaller batches. I choose to go with creating multiple batches per segment. So, for each segment I divided the datapoints into 6 pieces (10,000 datapoints per piece).</p>\n<p>Currently, we have 2854 features for our dataset. At first I processed all the possible features of the train dataset from the INGV competition using ts-fresh library (ComprehensiveFCParameters). This generated almost 8000 features. Then I removed highly correlated columns, and quasi-constant features. This brought our features down to 2854. I also applied a recursive feature elimination to take the top 501 features (500 seemed too goodie-to-shoe of a number). These columns are hard-coded in this notebook.</p>\n<p><strong>Little About Modelling</strong><br>\nFor making this prediction, I have used LGBMRegressor from the LightGBM (LGBM) framework. I am taking a two-fold approach, where I first use a single LGBM model for the entire dataset. This model is used on the test set to make an initial prediction. Then for multiple LGBM models is created for different ranges of time to eruption. Since these models concentrates on a specific range, they can be more specialized. Then finally, we have six output or prediction for each segment in our test dataset. We take the median of these outputs to get our final prediction.</p>",
      "rawMarkdown": "Hi Kaggle Community,\n\nI wanted to share the work I did for this competition for others so I can have some feedback on what I could have done better. I also wanted to do something for the newcomers in data science, so I cleaned up my competition  and the dataset, added some documentation, and made it available as public. \n\nI hope this also is helpful for the Volcanic Activity Research Community, and the machine learning research community in general. \n\n📖 Notebook - [Predicting 🌋 Eruption Full Tutorial\n](https://www.kaggle.com/ekhtiar/18th-place-predicting-eruption-full-tutorial)\n📈 Processed dataset - [INGV Processed Features With ts-fresh\n](https://www.kaggle.com/ekhtiar/ingv-parquet)\n\n**Little About The Processing...**\nFirst of all, I processed of all I processed a wide number of features for the original dataset using tsfresh. Tsfresh is an amazing package that automates feature generation of timeseries data. You can find the processed features dataset here.\n\nAs some features are very computation heavy, I either needed to downsample the 60000 datapoints per segment or process in smaller batches. I choose to go with creating multiple batches per segment. So, for each segment I divided the datapoints into 6 pieces (10,000 datapoints per piece).\n\nCurrently, we have 2854 features for our dataset. At first I processed all the possible features of the train dataset from the INGV competition using ts-fresh library (ComprehensiveFCParameters). This generated almost 8000 features. Then I removed highly correlated columns, and quasi-constant features. This brought our features down to 2854. I also applied a recursive feature elimination to take the top 501 features (500 seemed too goodie-to-shoe of a number). These columns are hard-coded in this notebook.\n\n**Little About Modelling**\nFor making this prediction, I have used LGBMRegressor from the LightGBM (LGBM) framework. I am taking a two-fold approach, where I first use a single LGBM model for the entire dataset. This model is used on the test set to make an initial prediction. Then for multiple LGBM models is created for different ranges of time to eruption. Since these models concentrates on a specific range, they can be more specialized. Then finally, we have six output or prediction for each segment in our test dataset. We take the median of these outputs to get our final prediction.\n\n\n",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1151971": "Hi Kaggle Community,\n\nI wanted to share the work I did for this competition for others so I can have some feedback on what I could have done better. I also wanted to do something for the newcomers in data science, so I cleaned up my competition  and the dataset, added some documentation, and made it available as public. \n\nI hope this also is helpful for the Volcanic Activity Research Community, and the machine learning research community in general. \n\n📖 Notebook - [Predicting 🌋 Eruption Full Tutorial\n](https://www.kaggle.com/ekhtiar/18th-place-predicting-eruption-full-tutorial)\n📈 Processed dataset - [INGV Processed Features With ts-fresh\n](https://www.kaggle.com/ekhtiar/ingv-parquet)\n\n**Little About The Processing...**\nFirst of all, I processed of all I processed a wide number of features for the original dataset using tsfresh. Tsfresh is an amazing package that automates feature generation of timeseries data. You can find the processed features dataset here.\n\nAs some features are very computation heavy, I either needed to downsample the 60000 datapoints per segment or process in smaller batches. I choose to go with creating multiple batches per segment. So, for each segment I divided the datapoints into 6 pieces (10,000 datapoints per piece).\n\nCurrently, we have 2854 features for our dataset. At first I processed all the possible features of the train dataset from the INGV competition using ts-fresh library (ComprehensiveFCParameters). This generated almost 8000 features. Then I removed highly correlated columns, and quasi-constant features. This brought our features down to 2854. I also applied a recursive feature elimination to take the top 501 features (500 seemed too goodie-to-shoe of a number). These columns are hard-coded in this notebook.\n\n**Little About Modelling**\nFor making this prediction, I have used LGBMRegressor from the LightGBM (LGBM) framework. I am taking a two-fold approach, where I first use a single LGBM model for the entire dataset. This model is used on the test set to make an initial prediction. Then for multiple LGBM models is created for different ranges of time to eruption. Since these models concentrates on a specific range, they can be more specialized. Then finally, we have six output or prediction for each segment in our test dataset. We take the median of these outputs to get our final prediction.\n\n\n"
  }
}