{
  "id": 211315,
  "title": "[1st Place] Mel Spectrogram + Blended ResNets 🌋🔥",
  "url": "/competitions/predict-volcanic-eruptions-ingv-oe/discussion/211315",
  "author_name": "Jie Feng",
  "post_date": "2021-01-14T15:48:17.038000",
  "votes": 16,
  "comment_count": 8,
  "views": 0,
  "content": "<p>First of all, thanks to INGV and Kaggle for hosting this competition!<br>\nEven though it isn’t ranked, we had a lot of fun working with this dataset.</p>\n<p><strong>1. Spectrograms</strong></p>\n<ul>\n<li>We used librosa mel spectrograms to generate 256x256 spectrograms</li>\n<li><a href=\"https://www.kaggle.com/guimingjiang/ingv-process-data-into-mel-spectrograms-256x256\" target=\"_blank\">Notebook Link</a></li>\n<li>Lower frequencies are given more space in the spectrogram</li>\n<li>256x256 images were fed into CNNs: ResNext, SEResNet, ResNeSt and their outputs were ensembled</li>\n<li>EfficientNets did not perform as well as ResNets in this data</li>\n<li>Validation MAE for the well performing ones is less than 1e6</li>\n<li>ResNet LB scores: 4100000 - 4700000<ul>\n<li>ResNeXt101 : 4121853</li>\n<li>ResNeSt269e: 4357483</li>\n<li>SeResNet152: 4737569 </li>\n<li>Other ResNet models are also in this range</li></ul></li>\n</ul>\n<p><strong>2. Tabular datasets</strong></p>\n<ul>\n<li>Thanks to <a href=\"https://www.kaggle.com/carpediemamigo\" target=\"_blank\">@carpediemamigo</a> and <a href=\"https://www.kaggle.com/josemori\" target=\"_blank\">@josemori</a> for their wonderful notebooks</li>\n<li>The inference from the tsfresh dataset and the tree model dataset from their notebooks were used in the final ensemble</li>\n<li>We were not able to make Neural Networks e.g. TabNet/MLP work well on this dataset.</li>\n</ul>\n<p><strong>3. Final Ensemble</strong></p>\n<ul>\n<li>Our final model was formed using a <strong>weighted blend</strong> of the following models:<ul>\n<li>ResNeXt101 ~ 50%</li>\n<li>ResNest269e ~ 20%</li>\n<li>Other ResNet-type models ~ 20%</li>\n<li>Tree Models ~ 10%</li></ul></li>\n<li>We have found that ensembling sufficient models, even based on the same data processing provides sufficient stability in results (not much significant difference in LB scores by exchanging 1 or 2 models)</li>\n<li>At the same time, ensembling sufficient models allows for good private LB stability.</li>\n</ul>\n<p>Thank you for reading! Hope this has been helpful.</p>",
  "messages": [
    {
      "id": 1153062,
      "postDate": "2021-01-14T15:48:17.037Z",
      "content": "<p>First of all, thanks to INGV and Kaggle for hosting this competition!<br>\nEven though it isn’t ranked, we had a lot of fun working with this dataset.</p>\n<p><strong>1. Spectrograms</strong></p>\n<ul>\n<li>We used librosa mel spectrograms to generate 256x256 spectrograms</li>\n<li><a href=\"https://www.kaggle.com/guimingjiang/ingv-process-data-into-mel-spectrograms-256x256\" target=\"_blank\">Notebook Link</a></li>\n<li>Lower frequencies are given more space in the spectrogram</li>\n<li>256x256 images were fed into CNNs: ResNext, SEResNet, ResNeSt and their outputs were ensembled</li>\n<li>EfficientNets did not perform as well as ResNets in this data</li>\n<li>Validation MAE for the well performing ones is less than 1e6</li>\n<li>ResNet LB scores: 4100000 - 4700000<ul>\n<li>ResNeXt101 : 4121853</li>\n<li>ResNeSt269e: 4357483</li>\n<li>SeResNet152: 4737569 </li>\n<li>Other ResNet models are also in this range</li></ul></li>\n</ul>\n<p><strong>2. Tabular datasets</strong></p>\n<ul>\n<li>Thanks to <a href=\"https://www.kaggle.com/carpediemamigo\" target=\"_blank\">@carpediemamigo</a> and <a href=\"https://www.kaggle.com/josemori\" target=\"_blank\">@josemori</a> for their wonderful notebooks</li>\n<li>The inference from the tsfresh dataset and the tree model dataset from their notebooks were used in the final ensemble</li>\n<li>We were not able to make Neural Networks e.g. TabNet/MLP work well on this dataset.</li>\n</ul>\n<p><strong>3. Final Ensemble</strong></p>\n<ul>\n<li>Our final model was formed using a <strong>weighted blend</strong> of the following models:<ul>\n<li>ResNeXt101 ~ 50%</li>\n<li>ResNest269e ~ 20%</li>\n<li>Other ResNet-type models ~ 20%</li>\n<li>Tree Models ~ 10%</li></ul></li>\n<li>We have found that ensembling sufficient models, even based on the same data processing provides sufficient stability in results (not much significant difference in LB scores by exchanging 1 or 2 models)</li>\n<li>At the same time, ensembling sufficient models allows for good private LB stability.</li>\n</ul>\n<p>Thank you for reading! Hope this has been helpful.</p>",
      "rawMarkdown": "\nFirst of all, thanks to INGV and Kaggle for hosting this competition!\nEven though it isn’t ranked, we had a lot of fun working with this dataset.\n  \n**1. Spectrograms**\n- We used librosa mel spectrograms to generate 256x256 spectrograms\n- [Notebook Link](https://www.kaggle.com/guimingjiang/ingv-process-data-into-mel-spectrograms-256x256)\n- Lower frequencies are given more space in the spectrogram\n- 256x256 images were fed into CNNs: ResNext, SEResNet, ResNeSt and their outputs were ensembled\n- EfficientNets did not perform as well as ResNets in this data\n- Validation MAE for the well performing ones is less than 1e6\n- ResNet LB scores: 4100000 - 4700000\n    - ResNeXt101 : 4121853\n    - ResNeSt269e: 4357483\n    - SeResNet152: 4737569 \n    - Other ResNet models are also in this range\n  \n**2. Tabular datasets**\n- Thanks to @carpediemamigo and @josemori for their wonderful notebooks\n- The inference from the tsfresh dataset and the tree model dataset from their notebooks were used in the final ensemble\n- We were not able to make Neural Networks e.g. TabNet/MLP work well on this dataset.\n  \n**3. Final Ensemble**\n- Our final model was formed using a **weighted blend** of the following models:\n    - ResNeXt101 ~ 50%\n    - ResNest269e ~ 20%\n    - Other ResNet-type models ~ 20%\n    - Tree Models ~ 10%\n- We have found that ensembling sufficient models, even based on the same data processing provides sufficient stability in results (not much significant difference in LB scores by exchanging 1 or 2 models)\n- At the same time, ensembling sufficient models allows for good private LB stability.\n  \nThank you for reading! Hope this has been helpful.",
      "votes": 16
    },
    {
      "id": 1175270,
      "postDate": "2021-01-29T03:54:53.937Z",
      "content": "<p>Very cool guys and very well done. How much time did the spectrogram generation take to run on both the train and test data? </p>",
      "rawMarkdown": "Very cool guys and very well done. How much time did the spectrogram generation take to run on both the train and test data? ",
      "votes": 1,
      "replies": [
        {
          "id": 1176333,
          "postDate": "2021-01-29T15:27:26.123Z",
          "content": "<p>Thanks! As for the spectrogram, the generation took an hour to run.</p>",
          "rawMarkdown": "Thanks! As for the spectrogram, the generation took an hour to run.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1157671,
      "postDate": "2021-01-18T04:54:10.633Z",
      "content": "<p>great work on the first place. i thought this was an interesting competition and data set and will learn from your solution.</p>",
      "rawMarkdown": "great work on the first place. i thought this was an interesting competition and data set and will learn from your solution.",
      "votes": 1
    },
    {
      "id": 1153106,
      "postDate": "2021-01-14T16:09:47.357Z",
      "content": "<p>Thanks INGV and Kaggle for this contest, and huge thanks to my teammate <a href=\"https://www.kaggle.com/jiefeng98\" target=\"_blank\">@jiefeng98</a> on working on this contest together</p>\n<p>Here's some other ideas we hoped to explore but didn't have the time/opportunity to:  </p>\n<p><strong>1. Training Augmentation</strong></p>\n<ul>\n<li>Since some seismic sensors for some datasets are NA values, we can synthetically expand the dataset by blanking out individual columns (sensors) at test time and effectively get ~10x training data</li>\n</ul>\n<p><strong>2. Test-Time Augmentation</strong></p>\n<ul>\n<li>Similar to the above, we can run test-time augmentation by running each train dataset through the model with individual columns blanked out</li>\n</ul>\n<p>Our initial experiments didn't really pan out for either but it could be a matter of further hyperparameter optimisation/methodology (e.g. weighted blend of TTA results)</p>\n<p>If anyone has managed to make this work, we hope you share your code!</p>",
      "rawMarkdown": "Thanks INGV and Kaggle for this contest, and huge thanks to my teammate @jiefeng98 on working on this contest together\n\nHere's some other ideas we hoped to explore but didn't have the time/opportunity to:  \n  \n**1. Training Augmentation**\n- Since some seismic sensors for some datasets are NA values, we can synthetically expand the dataset by blanking out individual columns (sensors) at test time and effectively get ~10x training data\n  \n**2. Test-Time Augmentation**\n- Similar to the above, we can run test-time augmentation by running each train dataset through the model with individual columns blanked out\n\nOur initial experiments didn't really pan out for either but it could be a matter of further hyperparameter optimisation/methodology (e.g. weighted blend of TTA results)\n\nIf anyone has managed to make this work, we hope you share your code!",
      "votes": 2
    },
    {
      "id": 1230701,
      "postDate": "2021-03-08T11:24:01.303Z",
      "content": "<p>Great work!<br>\nI got the the best result combining channels 4,5,&amp; 6 to one RGB image using mel-spectrograms and Efnet B5. </p>",
      "rawMarkdown": "Great work!\nI got the the best result combining channels 4,5,& 6 to one RGB image using mel-spectrograms and Efnet B5. "
    },
    {
      "id": 1228095,
      "postDate": "2021-03-06T05:33:46.107Z",
      "content": "<p>Thanks for the solution! Could you share the hyperparameters of the ResNet training if you don't mind, I am trying to reproduce your result but I can't get such low validation MAE like you.</p>",
      "rawMarkdown": "Thanks for the solution! Could you share the hyperparameters of the ResNet training if you don't mind, I am trying to reproduce your result but I can't get such low validation MAE like you."
    },
    {
      "id": 1181378,
      "postDate": "2021-02-01T20:50:38.490Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1181818,
          "postDate": "2021-02-02T06:15:33.270Z",
          "content": "<p>Thanks! </p>\n<p>You're right - the main technique here is processing the initial time-series tabular data (the sensor values) into a spectrogram. Afterwards, we can treat this as a regression problem using image data as an input, and then we can fine-tune strong pre-trained ResNet models.</p>\n<p>In this case, we convert the 10 column table (10 sensors) into a 10-channel image, where each channel represents the spectrogram of that sensor. During inference, we will convert the input table into 10 spectrograms. Most ResNets are trained on 3-channel images, so small modifications are needed to the model.</p>",
          "rawMarkdown": "Thanks! \n\nYou're right - the main technique here is processing the initial time-series tabular data (the sensor values) into a spectrogram. Afterwards, we can treat this as a regression problem using image data as an input, and then we can fine-tune strong pre-trained ResNet models.\n\nIn this case, we convert the 10 column table (10 sensors) into a 10-channel image, where each channel represents the spectrogram of that sensor. During inference, we will convert the input table into 10 spectrograms. Most ResNets are trained on 3-channel images, so small modifications are needed to the model.\n",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1175270,
      "author_name": "Matt S.",
      "author_url": "",
      "post_date": "2021-01-29T03:54:53.937000",
      "content": "<p>Very cool guys and very well done. How much time did the spectrogram generation take to run on both the train and test data? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1176333,
          "author_name": "Jie Feng",
          "author_url": "",
          "post_date": "2021-01-29T15:27:26.123000",
          "content": "<p>Thanks! As for the spectrogram, the generation took an hour to run.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1157671,
      "author_name": "Dave E",
      "author_url": "",
      "post_date": "2021-01-18T04:54:10.633000",
      "content": "<p>great work on the first place. i thought this was an interesting competition and data set and will learn from your solution.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1153106,
      "author_name": "Gui Ming Jiang",
      "author_url": "",
      "post_date": "2021-01-14T16:09:47.357000",
      "content": "<p>Thanks INGV and Kaggle for this contest, and huge thanks to my teammate <a href=\"https://www.kaggle.com/jiefeng98\" target=\"_blank\">@jiefeng98</a> on working on this contest together</p>\n<p>Here's some other ideas we hoped to explore but didn't have the time/opportunity to:  </p>\n<p><strong>1. Training Augmentation</strong></p>\n<ul>\n<li>Since some seismic sensors for some datasets are NA values, we can synthetically expand the dataset by blanking out individual columns (sensors) at test time and effectively get ~10x training data</li>\n</ul>\n<p><strong>2. Test-Time Augmentation</strong></p>\n<ul>\n<li>Similar to the above, we can run test-time augmentation by running each train dataset through the model with individual columns blanked out</li>\n</ul>\n<p>Our initial experiments didn't really pan out for either but it could be a matter of further hyperparameter optimisation/methodology (e.g. weighted blend of TTA results)</p>\n<p>If anyone has managed to make this work, we hope you share your code!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1230701,
      "author_name": "Michael Markzon",
      "author_url": "",
      "post_date": "2021-03-08T11:24:01.303000",
      "content": "<p>Great work!<br>\nI got the the best result combining channels 4,5,&amp; 6 to one RGB image using mel-spectrograms and Efnet B5. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1228095,
      "author_name": "jojo",
      "author_url": "",
      "post_date": "2021-03-06T05:33:46.107000",
      "content": "<p>Thanks for the solution! Could you share the hyperparameters of the ResNet training if you don't mind, I am trying to reproduce your result but I can't get such low validation MAE like you.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1181378,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-01T20:50:38.490000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1181818,
          "author_name": "Gui Ming Jiang",
          "author_url": "",
          "post_date": "2021-02-02T06:15:33.270000",
          "content": "<p>Thanks! </p>\n<p>You're right - the main technique here is processing the initial time-series tabular data (the sensor values) into a spectrogram. Afterwards, we can treat this as a regression problem using image data as an input, and then we can fine-tune strong pre-trained ResNet models.</p>\n<p>In this case, we convert the 10 column table (10 sensors) into a 10-channel image, where each channel represents the spectrogram of that sensor. During inference, we will convert the input table into 10 spectrograms. Most ResNets are trained on 3-channel images, so small modifications are needed to the model.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1153062": "\nFirst of all, thanks to INGV and Kaggle for hosting this competition!\nEven though it isn’t ranked, we had a lot of fun working with this dataset.\n  \n**1. Spectrograms**\n- We used librosa mel spectrograms to generate 256x256 spectrograms\n- [Notebook Link](https://www.kaggle.com/guimingjiang/ingv-process-data-into-mel-spectrograms-256x256)\n- Lower frequencies are given more space in the spectrogram\n- 256x256 images were fed into CNNs: ResNext, SEResNet, ResNeSt and their outputs were ensembled\n- EfficientNets did not perform as well as ResNets in this data\n- Validation MAE for the well performing ones is less than 1e6\n- ResNet LB scores: 4100000 - 4700000\n    - ResNeXt101 : 4121853\n    - ResNeSt269e: 4357483\n    - SeResNet152: 4737569 \n    - Other ResNet models are also in this range\n  \n**2. Tabular datasets**\n- Thanks to @carpediemamigo and @josemori for their wonderful notebooks\n- The inference from the tsfresh dataset and the tree model dataset from their notebooks were used in the final ensemble\n- We were not able to make Neural Networks e.g. TabNet/MLP work well on this dataset.\n  \n**3. Final Ensemble**\n- Our final model was formed using a **weighted blend** of the following models:\n    - ResNeXt101 ~ 50%\n    - ResNest269e ~ 20%\n    - Other ResNet-type models ~ 20%\n    - Tree Models ~ 10%\n- We have found that ensembling sufficient models, even based on the same data processing provides sufficient stability in results (not much significant difference in LB scores by exchanging 1 or 2 models)\n- At the same time, ensembling sufficient models allows for good private LB stability.\n  \nThank you for reading! Hope this has been helpful.",
    "1175270": "Very cool guys and very well done. How much time did the spectrogram generation take to run on both the train and test data? ",
    "1157671": "great work on the first place. i thought this was an interesting competition and data set and will learn from your solution.",
    "1153106": "Thanks INGV and Kaggle for this contest, and huge thanks to my teammate @jiefeng98 on working on this contest together\n\nHere's some other ideas we hoped to explore but didn't have the time/opportunity to:  \n  \n**1. Training Augmentation**\n- Since some seismic sensors for some datasets are NA values, we can synthetically expand the dataset by blanking out individual columns (sensors) at test time and effectively get ~10x training data\n  \n**2. Test-Time Augmentation**\n- Similar to the above, we can run test-time augmentation by running each train dataset through the model with individual columns blanked out\n\nOur initial experiments didn't really pan out for either but it could be a matter of further hyperparameter optimisation/methodology (e.g. weighted blend of TTA results)\n\nIf anyone has managed to make this work, we hope you share your code!",
    "1230701": "Great work!\nI got the the best result combining channels 4,5,& 6 to one RGB image using mel-spectrograms and Efnet B5. ",
    "1228095": "Thanks for the solution! Could you share the hyperparameters of the ResNet training if you don't mind, I am trying to reproduce your result but I can't get such low validation MAE like you.",
    "1181378": ""
  }
}