{
  "id": 220753,
  "title": "33rd place solution - SED model",
  "url": "/competitions/rfcx-species-audio-detection/discussion/220753",
  "author_name": "yuki",
  "post_date": "2021-02-19T12:03:17.082000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Congrats all the winners.<br>\nI would like to thank the organizers for hosting a fun competition.</p>\n<p>Our team's solution is an ensemble of seven models.<br>\nThe ensemble includes three SED models and four not-SED models.<br>\nI will discuss the single model with the best score among them. (Public: 0.919/Private: 0.927)</p>\n<h3>Model</h3>\n<p>My model is based on the SED model described in <a href=\"https://www.kaggle.com/shinmura0\" target=\"_blank\">@shinmura0</a>'s discussion.<br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007</a></p>\n<ul>\n<li>Feature Extractor:EfficientNet-B3</li>\n<li>Loss Function:BCELoss</li>\n<li>Optimizer:SGD</li>\n<li>Scheduler:CosineAnnealingLR</li>\n<li>LR:0.15</li>\n<li>Data Augmentation:denoise</li>\n<li>CV: 4Fold multilabel-stratifiedkfold</li>\n</ul>\n<p>My model is characterized by a very large learning rate, and as I reduced the learning rate from 0.15, the accuracy decreased.<br>\nThis is contrary to my experience.</p>\n<p>I thought that data expansion would be very effective, and I tried various data expansions, but most of them did not work.</p>\n<p>The only data enhancement that worked was the denoising that <a href=\"https://www.kaggle.com/takamichitoda\" target=\"_blank\">@takamichitoda</a> introduced in the discussion.<br>\nIn the discussion, he de-noises all the data, but it worked when extended with p=0.1.<br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/214019\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/214019</a></p>\n<h3>Training data</h3>\n<p>The training data was randomly cropped from a meruspectrogram at 60/9 seconds.<br>\nI experimented with crop sizes ranging from 60/2 seconds to 60/20 seconds, and found that 60/9 and 60/10 gave good results.<br>\nMost of the discussions used a 60/10 second crop, but I think a smaller size would have reduced the probability of including missing labels.</p>\n<h3>Validation/Test data</h3>\n<p>For the validation and test data, I used the same 60/9 seconds as the training date.</p>\n<p>I create a total of 17 images by sliding the size of 60/9 seconds by half.<br>\nMake a prediction for each of the 17 images, and use the max value of the probability of occurrence of each label as the prediction value.</p>\n<p><img src=\"https://i.imgur.com/XBc6x3R.png\" alt=\"\"></p>\n<p>The validation was very time consuming, as we needed to infer 17 images to validate 1 clip.</p>\n<h3>Didn't work.</h3>\n<p>Pseudo labels (train/test, soft label/hard label)<br>\nPrediction model for class3<br>\n　-&gt; The accuracy of 3 was extremely low, so I tried to create a model specific to 3, but it didn't work.</p>",
  "messages": [
    {
      "id": 1210399,
      "postDate": "2021-02-19T12:03:17.083Z",
      "content": "<p>Congrats all the winners.<br>\nI would like to thank the organizers for hosting a fun competition.</p>\n<p>Our team's solution is an ensemble of seven models.<br>\nThe ensemble includes three SED models and four not-SED models.<br>\nI will discuss the single model with the best score among them. (Public: 0.919/Private: 0.927)</p>\n<h3>Model</h3>\n<p>My model is based on the SED model described in <a href=\"https://www.kaggle.com/shinmura0\" target=\"_blank\">@shinmura0</a>'s discussion.<br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007</a></p>\n<ul>\n<li>Feature Extractor:EfficientNet-B3</li>\n<li>Loss Function:BCELoss</li>\n<li>Optimizer:SGD</li>\n<li>Scheduler:CosineAnnealingLR</li>\n<li>LR:0.15</li>\n<li>Data Augmentation:denoise</li>\n<li>CV: 4Fold multilabel-stratifiedkfold</li>\n</ul>\n<p>My model is characterized by a very large learning rate, and as I reduced the learning rate from 0.15, the accuracy decreased.<br>\nThis is contrary to my experience.</p>\n<p>I thought that data expansion would be very effective, and I tried various data expansions, but most of them did not work.</p>\n<p>The only data enhancement that worked was the denoising that <a href=\"https://www.kaggle.com/takamichitoda\" target=\"_blank\">@takamichitoda</a> introduced in the discussion.<br>\nIn the discussion, he de-noises all the data, but it worked when extended with p=0.1.<br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/214019\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/214019</a></p>\n<h3>Training data</h3>\n<p>The training data was randomly cropped from a meruspectrogram at 60/9 seconds.<br>\nI experimented with crop sizes ranging from 60/2 seconds to 60/20 seconds, and found that 60/9 and 60/10 gave good results.<br>\nMost of the discussions used a 60/10 second crop, but I think a smaller size would have reduced the probability of including missing labels.</p>\n<h3>Validation/Test data</h3>\n<p>For the validation and test data, I used the same 60/9 seconds as the training date.</p>\n<p>I create a total of 17 images by sliding the size of 60/9 seconds by half.<br>\nMake a prediction for each of the 17 images, and use the max value of the probability of occurrence of each label as the prediction value.</p>\n<p><img src=\"https://i.imgur.com/XBc6x3R.png\" alt=\"\"></p>\n<p>The validation was very time consuming, as we needed to infer 17 images to validate 1 clip.</p>\n<h3>Didn't work.</h3>\n<p>Pseudo labels (train/test, soft label/hard label)<br>\nPrediction model for class3<br>\n　-&gt; The accuracy of 3 was extremely low, so I tried to create a model specific to 3, but it didn't work.</p>",
      "rawMarkdown": "Congrats all the winners.\nI would like to thank the organizers for hosting a fun competition.\n\nOur team's solution is an ensemble of seven models.\nThe ensemble includes three SED models and four not-SED models.\nI will discuss the single model with the best score among them. (Public: 0.919/Private: 0.927)\n\n### Model\nMy model is based on the SED model described in @shinmura0's discussion.\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007\n\n* Feature Extractor:EfficientNet-B3\n* Loss Function:BCELoss\n* Optimizer:SGD\n* Scheduler:CosineAnnealingLR\n* LR:0.15\n* Data Augmentation:denoise\n* CV: 4Fold multilabel-stratifiedkfold\n\nMy model is characterized by a very large learning rate, and as I reduced the learning rate from 0.15, the accuracy decreased.\nThis is contrary to my experience.\n\nI thought that data expansion would be very effective, and I tried various data expansions, but most of them did not work.\n\nThe only data enhancement that worked was the denoising that @takamichitoda introduced in the discussion.\nIn the discussion, he de-noises all the data, but it worked when extended with p=0.1.\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/214019\n\n### Training data\nThe training data was randomly cropped from a meruspectrogram at 60/9 seconds.\nI experimented with crop sizes ranging from 60/2 seconds to 60/20 seconds, and found that 60/9 and 60/10 gave good results.\nMost of the discussions used a 60/10 second crop, but I think a smaller size would have reduced the probability of including missing labels.\n\n### Validation/Test data\nFor the validation and test data, I used the same 60/9 seconds as the training date.\n\nI create a total of 17 images by sliding the size of 60/9 seconds by half.\nMake a prediction for each of the 17 images, and use the max value of the probability of occurrence of each label as the prediction value.\n\n![](https://i.imgur.com/XBc6x3R.png)\n\nThe validation was very time consuming, as we needed to infer 17 images to validate 1 clip.\n\n### Didn't work.\nPseudo labels (train/test, soft label/hard label)\nPrediction model for class3\n　-> The accuracy of 3 was extremely low, so I tried to create a model specific to 3, but it didn't work.",
      "votes": 11
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1210399": "Congrats all the winners.\nI would like to thank the organizers for hosting a fun competition.\n\nOur team's solution is an ensemble of seven models.\nThe ensemble includes three SED models and four not-SED models.\nI will discuss the single model with the best score among them. (Public: 0.919/Private: 0.927)\n\n### Model\nMy model is based on the SED model described in @shinmura0's discussion.\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007\n\n* Feature Extractor:EfficientNet-B3\n* Loss Function:BCELoss\n* Optimizer:SGD\n* Scheduler:CosineAnnealingLR\n* LR:0.15\n* Data Augmentation:denoise\n* CV: 4Fold multilabel-stratifiedkfold\n\nMy model is characterized by a very large learning rate, and as I reduced the learning rate from 0.15, the accuracy decreased.\nThis is contrary to my experience.\n\nI thought that data expansion would be very effective, and I tried various data expansions, but most of them did not work.\n\nThe only data enhancement that worked was the denoising that @takamichitoda introduced in the discussion.\nIn the discussion, he de-noises all the data, but it worked when extended with p=0.1.\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/214019\n\n### Training data\nThe training data was randomly cropped from a meruspectrogram at 60/9 seconds.\nI experimented with crop sizes ranging from 60/2 seconds to 60/20 seconds, and found that 60/9 and 60/10 gave good results.\nMost of the discussions used a 60/10 second crop, but I think a smaller size would have reduced the probability of including missing labels.\n\n### Validation/Test data\nFor the validation and test data, I used the same 60/9 seconds as the training date.\n\nI create a total of 17 images by sliding the size of 60/9 seconds by half.\nMake a prediction for each of the 17 images, and use the max value of the probability of occurrence of each label as the prediction value.\n\n![](https://i.imgur.com/XBc6x3R.png)\n\nThe validation was very time consuming, as we needed to infer 17 images to validate 1 clip.\n\n### Didn't work.\nPseudo labels (train/test, soft label/hard label)\nPrediction model for class3\n　-> The accuracy of 3 was extremely low, so I tried to create a model specific to 3, but it didn't work."
  }
}