{
  "id": 327217,
  "title": "50th place solution",
  "url": "/competitions/birdclef-2022/discussion/327217",
  "author_name": "Node",
  "post_date": "2022-05-26T07:55:57.081000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thank you to every competitor and host that held this competition.<br>\nThis is my solution and what I tried.</p>\n<h2>Things that I watched out</h2>\n<p>As I have posted in <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/321668\" target=\"_blank\">this discussion</a>, I believe that three aspects of this competition needed to be considered.</p>\n<ol>\n<li><p>Get high CV score<br>\nThis is a given, but a better model should be used to improve cross-validation scores.<br>\nI have the impression that many people used the improved PANNS model used in the Notebook published by <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> in the 2021 Bird Competition.  I also used this model and reached CV:0.789 (f1 Score), Public:0.739, and Private:0.711.<br>\nThen, in the middle of the competition, I switched to use another model (that I'll explain below.) and improved to CV:0.805(f1 Score), Public:0.7634, Private:0.7210 (best with a single model).</p></li>\n<li><p>Measures against Domain Shift  <br>\nThere were several models with high CV but low LB. This is probably due to the existence of a domain shift between training data and test data. As a countermeasure, Augmentation, which inserts noise, was effective. In particular, Gaussian Noise and Pink Noise worked effectively.<br>\nModels trained with these augmentations got relatively low CV scores, but the Public scores improved.</p></li>\n<li><p>Selecting the Appropriate Threshold<br>\nIn this competition, since there were no Test Soundscapes (data collected under the same conditions as the test data), the threshold had to be set manually. This appropriate threshold was usually far from the value that was optimal during the Validation phase of the model, so tuning was an important step. Since the Public Score was the only available factor to choose threshold, we had no choice but to overfit the Public Score (despite this, LB did not shake as much as we expected).</p></li>\n</ol>\n<h2>Solution</h2>\n<p><img src=\"https://user-images.githubusercontent.com/69585273/170442235-ef15296d-33dd-475c-80fc-9e3482ca1278.png\" alt=\"solution\"></p>\n<h3>Single model</h3>\n<p>As figure shows, the sounddata was converted into a 20-second segmented mel-spectrogram, and then Augmentation (Normalize, Pink Noise) was added.<br>\nAugmentation (model1: Normalize&amp;PinkNoise, model2: Normalize&amp;GaussianNoise&amp;RandomVolume) was added to the input, and the features were extracted using ResNet101d(model1) and ResNeXt50d_32x4d(model2).<br>\nThe output is a batch size × channel × frequency × time (four-dimensional).<br>\nThen, the features are compressed in the frequency and time directions using GeM pooling, and the result was returned to a two-dimensional tensor of batch size × channel.<br>\nThen fc layer is applied to obtain a 152-dimensional output.  <br>\nI also applied MixUp as augmentation with a probability of 0.25. I believe that the robustness of the model has been improved, albeit only by a small margin.  <br>\nAs a loss function, I used Focal loss based on BCEWithLogitsLoss (<a href=\"https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/213075\" target=\"_blank\">https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/213075</a> ).  <br>\nI also used SAM Optimizer based on Adam as the Optimizer (<a href=\"https://github.com/davda54/sam\" target=\"_blank\">https://github.com/davda54/sam</a> ).<br>\nBy changing the Optimizer to SAM, the Public Score for the Sound Event Detection model improved by about 0.02 when the experiment was conducted under all conditions except for the Optimizer.</p>\n<h3>Ensemble</h3>\n<p>Instead of passing the model-by-model average of the output predictions through the threshold, the logical OR or logical product of the true/false values after passing through the threshold was<br>\nsubmitted as the final prediction result. By applying an ensemble that takes Logical OR, I obtained an output result with a Private score of 0.7292, from the predictions whose private score are 0.7176 and 0.7008 (I we believe they worked effectively)</p>\n<h2>what did not work</h2>\n<ul>\n<li>Use complex models as backbone (e.g. Swin Transformer)  </li>\n<li>PaSST (<a href=\"https://github.com/kkoutini/PaSST\" target=\"_blank\">https://github.com/kkoutini/PaSST</a>)  </li>\n<li>Forcibly apply Multi head attention to the obtained features </li>\n<li>Using the Psuedo label</li>\n</ul>",
  "messages": [
    {
      "id": 1801848,
      "postDate": "2022-05-26T07:55:57.080Z",
      "content": "<p>Thank you to every competitor and host that held this competition.<br>\nThis is my solution and what I tried.</p>\n<h2>Things that I watched out</h2>\n<p>As I have posted in <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/321668\" target=\"_blank\">this discussion</a>, I believe that three aspects of this competition needed to be considered.</p>\n<ol>\n<li><p>Get high CV score<br>\nThis is a given, but a better model should be used to improve cross-validation scores.<br>\nI have the impression that many people used the improved PANNS model used in the Notebook published by <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> in the 2021 Bird Competition.  I also used this model and reached CV:0.789 (f1 Score), Public:0.739, and Private:0.711.<br>\nThen, in the middle of the competition, I switched to use another model (that I'll explain below.) and improved to CV:0.805(f1 Score), Public:0.7634, Private:0.7210 (best with a single model).</p></li>\n<li><p>Measures against Domain Shift  <br>\nThere were several models with high CV but low LB. This is probably due to the existence of a domain shift between training data and test data. As a countermeasure, Augmentation, which inserts noise, was effective. In particular, Gaussian Noise and Pink Noise worked effectively.<br>\nModels trained with these augmentations got relatively low CV scores, but the Public scores improved.</p></li>\n<li><p>Selecting the Appropriate Threshold<br>\nIn this competition, since there were no Test Soundscapes (data collected under the same conditions as the test data), the threshold had to be set manually. This appropriate threshold was usually far from the value that was optimal during the Validation phase of the model, so tuning was an important step. Since the Public Score was the only available factor to choose threshold, we had no choice but to overfit the Public Score (despite this, LB did not shake as much as we expected).</p></li>\n</ol>\n<h2>Solution</h2>\n<p><img src=\"https://user-images.githubusercontent.com/69585273/170442235-ef15296d-33dd-475c-80fc-9e3482ca1278.png\" alt=\"solution\"></p>\n<h3>Single model</h3>\n<p>As figure shows, the sounddata was converted into a 20-second segmented mel-spectrogram, and then Augmentation (Normalize, Pink Noise) was added.<br>\nAugmentation (model1: Normalize&amp;PinkNoise, model2: Normalize&amp;GaussianNoise&amp;RandomVolume) was added to the input, and the features were extracted using ResNet101d(model1) and ResNeXt50d_32x4d(model2).<br>\nThe output is a batch size × channel × frequency × time (four-dimensional).<br>\nThen, the features are compressed in the frequency and time directions using GeM pooling, and the result was returned to a two-dimensional tensor of batch size × channel.<br>\nThen fc layer is applied to obtain a 152-dimensional output.  <br>\nI also applied MixUp as augmentation with a probability of 0.25. I believe that the robustness of the model has been improved, albeit only by a small margin.  <br>\nAs a loss function, I used Focal loss based on BCEWithLogitsLoss (<a href=\"https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/213075\" target=\"_blank\">https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/213075</a> ).  <br>\nI also used SAM Optimizer based on Adam as the Optimizer (<a href=\"https://github.com/davda54/sam\" target=\"_blank\">https://github.com/davda54/sam</a> ).<br>\nBy changing the Optimizer to SAM, the Public Score for the Sound Event Detection model improved by about 0.02 when the experiment was conducted under all conditions except for the Optimizer.</p>\n<h3>Ensemble</h3>\n<p>Instead of passing the model-by-model average of the output predictions through the threshold, the logical OR or logical product of the true/false values after passing through the threshold was<br>\nsubmitted as the final prediction result. By applying an ensemble that takes Logical OR, I obtained an output result with a Private score of 0.7292, from the predictions whose private score are 0.7176 and 0.7008 (I we believe they worked effectively)</p>\n<h2>what did not work</h2>\n<ul>\n<li>Use complex models as backbone (e.g. Swin Transformer)  </li>\n<li>PaSST (<a href=\"https://github.com/kkoutini/PaSST\" target=\"_blank\">https://github.com/kkoutini/PaSST</a>)  </li>\n<li>Forcibly apply Multi head attention to the obtained features </li>\n<li>Using the Psuedo label</li>\n</ul>",
      "rawMarkdown": "Thank you to every competitor and host that held this competition.\nThis is my solution and what I tried.\n\n## Things that I watched out\nAs I have posted in [this discussion](https://www.kaggle.com/competitions/birdclef-2022/discussion/321668), I believe that three aspects of this competition needed to be considered.\n\n1. Get high CV score\nThis is a given, but a better model should be used to improve cross-validation scores.\nI have the impression that many people used the improved PANNS model used in the Notebook published by @hidehisaarai1213 in the 2021 Bird Competition.  I also used this model and reached CV:0.789 (f1 Score), Public:0.739, and Private:0.711.\nThen, in the middle of the competition, I switched to use another model (that I'll explain below.) and improved to CV:0.805(f1 Score), Public:0.7634, Private:0.7210 (best with a single model).\n\n2.  Measures against Domain Shift  \nThere were several models with high CV but low LB. This is probably due to the existence of a domain shift between training data and test data. As a countermeasure, Augmentation, which inserts noise, was effective. In particular, Gaussian Noise and Pink Noise worked effectively.\nModels trained with these augmentations got relatively low CV scores, but the Public scores improved.\n\n3. Selecting the Appropriate Threshold\nIn this competition, since there were no Test Soundscapes (data collected under the same conditions as the test data), the threshold had to be set manually. This appropriate threshold was usually far from the value that was optimal during the Validation phase of the model, so tuning was an important step. Since the Public Score was the only available factor to choose threshold, we had no choice but to overfit the Public Score (despite this, LB did not shake as much as we expected).\n\n## Solution \n![solution](https://user-images.githubusercontent.com/69585273/170442235-ef15296d-33dd-475c-80fc-9e3482ca1278.png)\n\n### Single model\n\nAs figure shows, the sounddata was converted into a 20-second segmented mel-spectrogram, and then Augmentation (Normalize, Pink Noise) was added.\nAugmentation (model1: Normalize&PinkNoise, model2: Normalize&GaussianNoise&RandomVolume) was added to the input, and the features were extracted using ResNet101d(model1) and ResNeXt50d_32x4d(model2).\nThe output is a batch size × channel × frequency × time (four-dimensional).\nThen, the features are compressed in the frequency and time directions using GeM pooling, and the result was returned to a two-dimensional tensor of batch size × channel.\nThen fc layer is applied to obtain a 152-dimensional output.  \nI also applied MixUp as augmentation with a probability of 0.25. I believe that the robustness of the model has been improved, albeit only by a small margin.  \nAs a loss function, I used Focal loss based on BCEWithLogitsLoss (https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/213075 ).  \nI also used SAM Optimizer based on Adam as the Optimizer (https://github.com/davda54/sam ).\nBy changing the Optimizer to SAM, the Public Score for the Sound Event Detection model improved by about 0.02 when the experiment was conducted under all conditions except for the Optimizer.\n\n### Ensemble\nInstead of passing the model-by-model average of the output predictions through the threshold, the logical OR or logical product of the true/false values after passing through the threshold was\nsubmitted as the final prediction result. By applying an ensemble that takes Logical OR, I obtained an output result with a Private score of 0.7292, from the predictions whose private score are 0.7176 and 0.7008 (I we believe they worked effectively)\n\n## what did not work\n* Use complex models as backbone (e.g. Swin Transformer)  \n* PaSST (https://github.com/kkoutini/PaSST)  \n* Forcibly apply Multi head attention to the obtained features \n* Using the Psuedo label\n",
      "votes": 11
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1801848": "Thank you to every competitor and host that held this competition.\nThis is my solution and what I tried.\n\n## Things that I watched out\nAs I have posted in [this discussion](https://www.kaggle.com/competitions/birdclef-2022/discussion/321668), I believe that three aspects of this competition needed to be considered.\n\n1. Get high CV score\nThis is a given, but a better model should be used to improve cross-validation scores.\nI have the impression that many people used the improved PANNS model used in the Notebook published by @hidehisaarai1213 in the 2021 Bird Competition.  I also used this model and reached CV:0.789 (f1 Score), Public:0.739, and Private:0.711.\nThen, in the middle of the competition, I switched to use another model (that I'll explain below.) and improved to CV:0.805(f1 Score), Public:0.7634, Private:0.7210 (best with a single model).\n\n2.  Measures against Domain Shift  \nThere were several models with high CV but low LB. This is probably due to the existence of a domain shift between training data and test data. As a countermeasure, Augmentation, which inserts noise, was effective. In particular, Gaussian Noise and Pink Noise worked effectively.\nModels trained with these augmentations got relatively low CV scores, but the Public scores improved.\n\n3. Selecting the Appropriate Threshold\nIn this competition, since there were no Test Soundscapes (data collected under the same conditions as the test data), the threshold had to be set manually. This appropriate threshold was usually far from the value that was optimal during the Validation phase of the model, so tuning was an important step. Since the Public Score was the only available factor to choose threshold, we had no choice but to overfit the Public Score (despite this, LB did not shake as much as we expected).\n\n## Solution \n![solution](https://user-images.githubusercontent.com/69585273/170442235-ef15296d-33dd-475c-80fc-9e3482ca1278.png)\n\n### Single model\n\nAs figure shows, the sounddata was converted into a 20-second segmented mel-spectrogram, and then Augmentation (Normalize, Pink Noise) was added.\nAugmentation (model1: Normalize&PinkNoise, model2: Normalize&GaussianNoise&RandomVolume) was added to the input, and the features were extracted using ResNet101d(model1) and ResNeXt50d_32x4d(model2).\nThe output is a batch size × channel × frequency × time (four-dimensional).\nThen, the features are compressed in the frequency and time directions using GeM pooling, and the result was returned to a two-dimensional tensor of batch size × channel.\nThen fc layer is applied to obtain a 152-dimensional output.  \nI also applied MixUp as augmentation with a probability of 0.25. I believe that the robustness of the model has been improved, albeit only by a small margin.  \nAs a loss function, I used Focal loss based on BCEWithLogitsLoss (https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/213075 ).  \nI also used SAM Optimizer based on Adam as the Optimizer (https://github.com/davda54/sam ).\nBy changing the Optimizer to SAM, the Public Score for the Sound Event Detection model improved by about 0.02 when the experiment was conducted under all conditions except for the Optimizer.\n\n### Ensemble\nInstead of passing the model-by-model average of the output predictions through the threshold, the logical OR or logical product of the true/false values after passing through the threshold was\nsubmitted as the final prediction result. By applying an ensemble that takes Logical OR, I obtained an output result with a Private score of 0.7292, from the predictions whose private score are 0.7176 and 0.7008 (I we believe they worked effectively)\n\n## what did not work\n* Use complex models as backbone (e.g. Swin Transformer)  \n* PaSST (https://github.com/kkoutini/PaSST)  \n* Forcibly apply Multi head attention to the obtained features \n* Using the Psuedo label\n"
  }
}