{
  "id": 413209,
  "title": "14-th place solution",
  "url": "/competitions/birdclef-2023/discussion/413209",
  "author_name": "",
  "post_date": "2023-05-27T13:46:22.186541800Z",
  "votes": 14,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I’ve started this competition only a month before the end,  and I am very happy to finish with silver and 14-th place 😊  <br>\nFirst of all, thanks to the Kaggle Team and Cornell Lab of Ornithology for hosting this competition, to all participants, it was a very exciting experience. </p>\n<h2>My solution</h2>\n<p>Since this year  time for inference was limited to only 120 minutes on CPU,  it didn't make sense to use large models or a large image size. Also previous competitions results confirm this.  <br>\nI’ve tested different architectures and image sizes, and finally I’ve chosen size (1, 144, 244) for a 5-sec chunk. This size allowed me to ensemble up to 6 models, including SED. Hope, next year will be an opportunity to try ONNX and openvino for inference.</p>\n<h3>The solution includes:</h3>\n<p><strong>Prediction</strong> on MelSpectrogram with f_min=100 and f_max=15000.<br>\nFrequencies were chosen based on species in this specific competition. I've checked that for all of birds the voice lies within this interval.<br>\n<strong>Augmentation:</strong></p>\n<ul>\n<li>random mix-up with other records and background noise (“nocall”),</li>\n<li>pink noise, gaussian noise,</li>\n<li>random volume adjustment,</li>\n<li>random fade in, fade out,</li>\n<li>random low and high pass filters.  </li>\n</ul>\n<p><strong>Detection</strong> regions of interest (ROI) with librosa.effects.split():</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5732512%2F4186c584d604d24f63b5a16e72143fcc%2Ffull_mel.jpg?generation=1685192784649095&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5732512%2Fa367d2572733257ddc43652a0d5ce654%2Ffull_mel_with_mask.jpg?generation=1685192822856364&amp;alt=media\" alt=\"\"></p>\n<p><strong><em>!!! top_db</em></strong> parametert must be chosen very carefully<br>\nAfter extension and concatenation of split results, for part of the training I’ve used chunks whose center lies within the selected intervals. Not for all records this worked well, but for the most, that I've checked, the results were fine.   </p>\n<p><strong>Models:</strong></p>\n<ul>\n<li>EfficientNetV2-S with geographical coordinates, converted to <a href=\"https://www.kaggle.com/code/nataliayurasova/birdclef23-wgs-84-vs-cartesian-system\" target=\"_blank\">cartesian coordinate system</a> as additional input,</li>\n<li>SED-architecture  with  EfficientNetV2-S as backbone, with image size 1x144x244 per 5-sec audio chunk,<br>\nSED models were trained on 15-25 seconds clips.  </li>\n</ul>\n<p><strong>Loss:</strong><br>\nAs a loss function I’ve used the weighted sum of BCEWithLogitsLoss() for species and bird’s families prediction.</p>\n<p><strong>Additional data</strong> </p>\n<ul>\n<li>Soundscapes and  labeled tests <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/394358#2179605\" target=\"_blank\">from previous</a> competitions (I've extracted and used only parts without birds as \"nocall').</li>\n<li>Records from  <a href=\"https://xeno-canto.org/\" target=\"_blank\">https://xeno-canto.org/</a> .</li>\n</ul>\n<p><strong>The final ensemble</strong> includes 6 best models, probabilities were averaged with no pre- or post-processing.   </p>\n<p><strong>What doesn’t work:</strong></p>\n<ul>\n<li>Use of secondary labels.</li>\n<li>EfficientNetB0 performed worse even with bigger image size.</li>\n<li>Training on all data (starting from 5 sec) and validation on the first 5 seconds.  </li>\n<li>Increasing the probabilities if bird occurres more than once in soundscape (different thresholds and coefficients).</li>\n<li>Increasing the probabilities of the same bird in the nearest chunks (different thresholds and coefficients).</li>\n</ul>\n<p>Hope to see another BirdCLEF competition and you all next year!</p>",
  "messages": [
    {
      "id": "2277147",
      "postDate": "05/27/2023 13:46:22",
      "content": "<p>I’ve started this competition only a month before the end,  and I am very happy to finish with silver and 14-th place 😊  <br>\nFirst of all, thanks to the Kaggle Team and Cornell Lab of Ornithology for hosting this competition, to all participants, it was a very exciting experience. </p>\n<h2>My solution</h2>\n<p>Since this year  time for inference was limited to only 120 minutes on CPU,  it didn't make sense to use large models or a large image size. Also previous competitions results confirm this.  <br>\nI’ve tested different architectures and image sizes, and finally I’ve chosen size (1, 144, 244) for a 5-sec chunk. This size allowed me to ensemble up to 6 models, including SED. Hope, next year will be an opportunity to try ONNX and openvino for inference.</p>\n<h3>The solution includes:</h3>\n<p><strong>Prediction</strong> on MelSpectrogram with f_min=100 and f_max=15000.<br>\nFrequencies were chosen based on species in this specific competition. I've checked that for all of birds the voice lies within this interval.<br>\n<strong>Augmentation:</strong></p>\n<ul>\n<li>random mix-up with other records and background noise (“nocall”),</li>\n<li>pink noise, gaussian noise,</li>\n<li>random volume adjustment,</li>\n<li>random fade in, fade out,</li>\n<li>random low and high pass filters.  </li>\n</ul>\n<p><strong>Detection</strong> regions of interest (ROI) with librosa.effects.split():</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5732512%2F4186c584d604d24f63b5a16e72143fcc%2Ffull_mel.jpg?generation=1685192784649095&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5732512%2Fa367d2572733257ddc43652a0d5ce654%2Ffull_mel_with_mask.jpg?generation=1685192822856364&amp;alt=media\" alt=\"\"></p>\n<p><strong><em>!!! top_db</em></strong> parametert must be chosen very carefully<br>\nAfter extension and concatenation of split results, for part of the training I’ve used chunks whose center lies within the selected intervals. Not for all records this worked well, but for the most, that I've checked, the results were fine.   </p>\n<p><strong>Models:</strong></p>\n<ul>\n<li>EfficientNetV2-S with geographical coordinates, converted to <a href=\"https://www.kaggle.com/code/nataliayurasova/birdclef23-wgs-84-vs-cartesian-system\" target=\"_blank\">cartesian coordinate system</a> as additional input,</li>\n<li>SED-architecture  with  EfficientNetV2-S as backbone, with image size 1x144x244 per 5-sec audio chunk,<br>\nSED models were trained on 15-25 seconds clips.  </li>\n</ul>\n<p><strong>Loss:</strong><br>\nAs a loss function I’ve used the weighted sum of BCEWithLogitsLoss() for species and bird’s families prediction.</p>\n<p><strong>Additional data</strong> </p>\n<ul>\n<li>Soundscapes and  labeled tests <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/394358#2179605\" target=\"_blank\">from previous</a> competitions (I've extracted and used only parts without birds as \"nocall').</li>\n<li>Records from  <a href=\"https://xeno-canto.org/\" target=\"_blank\">https://xeno-canto.org/</a> .</li>\n</ul>\n<p><strong>The final ensemble</strong> includes 6 best models, probabilities were averaged with no pre- or post-processing.   </p>\n<p><strong>What doesn’t work:</strong></p>\n<ul>\n<li>Use of secondary labels.</li>\n<li>EfficientNetB0 performed worse even with bigger image size.</li>\n<li>Training on all data (starting from 5 sec) and validation on the first 5 seconds.  </li>\n<li>Increasing the probabilities if bird occurres more than once in soundscape (different thresholds and coefficients).</li>\n<li>Increasing the probabilities of the same bird in the nearest chunks (different thresholds and coefficients).</li>\n</ul>\n<p>Hope to see another BirdCLEF competition and you all next year!</p>",
      "rawMarkdown": "I’ve started this competition only a month before the end,  and I am very happy to finish with silver and 14-th place 😊  \nFirst of all, thanks to the Kaggle Team and Cornell Lab of Ornithology for hosting this competition, to all participants, it was a very exciting experience. \n## My solution\nSince this year  time for inference was limited to only 120 minutes on CPU,  it didn't make sense to use large models or a large image size. Also previous competitions results confirm this.  \nI’ve tested different architectures and image sizes, and finally I’ve chosen size (1, 144, 244) for a 5-sec chunk. This size allowed me to ensemble up to 6 models, including SED. Hope, next year will be an opportunity to try ONNX and openvino for inference.\n\n### The solution includes:\n**Prediction** on MelSpectrogram with f_min=100 and f_max=15000.\nFrequencies were chosen based on species in this specific competition. I've checked that for all of birds the voice lies within this interval.\n**Augmentation:**\n- random mix-up with other records and background noise (“nocall”),\n- pink noise, gaussian noise,\n- random volume adjustment,\n- random fade in, fade out,\n- random low and high pass filters.  \n\n**Detection** regions of interest (ROI) with librosa.effects.split():\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5732512%2F4186c584d604d24f63b5a16e72143fcc%2Ffull_mel.jpg?generation=1685192784649095&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5732512%2Fa367d2572733257ddc43652a0d5ce654%2Ffull_mel_with_mask.jpg?generation=1685192822856364&alt=media)\n\n***!!! top_db*** parametert must be chosen very carefully\nAfter extension and concatenation of split results, for part of the training I’ve used chunks whose center lies within the selected intervals. Not for all records this worked well, but for the most, that I've checked, the results were fine.   \n\n**Models:**\n - EfficientNetV2-S with geographical coordinates, converted to [cartesian coordinate system](https://www.kaggle.com/code/nataliayurasova/birdclef23-wgs-84-vs-cartesian-system) as additional input,\n - SED-architecture  with  EfficientNetV2-S as backbone, with image size 1x144x244 per 5-sec audio chunk,\nSED models were trained on 15-25 seconds clips.  \n\n**Loss:**\nAs a loss function I’ve used the weighted sum of BCEWithLogitsLoss() for species and bird’s families prediction.\n\n**Additional data** \n - Soundscapes and  labeled tests [from previous](https://www.kaggle.com/competitions/birdclef-2023/discussion/394358#2179605) competitions (I've extracted and used only parts without birds as \"nocall').\n - Records from  https://xeno-canto.org/ .\n\n**The final ensemble** includes 6 best models, probabilities were averaged with no pre- or post-processing.   \n\n**What doesn’t work:**\n - Use of secondary labels.\n - EfficientNetB0 performed worse even with bigger image size.\n - Training on all data (starting from 5 sec) and validation on the first 5 seconds.  \n - Increasing the probabilities if bird occurres more than once in soundscape (different thresholds and coefficients).\n - Increasing the probabilities of the same bird in the nearest chunks (different thresholds and coefficients).\n\nHope to see another BirdCLEF competition and you all next year!",
      "votes": null
    },
    {
      "id": "2277736",
      "postDate": "05/28/2023 02:45:18",
      "content": "<p>Thanks for sharing! Good good material for study!!</p>",
      "rawMarkdown": "Thanks for sharing! Good good material for study!!",
      "votes": null
    },
    {
      "id": "2278463",
      "postDate": "05/28/2023 17:51:49",
      "content": "<p>I was really curious about your solution, thanks for sharing. Your leaderboard climbing was amazing and after reading your post, it’s a shame your final position, you deserved a gold medal. This was the first solution I've read that uses something to detect the sound intervals and coordinates. Of these 2 approaches, what do you think was the most important one?</p>",
      "rawMarkdown": "I was really curious about your solution, thanks for sharing. Your leaderboard climbing was amazing and after reading your post, it’s a shame your final position, you deserved a gold medal. This was the first solution I've read that uses something to detect the sound intervals and coordinates. Of these 2 approaches, what do you think was the most important one?",
      "votes": null
    },
    {
      "id": "2281283",
      "postDate": "05/30/2023 18:26:39",
      "content": "<p>Thank you! On the one hand, detection of ROI helps a lot, on the other,  some of the quiet regions with primary labels weren’t used for training, although SED architecture more or less reduces this problem (SED models with ROI also performed better).<br>\nI'm not a professional Data Scientist,  I'm learning everything by myself and I've started to dig into SED models only 10 days  before the end (I don't use code that I don’t understand🤔), hope, next year i will be useful😊</p>",
      "rawMarkdown": "Thank you! On the one hand, detection of ROI helps a lot, on the other,  some of the quiet regions with primary labels weren’t used for training, although SED architecture more or less reduces this problem (SED models with ROI also performed better).\nI'm not a professional Data Scientist,  I'm learning everything by myself and I've started to dig into SED models only 10 days  before the end (I don't use code that I don’t understand🤔), hope, next year i will be useful😊",
      "votes": null
    },
    {
      "id": "2281284",
      "postDate": "05/30/2023 18:27:18",
      "content": "<p>Thank you🙂</p>",
      "rawMarkdown": "Thank you🙂",
      "votes": null
    },
    {
      "id": "2282617",
      "postDate": "05/31/2023 17:35:10",
      "content": "<p>Thanks for the info, I'll definitively try ROI if there's another Birdclef competition next year. This was my first audio competition, so SED models are also new to me too, and I've not fully understood the code. Best of luck in future competitions!</p>",
      "rawMarkdown": "Thanks for the info, I'll definitively try ROI if there's another Birdclef competition next year. This was my first audio competition, so SED models are also new to me too, and I've not fully understood the code. Best of luck in future competitions!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2277736,
      "author_name": "jeongtaewoo",
      "author_url": "",
      "post_date": "05/28/2023 02:45:18",
      "content": "<p>Thanks for sharing! Good good material for study!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2281284,
          "author_name": "nataliayurasova",
          "author_url": "",
          "post_date": "05/30/2023 18:27:18",
          "content": "<p>Thank you🙂</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2278463,
      "author_name": "maxdiazbattan",
      "author_url": "",
      "post_date": "05/28/2023 17:51:49",
      "content": "<p>I was really curious about your solution, thanks for sharing. Your leaderboard climbing was amazing and after reading your post, it’s a shame your final position, you deserved a gold medal. This was the first solution I've read that uses something to detect the sound intervals and coordinates. Of these 2 approaches, what do you think was the most important one?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2281283,
          "author_name": "nataliayurasova",
          "author_url": "",
          "post_date": "05/30/2023 18:26:39",
          "content": "<p>Thank you! On the one hand, detection of ROI helps a lot, on the other,  some of the quiet regions with primary labels weren’t used for training, although SED architecture more or less reduces this problem (SED models with ROI also performed better).<br>\nI'm not a professional Data Scientist,  I'm learning everything by myself and I've started to dig into SED models only 10 days  before the end (I don't use code that I don’t understand🤔), hope, next year i will be useful😊</p>",
          "votes": null,
          "replies": [
            {
              "id": 2282617,
              "author_name": "maxdiazbattan",
              "author_url": "",
              "post_date": "05/31/2023 17:35:10",
              "content": "<p>Thanks for the info, I'll definitively try ROI if there's another Birdclef competition next year. This was my first audio competition, so SED models are also new to me too, and I've not fully understood the code. Best of luck in future competitions!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2277147": "I’ve started this competition only a month before the end,  and I am very happy to finish with silver and 14-th place 😊  \nFirst of all, thanks to the Kaggle Team and Cornell Lab of Ornithology for hosting this competition, to all participants, it was a very exciting experience. \n## My solution\nSince this year  time for inference was limited to only 120 minutes on CPU,  it didn't make sense to use large models or a large image size. Also previous competitions results confirm this.  \nI’ve tested different architectures and image sizes, and finally I’ve chosen size (1, 144, 244) for a 5-sec chunk. This size allowed me to ensemble up to 6 models, including SED. Hope, next year will be an opportunity to try ONNX and openvino for inference.\n\n### The solution includes:\n**Prediction** on MelSpectrogram with f_min=100 and f_max=15000.\nFrequencies were chosen based on species in this specific competition. I've checked that for all of birds the voice lies within this interval.\n**Augmentation:**\n- random mix-up with other records and background noise (“nocall”),\n- pink noise, gaussian noise,\n- random volume adjustment,\n- random fade in, fade out,\n- random low and high pass filters.  \n\n**Detection** regions of interest (ROI) with librosa.effects.split():\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5732512%2F4186c584d604d24f63b5a16e72143fcc%2Ffull_mel.jpg?generation=1685192784649095&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5732512%2Fa367d2572733257ddc43652a0d5ce654%2Ffull_mel_with_mask.jpg?generation=1685192822856364&alt=media)\n\n***!!! top_db*** parametert must be chosen very carefully\nAfter extension and concatenation of split results, for part of the training I’ve used chunks whose center lies within the selected intervals. Not for all records this worked well, but for the most, that I've checked, the results were fine.   \n\n**Models:**\n - EfficientNetV2-S with geographical coordinates, converted to [cartesian coordinate system](https://www.kaggle.com/code/nataliayurasova/birdclef23-wgs-84-vs-cartesian-system) as additional input,\n - SED-architecture  with  EfficientNetV2-S as backbone, with image size 1x144x244 per 5-sec audio chunk,\nSED models were trained on 15-25 seconds clips.  \n\n**Loss:**\nAs a loss function I’ve used the weighted sum of BCEWithLogitsLoss() for species and bird’s families prediction.\n\n**Additional data** \n - Soundscapes and  labeled tests [from previous](https://www.kaggle.com/competitions/birdclef-2023/discussion/394358#2179605) competitions (I've extracted and used only parts without birds as \"nocall').\n - Records from  https://xeno-canto.org/ .\n\n**The final ensemble** includes 6 best models, probabilities were averaged with no pre- or post-processing.   \n\n**What doesn’t work:**\n - Use of secondary labels.\n - EfficientNetB0 performed worse even with bigger image size.\n - Training on all data (starting from 5 sec) and validation on the first 5 seconds.  \n - Increasing the probabilities if bird occurres more than once in soundscape (different thresholds and coefficients).\n - Increasing the probabilities of the same bird in the nearest chunks (different thresholds and coefficients).\n\nHope to see another BirdCLEF competition and you all next year!",
    "2277736": "Thanks for sharing! Good good material for study!!",
    "2278463": "I was really curious about your solution, thanks for sharing. Your leaderboard climbing was amazing and after reading your post, it’s a shame your final position, you deserved a gold medal. This was the first solution I've read that uses something to detect the sound intervals and coordinates. Of these 2 approaches, what do you think was the most important one?",
    "2281283": "Thank you! On the one hand, detection of ROI helps a lot, on the other,  some of the quiet regions with primary labels weren’t used for training, although SED architecture more or less reduces this problem (SED models with ROI also performed better).\nI'm not a professional Data Scientist,  I'm learning everything by myself and I've started to dig into SED models only 10 days  before the end (I don't use code that I don’t understand🤔), hope, next year i will be useful😊",
    "2281284": "Thank you🙂",
    "2282617": "Thanks for the info, I'll definitively try ROI if there's another Birdclef competition next year. This was my first audio competition, so SED models are also new to me too, and I've not fully understood the code. Best of luck in future competitions!"
  },
  "source": "meta"
}