{
  "id": 412903,
  "title": "5th place solution",
  "url": "/competitions/birdclef-2023/writeups/yevhenii-maslov-5th-place-solution",
  "author_name": "",
  "post_date": "2023-06-08T13:00:37.260Z",
  "votes": 23,
  "comment_count": 7,
  "views": 0,
  "content": "<p>First of all, thanks to the Cornell Lab of Ornithology and the Kaggle Team for hosting this competition. It was a great opportunity to learn something new.</p>\n<p>In this post, I want to present a summary of my solution.</p>\n<h3>Datasets</h3>\n<ul>\n<li>2023/2022/2021 competition data</li>\n<li>Additional Xeno-Canto data containing 2023 comp. species in foreground and background</li>\n<li>ESC50 noise</li>\n<li>No-call noise from 2021 competition data</li>\n</ul>\n<h3>Models</h3>\n<p>I used SED architecture (same as in <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> 4th place <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243293\" target=\"_blank\">solution</a>) with following backbones:</p>\n<ul>\n<li>tf_efficientnet_b1_ns</li>\n<li>tf_efficientnet_b2_ns</li>\n<li>tf_efficientnet_b3_ns</li>\n<li>tf_efficientnetv2_s_in21k</li>\n</ul>\n<h3>Training</h3>\n<p>I trained all models in two steps:</p>\n<ul>\n<li>Pretrain with 2022/2021 data. I used only white noise (p=0.5) for this step</li>\n<li>Finetune on 2023 data with the following augmentations:<ul>\n<li>For waveform - Mixup (p=1) and OneOf([White noise, pink noise, brown noise, noise injection, esc50 noise, no-call noise]) (p=0.5)</li>\n<li>For spectrogram - Two time masks (p=0.5 each) and one freq mask (p=0.5)</li></ul></li>\n</ul>\n<p>Training details:</p>\n<ul>\n<li>All models were trained on 5-sec clips</li>\n<li>4-fold stratified CV split</li>\n<li>I used both primary and secondary labels</li>\n<li>BCEWithLogitsLoss with weight for each sample based on the rating</li>\n<li>AdamW - 5e-4 lr, 1e-3 weight decay for most of the models</li>\n<li>CosineLRScheduler with default parameters</li>\n<li>40 epochs - the best score was almost always in the last epoch, so in addition to 4 folds, I also trained the 5th model using all available data. This full-fit version was consistently better than one-fold by 0.002-0.003 public and private LB.</li>\n<li>Some models were finetuned on soft/hard pseudolabels. </li>\n</ul>\n<p>My best model was trained with a combination of actual labels for competition data and hard pseudo labels for XC data. It has 0.72714/0.81836 private/public LB and 10 minutes submission time. </p>\n<h3>Inference</h3>\n<p>The probabilities from 15 models (folds) were averaged. 6 of these models have the same first three layers of the backbone, so the number of models is not very fair)</p>\n<p>I've also used ONNX runtime and ThreadPoolExecutor - thanks to the <a href=\"https://www.kaggle.com/leonshangguan\" target=\"_blank\">@leonshangguan</a> for sharing this <a href=\"https://www.kaggle.com/code/leonshangguan/faster-eb0-sed-model-inference\" target=\"_blank\">notebook</a>.</p>\n<p>Inference kernel:  <a href=\"https://www.kaggle.com/evgeniimaslov2/birdclef-5th-place-code\" target=\"_blank\">https://www.kaggle.com/evgeniimaslov2/birdclef-5th-place-code</a><br>\nGithub: <a href=\"https://github.com/yevmaslov/birdclef-2023-5th-place-solution\" target=\"_blank\">https://github.com/yevmaslov/birdclef-2023-5th-place-solution</a></p>",
  "messages": [
    {
      "id": "2274217",
      "postDate": "05/25/2023 17:47:20",
      "content": "<p>First of all, thanks to the Cornell Lab of Ornithology and the Kaggle Team for hosting this competition. It was a great opportunity to learn something new.</p>\n<p>In this post, I want to present a summary of my solution.</p>\n<h3>Datasets</h3>\n<ul>\n<li>2023/2022/2021 competition data</li>\n<li>Additional Xeno-Canto data containing 2023 comp. species in foreground and background</li>\n<li>ESC50 noise</li>\n<li>No-call noise from 2021 competition data</li>\n</ul>\n<h3>Models</h3>\n<p>I used SED architecture (same as in <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> 4th place <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243293\" target=\"_blank\">solution</a>) with following backbones:</p>\n<ul>\n<li>tf_efficientnet_b1_ns</li>\n<li>tf_efficientnet_b2_ns</li>\n<li>tf_efficientnet_b3_ns</li>\n<li>tf_efficientnetv2_s_in21k</li>\n</ul>\n<h3>Training</h3>\n<p>I trained all models in two steps:</p>\n<ul>\n<li>Pretrain with 2022/2021 data. I used only white noise (p=0.5) for this step</li>\n<li>Finetune on 2023 data with the following augmentations:<ul>\n<li>For waveform - Mixup (p=1) and OneOf([White noise, pink noise, brown noise, noise injection, esc50 noise, no-call noise]) (p=0.5)</li>\n<li>For spectrogram - Two time masks (p=0.5 each) and one freq mask (p=0.5)</li></ul></li>\n</ul>\n<p>Training details:</p>\n<ul>\n<li>All models were trained on 5-sec clips</li>\n<li>4-fold stratified CV split</li>\n<li>I used both primary and secondary labels</li>\n<li>BCEWithLogitsLoss with weight for each sample based on the rating</li>\n<li>AdamW - 5e-4 lr, 1e-3 weight decay for most of the models</li>\n<li>CosineLRScheduler with default parameters</li>\n<li>40 epochs - the best score was almost always in the last epoch, so in addition to 4 folds, I also trained the 5th model using all available data. This full-fit version was consistently better than one-fold by 0.002-0.003 public and private LB.</li>\n<li>Some models were finetuned on soft/hard pseudolabels. </li>\n</ul>\n<p>My best model was trained with a combination of actual labels for competition data and hard pseudo labels for XC data. It has 0.72714/0.81836 private/public LB and 10 minutes submission time. </p>\n<h3>Inference</h3>\n<p>The probabilities from 15 models (folds) were averaged. 6 of these models have the same first three layers of the backbone, so the number of models is not very fair)</p>\n<p>I've also used ONNX runtime and ThreadPoolExecutor - thanks to the <a href=\"https://www.kaggle.com/leonshangguan\" target=\"_blank\">@leonshangguan</a> for sharing this <a href=\"https://www.kaggle.com/code/leonshangguan/faster-eb0-sed-model-inference\" target=\"_blank\">notebook</a>.</p>\n<p>Inference kernel:  <a href=\"https://www.kaggle.com/evgeniimaslov2/birdclef-5th-place-code\" target=\"_blank\">https://www.kaggle.com/evgeniimaslov2/birdclef-5th-place-code</a><br>\nGithub: <a href=\"https://github.com/yevmaslov/birdclef-2023-5th-place-solution\" target=\"_blank\">https://github.com/yevmaslov/birdclef-2023-5th-place-solution</a></p>",
      "rawMarkdown": "First of all, thanks to the Cornell Lab of Ornithology and the Kaggle Team for hosting this competition. It was a great opportunity to learn something new.\n\nIn this post, I want to present a summary of my solution.\n\n\n\n### Datasets\n\n* 2023/2022/2021 competition data\n* Additional Xeno-Canto data containing 2023 comp. species in foreground and background\n* ESC50 noise\n* No-call noise from 2021 competition data\n\n\n\n### Models\n\nI used SED architecture (same as in @tattaka 4th place [solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243293)) with following backbones:\n\n* tf_efficientnet_b1_ns\n* tf_efficientnet_b2_ns\n* tf_efficientnet_b3_ns\n* tf_efficientnetv2_s_in21k\n\n\n\n### Training\n\nI trained all models in two steps:\n* Pretrain with 2022/2021 data. I used only white noise (p=0.5) for this step\n* Finetune on 2023 data with the following augmentations:\n\t* For waveform - Mixup (p=1) and OneOf([White noise, pink noise, brown noise, noise injection, esc50 noise, no-call noise]) (p=0.5)\n\t* For spectrogram - Two time masks (p=0.5 each) and one freq mask (p=0.5)\n\n\nTraining details:\n* All models were trained on 5-sec clips\n* 4-fold stratified CV split\n* I used both primary and secondary labels\n* BCEWithLogitsLoss with weight for each sample based on the rating\n* AdamW - 5e-4 lr, 1e-3 weight decay for most of the models\n* CosineLRScheduler with default parameters\n* 40 epochs - the best score was almost always in the last epoch, so in addition to 4 folds, I also trained the 5th model using all available data. This full-fit version was consistently better than one-fold by 0.002-0.003 public and private LB.\n* Some models were finetuned on soft/hard pseudolabels. \n\n\nMy best model was trained with a combination of actual labels for competition data and hard pseudo labels for XC data. It has 0.72714/0.81836 private/public LB and 10 minutes submission time. \n\n\n\n### Inference\n\nThe probabilities from 15 models (folds) were averaged. 6 of these models have the same first three layers of the backbone, so the number of models is not very fair)\n\nI've also used ONNX runtime and ThreadPoolExecutor - thanks to the @leonshangguan for sharing this [notebook](https://www.kaggle.com/code/leonshangguan/faster-eb0-sed-model-inference).\n\n\n\nInference kernel:  https://www.kaggle.com/evgeniimaslov2/birdclef-5th-place-code\nGithub: https://github.com/yevmaslov/birdclef-2023-5th-place-solution",
      "votes": null
    },
    {
      "id": "2275058",
      "postDate": "05/26/2023 13:08:24",
      "content": "<p>Thanks - that's very helpful! Any chance you could also put your training code on github and post a link to that?</p>",
      "rawMarkdown": "Thanks - that's very helpful! Any chance you could also put your training code on github and post a link to that?",
      "votes": null
    },
    {
      "id": "2276297",
      "postDate": "05/26/2023 17:03:33",
      "content": "<p>Hi! Sure, I'll update this thread with a link once I've finished cleaning up the code.</p>",
      "rawMarkdown": "Hi! Sure, I'll update this thread with a link once I've finished cleaning up the code.",
      "votes": null
    },
    {
      "id": "2276476",
      "postDate": "05/26/2023 21:11:56",
      "content": "<p>Hi! Many thanks for the details of your solution. I have a question on training \"using all available data\". When you had no validation data, did you have a specific criteria for early stopping or did you still train all 40 epochs? </p>",
      "rawMarkdown": "Hi! Many thanks for the details of your solution. I have a question on training \"using all available data\". When you had no validation data, did you have a specific criteria for early stopping or did you still train all 40 epochs?",
      "votes": null
    },
    {
      "id": "2276676",
      "postDate": "05/27/2023 05:55:55",
      "content": "<p>Hi! I trained all 40 epochs</p>",
      "rawMarkdown": "Hi! I trained all 40 epochs",
      "votes": null
    },
    {
      "id": "2286241",
      "postDate": "06/03/2023 10:25:28",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/evgeniimaslov2\" target=\"_blank\">@evgeniimaslov2</a> and thanks for sharing this writeup! 🙏🎉</p>",
      "rawMarkdown": "Congrats @evgeniimaslov2 and thanks for sharing this writeup! 🙏🎉",
      "votes": null
    },
    {
      "id": "2287185",
      "postDate": "06/04/2023 07:41:47",
      "content": "<p>Congratulation dear, many thanks for share details </p>",
      "rawMarkdown": "Congratulation dear, many thanks for share details",
      "votes": null
    },
    {
      "id": "2292579",
      "postDate": "06/08/2023 13:02:52",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/janhuus\" target=\"_blank\">@janhuus</a>  I've uploaded the training code to github, hope it's still relevant to you</p>",
      "rawMarkdown": "Hi @janhuus  I've uploaded the training code to github, hope it's still relevant to you",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2275058,
      "author_name": "janhuus",
      "author_url": "",
      "post_date": "05/26/2023 13:08:24",
      "content": "<p>Thanks - that's very helpful! Any chance you could also put your training code on github and post a link to that?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2276297,
          "author_name": "evgeniimaslov2",
          "author_url": "",
          "post_date": "05/26/2023 17:03:33",
          "content": "<p>Hi! Sure, I'll update this thread with a link once I've finished cleaning up the code.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2292579,
          "author_name": "evgeniimaslov2",
          "author_url": "",
          "post_date": "06/08/2023 13:02:52",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/janhuus\" target=\"_blank\">@janhuus</a>  I've uploaded the training code to github, hope it's still relevant to you</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2276476,
      "author_name": "hakandogan",
      "author_url": "",
      "post_date": "05/26/2023 21:11:56",
      "content": "<p>Hi! Many thanks for the details of your solution. I have a question on training \"using all available data\". When you had no validation data, did you have a specific criteria for early stopping or did you still train all 40 epochs? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2276676,
          "author_name": "evgeniimaslov2",
          "author_url": "",
          "post_date": "05/27/2023 05:55:55",
          "content": "<p>Hi! I trained all 40 epochs</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2286241,
      "author_name": "pardeep19singh",
      "author_url": "",
      "post_date": "06/03/2023 10:25:28",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/evgeniimaslov2\" target=\"_blank\">@evgeniimaslov2</a> and thanks for sharing this writeup! 🙏🎉</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2287185,
      "author_name": "awaisasghar",
      "author_url": "",
      "post_date": "06/04/2023 07:41:47",
      "content": "<p>Congratulation dear, many thanks for share details </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2274217": "First of all, thanks to the Cornell Lab of Ornithology and the Kaggle Team for hosting this competition. It was a great opportunity to learn something new.\n\nIn this post, I want to present a summary of my solution.\n\n\n\n### Datasets\n\n* 2023/2022/2021 competition data\n* Additional Xeno-Canto data containing 2023 comp. species in foreground and background\n* ESC50 noise\n* No-call noise from 2021 competition data\n\n\n\n### Models\n\nI used SED architecture (same as in @tattaka 4th place [solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243293)) with following backbones:\n\n* tf_efficientnet_b1_ns\n* tf_efficientnet_b2_ns\n* tf_efficientnet_b3_ns\n* tf_efficientnetv2_s_in21k\n\n\n\n### Training\n\nI trained all models in two steps:\n* Pretrain with 2022/2021 data. I used only white noise (p=0.5) for this step\n* Finetune on 2023 data with the following augmentations:\n\t* For waveform - Mixup (p=1) and OneOf([White noise, pink noise, brown noise, noise injection, esc50 noise, no-call noise]) (p=0.5)\n\t* For spectrogram - Two time masks (p=0.5 each) and one freq mask (p=0.5)\n\n\nTraining details:\n* All models were trained on 5-sec clips\n* 4-fold stratified CV split\n* I used both primary and secondary labels\n* BCEWithLogitsLoss with weight for each sample based on the rating\n* AdamW - 5e-4 lr, 1e-3 weight decay for most of the models\n* CosineLRScheduler with default parameters\n* 40 epochs - the best score was almost always in the last epoch, so in addition to 4 folds, I also trained the 5th model using all available data. This full-fit version was consistently better than one-fold by 0.002-0.003 public and private LB.\n* Some models were finetuned on soft/hard pseudolabels. \n\n\nMy best model was trained with a combination of actual labels for competition data and hard pseudo labels for XC data. It has 0.72714/0.81836 private/public LB and 10 minutes submission time. \n\n\n\n### Inference\n\nThe probabilities from 15 models (folds) were averaged. 6 of these models have the same first three layers of the backbone, so the number of models is not very fair)\n\nI've also used ONNX runtime and ThreadPoolExecutor - thanks to the @leonshangguan for sharing this [notebook](https://www.kaggle.com/code/leonshangguan/faster-eb0-sed-model-inference).\n\n\n\nInference kernel:  https://www.kaggle.com/evgeniimaslov2/birdclef-5th-place-code\nGithub: https://github.com/yevmaslov/birdclef-2023-5th-place-solution",
    "2275058": "Thanks - that's very helpful! Any chance you could also put your training code on github and post a link to that?",
    "2276297": "Hi! Sure, I'll update this thread with a link once I've finished cleaning up the code.",
    "2276476": "Hi! Many thanks for the details of your solution. I have a question on training \"using all available data\". When you had no validation data, did you have a specific criteria for early stopping or did you still train all 40 epochs?",
    "2276676": "Hi! I trained all 40 epochs",
    "2286241": "Congrats @evgeniimaslov2 and thanks for sharing this writeup! 🙏🎉",
    "2287185": "Congratulation dear, many thanks for share details",
    "2292579": "Hi @janhuus  I've uploaded the training code to github, hope it's still relevant to you"
  },
  "source": "meta"
}