{
  "id": 326968,
  "title": "9th place summary",
  "url": "/competitions/birdclef-2022/writeups/fly-9th-place-summary",
  "author_name": "",
  "post_date": "2022-05-25T04:54:27.930Z",
  "votes": 31,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Thanks to Cornell Lab of Ornithology and Kaggle for hosting this interesting competition. I also would like to thank my teammates <a href=\"https://www.kaggle.com/truonghoang\" target=\"_blank\">@truonghoang</a> and <a href=\"https://www.kaggle.com/gwanghan\" target=\"_blank\">@gwanghan</a> for the great collaborative teamwork during the competition.</p>\n<p>Our solution is an ensemble of 1 SED model and 1 melspectrogram classification model which score 0.76 and 0.77 respectively in private LB. ensemble gave us 0.79.</p>\n<h2>SED</h2>\n<ul>\n<li>Backbone: tf_efficientnet_b5_ns.</li>\n<li>Train only on the first 5s</li>\n<li>Inference 5s</li>\n<li>Since the backbone is powerful, we use quite a lot augmentation in the hope that it will help closing the domain gap between training data and the test data. The augmentation include audio augmentation, melspec augmentation, mixup 2-3 samples and cutout.</li>\n</ul>\n<h2>Melspectrogram classification</h2>\n<ul>\n<li>Backbone: resnest50d_1s4x24d</li>\n<li>Train on 7s random crop.</li>\n<li>Inference 7s</li>\n<li>Augmentation: we use simple augmentation</li>\n</ul>\n<pre><code>    NoiseInjection(max_noise_level=0.04, sr=SAMPLE_RATE),\n    PitchShift(max_range=3, sr=SAMPLE_RATE),\n    RandomVolume(limit=4),\n</code></pre>\n<ul>\n<li>Training: <ul>\n<li>Round 1: We train several models and ensemble to create an oof for each 7s crop.</li>\n<li>Round 2: We conbine groundtruth label and oof to modify the target of each 7s, then train on the modified target.</li></ul></li>\n</ul>\n<pre><code>  if oof_prob[primary_bird]&gt;0.5:\n      target[primary_bird] = 1.0\n  elif prob[primary_bird] &gt; 0.1:\n      target[primary_bird] = 0.9975\n  elif prob[primary_bird] &gt; 0.01:\n      target[primary_bird] = 0.5\n  else:\n      target[primary_bird] = 0.2\n\n  if oof_prob[secondary_bird]&gt;0.5:\n      target[secondary_bird] = 0.9975\n  elif oof_prob[secondary_bird] &gt; 0.1:\n      target[secondary_bird] = 0.8\n  else:\n      target[secondary_bird] = 0.0025\n</code></pre>\n<h2>Ensemble</h2>\n<p>We found that the public lb score is very sensitive to threshold, the optimal threshold for each model are very different, so doing weight average does not bring much improvement in our ensemble. We end up multiplying their probability and set the top 33% highest confidence score in the test set as True, the rest is False.</p>\n<p>P/S: Congrats my hard-working teammate <a href=\"https://www.kaggle.com/truonghoang\" target=\"_blank\">@truonghoang</a> on becomming competition master</p>",
  "messages": [
    {
      "id": "1800612",
      "postDate": "05/25/2022 04:49:41",
      "content": "<p>Thanks to Cornell Lab of Ornithology and Kaggle for hosting this interesting competition. I also would like to thank my teammates <a href=\"https://www.kaggle.com/truonghoang\" target=\"_blank\">@truonghoang</a> and <a href=\"https://www.kaggle.com/gwanghan\" target=\"_blank\">@gwanghan</a> for the great collaborative teamwork during the competition.</p>\n<p>Our solution is an ensemble of 1 SED model and 1 melspectrogram classification model which score 0.76 and 0.77 respectively in private LB. ensemble gave us 0.79.</p>\n<h2>SED</h2>\n<ul>\n<li>Backbone: tf_efficientnet_b5_ns.</li>\n<li>Train only on the first 5s</li>\n<li>Inference 5s</li>\n<li>Since the backbone is powerful, we use quite a lot augmentation in the hope that it will help closing the domain gap between training data and the test data. The augmentation include audio augmentation, melspec augmentation, mixup 2-3 samples and cutout.</li>\n</ul>\n<h2>Melspectrogram classification</h2>\n<ul>\n<li>Backbone: resnest50d_1s4x24d</li>\n<li>Train on 7s random crop.</li>\n<li>Inference 7s</li>\n<li>Augmentation: we use simple augmentation</li>\n</ul>\n<pre><code>    NoiseInjection(max_noise_level=0.04, sr=SAMPLE_RATE),\n    PitchShift(max_range=3, sr=SAMPLE_RATE),\n    RandomVolume(limit=4),\n</code></pre>\n<ul>\n<li>Training: <ul>\n<li>Round 1: We train several models and ensemble to create an oof for each 7s crop.</li>\n<li>Round 2: We conbine groundtruth label and oof to modify the target of each 7s, then train on the modified target.</li></ul></li>\n</ul>\n<pre><code>  if oof_prob[primary_bird]&gt;0.5:\n      target[primary_bird] = 1.0\n  elif prob[primary_bird] &gt; 0.1:\n      target[primary_bird] = 0.9975\n  elif prob[primary_bird] &gt; 0.01:\n      target[primary_bird] = 0.5\n  else:\n      target[primary_bird] = 0.2\n\n  if oof_prob[secondary_bird]&gt;0.5:\n      target[secondary_bird] = 0.9975\n  elif oof_prob[secondary_bird] &gt; 0.1:\n      target[secondary_bird] = 0.8\n  else:\n      target[secondary_bird] = 0.0025\n</code></pre>\n<h2>Ensemble</h2>\n<p>We found that the public lb score is very sensitive to threshold, the optimal threshold for each model are very different, so doing weight average does not bring much improvement in our ensemble. We end up multiplying their probability and set the top 33% highest confidence score in the test set as True, the rest is False.</p>\n<p>P/S: Congrats my hard-working teammate <a href=\"https://www.kaggle.com/truonghoang\" target=\"_blank\">@truonghoang</a> on becomming competition master</p>",
      "rawMarkdown": "Thanks to Cornell Lab of Ornithology and Kaggle for hosting this interesting competition. I also would like to thank my teammates @truonghoang and @gwanghan for the great collaborative teamwork during the competition.\n\nOur solution is an ensemble of 1 SED model and 1 melspectrogram classification model which score 0.76 and 0.77 respectively in private LB. ensemble gave us 0.79.\n\n## SED\n- Backbone: tf_efficientnet_b5_ns.\n- Train only on the first 5s\n- Inference 5s\n- Since the backbone is powerful, we use quite a lot augmentation in the hope that it will help closing the domain gap between training data and the test data. The augmentation include audio augmentation, melspec augmentation, mixup 2-3 samples and cutout.\n\n## Melspectrogram classification\n- Backbone: resnest50d_1s4x24d\n- Train on 7s random crop.\n- Inference 7s\n- Augmentation: we use simple augmentation\n```\n    NoiseInjection(max_noise_level=0.04, sr=SAMPLE_RATE),\n    PitchShift(max_range=3, sr=SAMPLE_RATE),\n    RandomVolume(limit=4),\n```\n- Training: \n  - Round 1: We train several models and ensemble to create an oof for each 7s crop.\n  - Round 2: We conbine groundtruth label and oof to modify the target of each 7s, then train on the modified target.\n  ```\n  if oof_prob[primary_bird]>0.5:\n      target[primary_bird] = 1.0\n  elif prob[primary_bird] > 0.1:\n      target[primary_bird] = 0.9975\n  elif prob[primary_bird] > 0.01:\n      target[primary_bird] = 0.5\n  else:\n      target[primary_bird] = 0.2\n      \n  if oof_prob[secondary_bird]>0.5:\n      target[secondary_bird] = 0.9975\n  elif oof_prob[secondary_bird] > 0.1:\n      target[secondary_bird] = 0.8\n  else:\n      target[secondary_bird] = 0.0025\n  ```\n## Ensemble\n We found that the public lb score is very sensitive to threshold, the optimal threshold for each model are very different, so doing weight average does not bring much improvement in our ensemble. We end up multiplying their probability and set the top 33% highest confidence score in the test set as True, the rest is False.\n\nP/S: Congrats my hard-working teammate @truonghoang on becomming competition master",
      "votes": null
    },
    {
      "id": "1800616",
      "postDate": "05/25/2022 05:00:48",
      "content": "<p>Thanks for sharing and congratulations！</p>",
      "rawMarkdown": "Thanks for sharing and congratulations！",
      "votes": null
    },
    {
      "id": "1800624",
      "postDate": "05/25/2022 05:11:44",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/nvnnghia\" target=\"_blank\">@nvnnghia</a> <a href=\"https://www.kaggle.com/truonghoang\" target=\"_blank\">@truonghoang</a> ! Well done and nice idea</p>",
      "rawMarkdown": "Congrats @nvnnghia @truonghoang ! Well done and nice idea",
      "votes": null
    },
    {
      "id": "1800633",
      "postDate": "05/25/2022 05:32:24",
      "content": "<p>congratulations on the gold! I have a few questions, if you don't mind:</p>\n<ul>\n<li>There are 21 birds that are evaluated. how did you come up with 33%? Do you mean that upper-33%-scoring birds were selected for each sample? Or do you mean that your model infers 33% of the whole set includes an \"apapan\" and the rest does not?</li>\n<li>For the SED backbone, what prompted you to use b5 instead of, for instance, b0 or b2, v2? was it because it simply seemed to generalize better?</li>\n</ul>",
      "rawMarkdown": "congratulations on the gold! I have a few questions, if you don't mind:\n- There are 21 birds that are evaluated. how did you come up with 33%? Do you mean that upper-33%-scoring birds were selected for each sample? Or do you mean that your model infers 33% of the whole set includes an \"apapan\" and the rest does not?\n- For the SED backbone, what prompted you to use b5 instead of, for instance, b0 or b2, v2? was it because it simply seemed to generalize better?",
      "votes": null
    },
    {
      "id": "1800638",
      "postDate": "05/25/2022 05:39:35",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/nvnnghia\" target=\"_blank\">@nvnnghia</a> <a href=\"https://www.kaggle.com/truonghoang\" target=\"_blank\">@truonghoang</a> on results!</p>",
      "rawMarkdown": "Congrats @nvnnghia @truonghoang on results!",
      "votes": null
    },
    {
      "id": "1800649",
      "postDate": "05/25/2022 05:49:18",
      "content": "<ul>\n<li>so I add another column (conf) in submission.csv, then the top 33% highest conf will has target=True.<br>\n<img src=\"https://i.ibb.co/dPxwghm/Screenshot-from-2022-05-25-14-40-05.png\" alt=\"\"></li>\n<li>I tried b0, b4, b5, b7, and b5 gave better cv-lb. So I use it, I had no time to test other backbone.</li>\n</ul>",
      "rawMarkdown": "so I add another column (conf) in submission.csv, then the top 33% highest conf will has target=True.\n![](https://i.ibb.co/dPxwghm/Screenshot-from-2022-05-25-14-40-05.png)\n- I tried b0, b4, b5, b7, and b5 gave better cv-lb. So I use it, I had no time to test other backbone.",
      "votes": null
    },
    {
      "id": "1800706",
      "postDate": "05/25/2022 06:38:47",
      "content": "<p>This makes everything clear. Thank you!!</p>",
      "rawMarkdown": "This makes everything clear. Thank you!!",
      "votes": null
    },
    {
      "id": "1800806",
      "postDate": "05/25/2022 08:23:39",
      "content": "<p>Congrats to all of you and thanks for the summary 👍</p>",
      "rawMarkdown": "Congrats to all of you and thanks for the summary 👍",
      "votes": null
    },
    {
      "id": "1801055",
      "postDate": "05/25/2022 11:48:55",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/nvnnghia\" target=\"_blank\">@nvnnghia</a> <a href=\"https://www.kaggle.com/truonghoang\" target=\"_blank\">@truonghoang</a> <a href=\"https://www.kaggle.com/gwanghan\" target=\"_blank\">@gwanghan</a> ! :)</p>",
      "rawMarkdown": "Congrats @nvnnghia @truonghoang @gwanghan ! :)",
      "votes": null
    },
    {
      "id": "1801670",
      "postDate": "05/26/2022 03:00:30",
      "content": "<p>Congrats. And <a href=\"https://www.kaggle.com/truonghoang\" target=\"_blank\">@truonghoang</a> became competition master!</p>",
      "rawMarkdown": "Congrats. And @truonghoang became competition master!",
      "votes": null
    },
    {
      "id": "1801835",
      "postDate": "05/26/2022 07:38:06",
      "content": "<p>Thanks. Congrats on your solo gold and you became competition master too!!</p>",
      "rawMarkdown": "Thanks. Congrats on your solo gold and you became competition master too!!",
      "votes": null
    },
    {
      "id": "1803197",
      "postDate": "05/27/2022 14:54:50",
      "content": "<p>congratulations  <a href=\"https://www.kaggle.com/nvnnghia\" target=\"_blank\">@nvnnghia</a> <a href=\"https://www.kaggle.com/truonghoang\" target=\"_blank\">@truonghoang</a> ! Thanks for the summary</p>",
      "rawMarkdown": "congratulations  @nvnnghia @truonghoang ! Thanks for the summary",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1800616,
      "author_name": "chenbaoying",
      "author_url": "",
      "post_date": "05/25/2022 05:00:48",
      "content": "<p>Thanks for sharing and congratulations！</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1800624,
      "author_name": "doanquanvietnamca",
      "author_url": "",
      "post_date": "05/25/2022 05:11:44",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/nvnnghia\" target=\"_blank\">@nvnnghia</a> <a href=\"https://www.kaggle.com/truonghoang\" target=\"_blank\">@truonghoang</a> ! Well done and nice idea</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1800633,
      "author_name": "jayeonyi",
      "author_url": "",
      "post_date": "05/25/2022 05:32:24",
      "content": "<p>congratulations on the gold! I have a few questions, if you don't mind:</p>\n<ul>\n<li>There are 21 birds that are evaluated. how did you come up with 33%? Do you mean that upper-33%-scoring birds were selected for each sample? Or do you mean that your model infers 33% of the whole set includes an \"apapan\" and the rest does not?</li>\n<li>For the SED backbone, what prompted you to use b5 instead of, for instance, b0 or b2, v2? was it because it simply seemed to generalize better?</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1800649,
          "author_name": "nvnnghia",
          "author_url": "",
          "post_date": "05/25/2022 05:49:18",
          "content": "<ul>\n<li>so I add another column (conf) in submission.csv, then the top 33% highest conf will has target=True.<br>\n<img src=\"https://i.ibb.co/dPxwghm/Screenshot-from-2022-05-25-14-40-05.png\" alt=\"\"></li>\n<li>I tried b0, b4, b5, b7, and b5 gave better cv-lb. So I use it, I had no time to test other backbone.</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1800706,
          "author_name": "jayeonyi",
          "author_url": "",
          "post_date": "05/25/2022 06:38:47",
          "content": "<p>This makes everything clear. Thank you!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1800638,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "05/25/2022 05:39:35",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/nvnnghia\" target=\"_blank\">@nvnnghia</a> <a href=\"https://www.kaggle.com/truonghoang\" target=\"_blank\">@truonghoang</a> on results!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1800806,
      "author_name": "tjamali",
      "author_url": "",
      "post_date": "05/25/2022 08:23:39",
      "content": "<p>Congrats to all of you and thanks for the summary 👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1801055,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "05/25/2022 11:48:55",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/nvnnghia\" target=\"_blank\">@nvnnghia</a> <a href=\"https://www.kaggle.com/truonghoang\" target=\"_blank\">@truonghoang</a> <a href=\"https://www.kaggle.com/gwanghan\" target=\"_blank\">@gwanghan</a> ! :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1801670,
      "author_name": "shinmurashinmura",
      "author_url": "",
      "post_date": "05/26/2022 03:00:30",
      "content": "<p>Congrats. And <a href=\"https://www.kaggle.com/truonghoang\" target=\"_blank\">@truonghoang</a> became competition master!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1801835,
          "author_name": "truonghoang",
          "author_url": "",
          "post_date": "05/26/2022 07:38:06",
          "content": "<p>Thanks. Congrats on your solo gold and you became competition master too!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1803197,
      "author_name": "tejasurya",
      "author_url": "",
      "post_date": "05/27/2022 14:54:50",
      "content": "<p>congratulations  <a href=\"https://www.kaggle.com/nvnnghia\" target=\"_blank\">@nvnnghia</a> <a href=\"https://www.kaggle.com/truonghoang\" target=\"_blank\">@truonghoang</a> ! Thanks for the summary</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1800612": "Thanks to Cornell Lab of Ornithology and Kaggle for hosting this interesting competition. I also would like to thank my teammates @truonghoang and @gwanghan for the great collaborative teamwork during the competition.\n\nOur solution is an ensemble of 1 SED model and 1 melspectrogram classification model which score 0.76 and 0.77 respectively in private LB. ensemble gave us 0.79.\n\n## SED\n- Backbone: tf_efficientnet_b5_ns.\n- Train only on the first 5s\n- Inference 5s\n- Since the backbone is powerful, we use quite a lot augmentation in the hope that it will help closing the domain gap between training data and the test data. The augmentation include audio augmentation, melspec augmentation, mixup 2-3 samples and cutout.\n\n## Melspectrogram classification\n- Backbone: resnest50d_1s4x24d\n- Train on 7s random crop.\n- Inference 7s\n- Augmentation: we use simple augmentation\n```\n    NoiseInjection(max_noise_level=0.04, sr=SAMPLE_RATE),\n    PitchShift(max_range=3, sr=SAMPLE_RATE),\n    RandomVolume(limit=4),\n```\n- Training: \n  - Round 1: We train several models and ensemble to create an oof for each 7s crop.\n  - Round 2: We conbine groundtruth label and oof to modify the target of each 7s, then train on the modified target.\n  ```\n  if oof_prob[primary_bird]>0.5:\n      target[primary_bird] = 1.0\n  elif prob[primary_bird] > 0.1:\n      target[primary_bird] = 0.9975\n  elif prob[primary_bird] > 0.01:\n      target[primary_bird] = 0.5\n  else:\n      target[primary_bird] = 0.2\n      \n  if oof_prob[secondary_bird]>0.5:\n      target[secondary_bird] = 0.9975\n  elif oof_prob[secondary_bird] > 0.1:\n      target[secondary_bird] = 0.8\n  else:\n      target[secondary_bird] = 0.0025\n  ```\n## Ensemble\n We found that the public lb score is very sensitive to threshold, the optimal threshold for each model are very different, so doing weight average does not bring much improvement in our ensemble. We end up multiplying their probability and set the top 33% highest confidence score in the test set as True, the rest is False.\n\nP/S: Congrats my hard-working teammate @truonghoang on becomming competition master",
    "1800616": "Thanks for sharing and congratulations！",
    "1800624": "Congrats @nvnnghia @truonghoang ! Well done and nice idea",
    "1800633": "congratulations on the gold! I have a few questions, if you don't mind:\n- There are 21 birds that are evaluated. how did you come up with 33%? Do you mean that upper-33%-scoring birds were selected for each sample? Or do you mean that your model infers 33% of the whole set includes an \"apapan\" and the rest does not?\n- For the SED backbone, what prompted you to use b5 instead of, for instance, b0 or b2, v2? was it because it simply seemed to generalize better?",
    "1800638": "Congrats @nvnnghia @truonghoang on results!",
    "1800649": "so I add another column (conf) in submission.csv, then the top 33% highest conf will has target=True.\n![](https://i.ibb.co/dPxwghm/Screenshot-from-2022-05-25-14-40-05.png)\n- I tried b0, b4, b5, b7, and b5 gave better cv-lb. So I use it, I had no time to test other backbone.",
    "1800706": "This makes everything clear. Thank you!!",
    "1800806": "Congrats to all of you and thanks for the summary 👍",
    "1801055": "Congrats @nvnnghia @truonghoang @gwanghan ! :)",
    "1801670": "Congrats. And @truonghoang became competition master!",
    "1801835": "Thanks. Congrats on your solo gold and you became competition master too!!",
    "1803197": "congratulations  @nvnnghia @truonghoang ! Thanks for the summary"
  },
  "source": "meta"
}