{
  "id": 326966,
  "title": "65th place writeup - bronze",
  "url": "/competitions/birdclef-2022/writeups/enemy-pidgey-used-sand-attack-65th-place-writeup-b",
  "author_name": "",
  "post_date": "2022-05-27T18:06:51.823Z",
  "votes": 14,
  "comment_count": 2,
  "views": 0,
  "content": "<p>This was an awesome competition and I learned a lot from it. This is due to the hosts, to my teammates, and the kagglers who shared ideas!</p>\n<p>A few things about our solution:</p>\n<ul>\n<li>Ensemble of SED models with two different backbones: seresnext26tn_32x4d and  tf_efficientnet_b0_ns, both from timm, 5 folds each, 30 epochs each run</li>\n<li>Each backbone used a different seed</li>\n<li>Both backbones used primary and secondary labels in training</li>\n<li>Everything was trained on google colab notebooks</li>\n<li>The optimal threshold for the ensemble was 0.25</li>\n<li>Audio augmentations and image augmentations</li>\n<li>Cutmix and mixup</li>\n<li>CosineAnnealingLR (Tmax=7 for seresnext and Tmax=4 for effnet). Couldn't find any improvements using other schedulers</li>\n<li>seresnext was trained with the first 5 seconds of each audio file</li>\n<li>tf_efficientnet_b0_ns was trained with seconds 5 to 10 when audio file has more than 10 seconds and with first 5 seconds otherwise</li>\n<li>We used <a href=\"https://arxiv.org/pdf/1708.02002.pdf\" target=\"_blank\">focal loss</a>, tuned its parameters, and found that using alpha = 0.75 and gamma = 5 gave a huge boost on our models</li>\n<li>From the moment we started using gamma = 5 in focal loss we had good CV/LB correlation for most cases, using <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/314999\" target=\"_blank\">method</a></li>\n<li>We also monitored another computation of f1 score, similar to the one above but only for scored birds. So we could track the learning of scored birds, which was slower than the other birds, but in the end of training both CV strategies and calculated f1 were quite similar</li>\n</ul>\n<p>Things that didn't work:</p>\n<ul>\n<li>Noise reduction</li>\n<li>PCEN (Per-Channel Energy Normalization)</li>\n<li><a href=\"https://arxiv.org/abs/1711.02512\" target=\"_blank\">GeM</a></li>\n<li>AdamW</li>\n<li>Pseudo labels via oof predictions (we tried this a lot but couldn't make it work)</li>\n<li>Reducing dataset size or selecting samples from it</li>\n</ul>\n<p>Again, thanks to my teammates who contributed a lot to this score:<br>\n<a href=\"https://www.kaggle.com/gabrielvinicius\" target=\"_blank\">@gabrielvinicius</a> <br>\n<a href=\"https://www.kaggle.com/lucasdmr\" target=\"_blank\">@lucasdmr</a> <br>\n<a href=\"https://www.kaggle.com/paulojunqueira\" target=\"_blank\">@paulojunqueira</a> <br>\n<a href=\"https://www.kaggle.com/felipemandrade\" target=\"_blank\">@felipemandrade</a></p>\n<p>We put a lot of time, effort and dedication to achieve this result.</p>\n<p>In my opinion, what made this problem really hard was the weak labels. We tried to deal with this using pseudo-labels but we weren't successful on this.</p>",
  "messages": [
    {
      "id": "1800601",
      "postDate": "05/25/2022 04:31:16",
      "content": "<p>This was an awesome competition and I learned a lot from it. This is due to the hosts, to my teammates, and the kagglers who shared ideas!</p>\n<p>A few things about our solution:</p>\n<ul>\n<li>Ensemble of SED models with two different backbones: seresnext26tn_32x4d and  tf_efficientnet_b0_ns, both from timm, 5 folds each, 30 epochs each run</li>\n<li>Each backbone used a different seed</li>\n<li>Both backbones used primary and secondary labels in training</li>\n<li>Everything was trained on google colab notebooks</li>\n<li>The optimal threshold for the ensemble was 0.25</li>\n<li>Audio augmentations and image augmentations</li>\n<li>Cutmix and mixup</li>\n<li>CosineAnnealingLR (Tmax=7 for seresnext and Tmax=4 for effnet). Couldn't find any improvements using other schedulers</li>\n<li>seresnext was trained with the first 5 seconds of each audio file</li>\n<li>tf_efficientnet_b0_ns was trained with seconds 5 to 10 when audio file has more than 10 seconds and with first 5 seconds otherwise</li>\n<li>We used <a href=\"https://arxiv.org/pdf/1708.02002.pdf\" target=\"_blank\">focal loss</a>, tuned its parameters, and found that using alpha = 0.75 and gamma = 5 gave a huge boost on our models</li>\n<li>From the moment we started using gamma = 5 in focal loss we had good CV/LB correlation for most cases, using <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/314999\" target=\"_blank\">method</a></li>\n<li>We also monitored another computation of f1 score, similar to the one above but only for scored birds. So we could track the learning of scored birds, which was slower than the other birds, but in the end of training both CV strategies and calculated f1 were quite similar</li>\n</ul>\n<p>Things that didn't work:</p>\n<ul>\n<li>Noise reduction</li>\n<li>PCEN (Per-Channel Energy Normalization)</li>\n<li><a href=\"https://arxiv.org/abs/1711.02512\" target=\"_blank\">GeM</a></li>\n<li>AdamW</li>\n<li>Pseudo labels via oof predictions (we tried this a lot but couldn't make it work)</li>\n<li>Reducing dataset size or selecting samples from it</li>\n</ul>\n<p>Again, thanks to my teammates who contributed a lot to this score:<br>\n<a href=\"https://www.kaggle.com/gabrielvinicius\" target=\"_blank\">@gabrielvinicius</a> <br>\n<a href=\"https://www.kaggle.com/lucasdmr\" target=\"_blank\">@lucasdmr</a> <br>\n<a href=\"https://www.kaggle.com/paulojunqueira\" target=\"_blank\">@paulojunqueira</a> <br>\n<a href=\"https://www.kaggle.com/felipemandrade\" target=\"_blank\">@felipemandrade</a></p>\n<p>We put a lot of time, effort and dedication to achieve this result.</p>\n<p>In my opinion, what made this problem really hard was the weak labels. We tried to deal with this using pseudo-labels but we weren't successful on this.</p>",
      "rawMarkdown": "This was an awesome competition and I learned a lot from it. This is due to the hosts, to my teammates, and the kagglers who shared ideas!\n\nA few things about our solution:\n- Ensemble of SED models with two different backbones: seresnext26tn_32x4d and  tf_efficientnet_b0_ns, both from timm, 5 folds each, 30 epochs each run\n- Each backbone used a different seed\n- Both backbones used primary and secondary labels in training\n- Everything was trained on google colab notebooks\n- The optimal threshold for the ensemble was 0.25\n- Audio augmentations and image augmentations\n- Cutmix and mixup\n- CosineAnnealingLR (Tmax=7 for seresnext and Tmax=4 for effnet). Couldn't find any improvements using other schedulers\n- seresnext was trained with the first 5 seconds of each audio file\n- tf_efficientnet_b0_ns was trained with seconds 5 to 10 when audio file has more than 10 seconds and with first 5 seconds otherwise\n- We used [focal loss](https://arxiv.org/pdf/1708.02002.pdf), tuned its parameters, and found that using alpha = 0.75 and gamma = 5 gave a huge boost on our models\n- From the moment we started using gamma = 5 in focal loss we had good CV/LB correlation for most cases, using @dschettler8845 [method](https://www.kaggle.com/competitions/birdclef-2022/discussion/314999)\n- We also monitored another computation of f1 score, similar to the one above but only for scored birds. So we could track the learning of scored birds, which was slower than the other birds, but in the end of training both CV strategies and calculated f1 were quite similar\n\nThings that didn't work:\n- Noise reduction\n- PCEN (Per-Channel Energy Normalization)\n- [GeM](https://arxiv.org/abs/1711.02512)\n- AdamW\n- Pseudo labels via oof predictions (we tried this a lot but couldn't make it work)\n- Reducing dataset size or selecting samples from it\n\nAgain, thanks to my teammates who contributed a lot to this score:\n@gabrielvinicius \n@lucasdmr \n@paulojunqueira \n@felipemandrade\n\nWe put a lot of time, effort and dedication to achieve this result.\n\nIn my opinion, what made this problem really hard was the weak labels. We tried to deal with this using pseudo-labels but we weren't successful on this.",
      "votes": null
    },
    {
      "id": "1801051",
      "postDate": "05/25/2022 11:43:58",
      "content": "<p>It was a pleasure to work with you guys. I have learned a lot!</p>",
      "rawMarkdown": "It was a pleasure to work with you guys. I have learned a lot!",
      "votes": null
    },
    {
      "id": "1801066",
      "postDate": "05/25/2022 12:04:55",
      "content": "<p>Great job, guys! You rock! </p>",
      "rawMarkdown": "Great job, guys! You rock!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1801051,
      "author_name": "paulojunqueira",
      "author_url": "",
      "post_date": "05/25/2022 11:43:58",
      "content": "<p>It was a pleasure to work with you guys. I have learned a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1801066,
      "author_name": "marcusdipaula",
      "author_url": "",
      "post_date": "05/25/2022 12:04:55",
      "content": "<p>Great job, guys! You rock! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1800601": "This was an awesome competition and I learned a lot from it. This is due to the hosts, to my teammates, and the kagglers who shared ideas!\n\nA few things about our solution:\n- Ensemble of SED models with two different backbones: seresnext26tn_32x4d and  tf_efficientnet_b0_ns, both from timm, 5 folds each, 30 epochs each run\n- Each backbone used a different seed\n- Both backbones used primary and secondary labels in training\n- Everything was trained on google colab notebooks\n- The optimal threshold for the ensemble was 0.25\n- Audio augmentations and image augmentations\n- Cutmix and mixup\n- CosineAnnealingLR (Tmax=7 for seresnext and Tmax=4 for effnet). Couldn't find any improvements using other schedulers\n- seresnext was trained with the first 5 seconds of each audio file\n- tf_efficientnet_b0_ns was trained with seconds 5 to 10 when audio file has more than 10 seconds and with first 5 seconds otherwise\n- We used [focal loss](https://arxiv.org/pdf/1708.02002.pdf), tuned its parameters, and found that using alpha = 0.75 and gamma = 5 gave a huge boost on our models\n- From the moment we started using gamma = 5 in focal loss we had good CV/LB correlation for most cases, using @dschettler8845 [method](https://www.kaggle.com/competitions/birdclef-2022/discussion/314999)\n- We also monitored another computation of f1 score, similar to the one above but only for scored birds. So we could track the learning of scored birds, which was slower than the other birds, but in the end of training both CV strategies and calculated f1 were quite similar\n\nThings that didn't work:\n- Noise reduction\n- PCEN (Per-Channel Energy Normalization)\n- [GeM](https://arxiv.org/abs/1711.02512)\n- AdamW\n- Pseudo labels via oof predictions (we tried this a lot but couldn't make it work)\n- Reducing dataset size or selecting samples from it\n\nAgain, thanks to my teammates who contributed a lot to this score:\n@gabrielvinicius \n@lucasdmr \n@paulojunqueira \n@felipemandrade\n\nWe put a lot of time, effort and dedication to achieve this result.\n\nIn my opinion, what made this problem really hard was the weak labels. We tried to deal with this using pseudo-labels but we weren't successful on this.",
    "1801051": "It was a pleasure to work with you guys. I have learned a lot!",
    "1801066": "Great job, guys! You rock!"
  },
  "source": "meta"
}