{
  "id": 326933,
  "title": "Simple 17th Place Solution [0.81 Public, 0.77 Private]",
  "url": "/competitions/birdclef-2022/writeups/ari-enzo-simple-17th-place-solution-0-81-public-0-",
  "author_name": "",
  "post_date": "2022-05-25T21:05:28.993Z",
  "votes": 30,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Thanks to Kaggle, competition hosts, and fellow competitors for this very interesting competition. We joined this competition in the last month and had to work hard to understand the competition as neither <a href=\"https://www.kaggle.com/neomaoro\" target=\"_blank\">@neomaoro</a> or I have done anything with audio before this. Both <a href=\"https://www.kaggle.com/neomaoro\" target=\"_blank\">@neomaoro</a> and I worked equally hard on this competition.</p>\n<p><strong>TLDR</strong></p>\n<p>Our solution is based heavily on <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>'s <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">2nd place solution </a> from last year's birdclef competition. We made some modifications to this pipeline, the most significant being resizing spectrograms to 256 * 512 which gave a boost of 0.01 in public and private leaderboard.</p>\n<p><strong>Submission Notebook</strong><br>\n<a href=\"https://www.kaggle.com/code/vexxingbanana/18th-place-solution-0-77-private-0-81-public\" target=\"_blank\">https://www.kaggle.com/code/vexxingbanana/18th-place-solution-0-77-private-0-81-public</a></p>\n<p><strong>Code Pipeline and Data Setup</strong></p>\n<p>We heavily based our notebooks and python files for training and submitting on <a href=\"https://www.kaggle.com/julian3833\" target=\"_blank\">@julian3833</a>'s <a href=\"https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-train-0-66\" target=\"_blank\">training notebook</a> and <a href=\"https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-submit-0-66\" target=\"_blank\">inference notebook</a>. We setup training in the cloud on 2 A40 GPUs.</p>\n<p><strong>Bird Classifier</strong></p>\n<p>Our models were very similar and trained the same as last year's 2nd place solution with 30 second clips of train_audio resized into 6 x 5 second parts. We also used primary and secondary labels as targets. For inference, we fed in 6 x 5 second snippets into the model.</p>\n<p>We used the following backbones: eca_nfnet_l0, eca_nfnet_l1, tf_efficientnetv2_m_in21k, seresnext50_32x4d, and resnest50d_4s2x40d. Eca_nfnet_l0 and eca_nfnet_l1 achieved the best private leaderboard scores with 5 fold scores of 0.76 and 0.75 respectively. </p>\n<p>During training, we started with the 2nd place solution's training strategy then added some slight modifications. The modifications are as follows:</p>\n<ul>\n<li>Epochs: 20 -&gt; 25</li>\n<li>Optimizer: Adam -&gt; AdamW with learning rate 1e-3 and weight decay 1e-6</li>\n<li><strong>Resizing: After making a spectrogram, the spectrogram was resized to 256 x 512 using torchvision.</strong></li>\n<li>Background noise, label smoothing removed.</li>\n</ul>\n<p><strong>Validation</strong></p>\n<p>We couldn't find a way to make a good validation pipeline based on soundscapes since there were no soundscapes with labels given for this competition. Therefore, we simply used validation loss and public leaderboard as validation. </p>\n<p><strong>Ensembling</strong></p>\n<p>The ensembling of our models was pretty straightforward. We trained 5 fold models of each backbone and then averaged all of the models' predictions together. We tried out a few other ensembling methods such as purely voting and a combination of voting and averaging but found averaging all models to be the best performing. </p>\n<p><strong>What Did Not Work</strong></p>\n<p>Throughout the competition, we had ups and downs, many successes and some things that didn't work well. Here are some experiments which didn't work out for us:</p>\n<ul>\n<li>Applying PCEN</li>\n<li>Using 2021 Data  </li>\n<li>Attention head</li>\n<li>Adding augmentations such as pink noise, gaussian noise, etc.</li>\n<li>SED models</li>\n<li>Adjusting mel spectrograms values to increase their size (window_size 1024, hop_size 320)</li>\n<li>Using recordings with either rating 0 or ratings &gt;= 2 and label smoothing of 0.01</li>\n<li>More epochs</li>\n</ul>",
  "messages": [
    {
      "id": "1800468",
      "postDate": "05/25/2022 01:52:53",
      "content": "<p>Thanks to Kaggle, competition hosts, and fellow competitors for this very interesting competition. We joined this competition in the last month and had to work hard to understand the competition as neither <a href=\"https://www.kaggle.com/neomaoro\" target=\"_blank\">@neomaoro</a> or I have done anything with audio before this. Both <a href=\"https://www.kaggle.com/neomaoro\" target=\"_blank\">@neomaoro</a> and I worked equally hard on this competition.</p>\n<p><strong>TLDR</strong></p>\n<p>Our solution is based heavily on <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>'s <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">2nd place solution </a> from last year's birdclef competition. We made some modifications to this pipeline, the most significant being resizing spectrograms to 256 * 512 which gave a boost of 0.01 in public and private leaderboard.</p>\n<p><strong>Submission Notebook</strong><br>\n<a href=\"https://www.kaggle.com/code/vexxingbanana/18th-place-solution-0-77-private-0-81-public\" target=\"_blank\">https://www.kaggle.com/code/vexxingbanana/18th-place-solution-0-77-private-0-81-public</a></p>\n<p><strong>Code Pipeline and Data Setup</strong></p>\n<p>We heavily based our notebooks and python files for training and submitting on <a href=\"https://www.kaggle.com/julian3833\" target=\"_blank\">@julian3833</a>'s <a href=\"https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-train-0-66\" target=\"_blank\">training notebook</a> and <a href=\"https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-submit-0-66\" target=\"_blank\">inference notebook</a>. We setup training in the cloud on 2 A40 GPUs.</p>\n<p><strong>Bird Classifier</strong></p>\n<p>Our models were very similar and trained the same as last year's 2nd place solution with 30 second clips of train_audio resized into 6 x 5 second parts. We also used primary and secondary labels as targets. For inference, we fed in 6 x 5 second snippets into the model.</p>\n<p>We used the following backbones: eca_nfnet_l0, eca_nfnet_l1, tf_efficientnetv2_m_in21k, seresnext50_32x4d, and resnest50d_4s2x40d. Eca_nfnet_l0 and eca_nfnet_l1 achieved the best private leaderboard scores with 5 fold scores of 0.76 and 0.75 respectively. </p>\n<p>During training, we started with the 2nd place solution's training strategy then added some slight modifications. The modifications are as follows:</p>\n<ul>\n<li>Epochs: 20 -&gt; 25</li>\n<li>Optimizer: Adam -&gt; AdamW with learning rate 1e-3 and weight decay 1e-6</li>\n<li><strong>Resizing: After making a spectrogram, the spectrogram was resized to 256 x 512 using torchvision.</strong></li>\n<li>Background noise, label smoothing removed.</li>\n</ul>\n<p><strong>Validation</strong></p>\n<p>We couldn't find a way to make a good validation pipeline based on soundscapes since there were no soundscapes with labels given for this competition. Therefore, we simply used validation loss and public leaderboard as validation. </p>\n<p><strong>Ensembling</strong></p>\n<p>The ensembling of our models was pretty straightforward. We trained 5 fold models of each backbone and then averaged all of the models' predictions together. We tried out a few other ensembling methods such as purely voting and a combination of voting and averaging but found averaging all models to be the best performing. </p>\n<p><strong>What Did Not Work</strong></p>\n<p>Throughout the competition, we had ups and downs, many successes and some things that didn't work well. Here are some experiments which didn't work out for us:</p>\n<ul>\n<li>Applying PCEN</li>\n<li>Using 2021 Data  </li>\n<li>Attention head</li>\n<li>Adding augmentations such as pink noise, gaussian noise, etc.</li>\n<li>SED models</li>\n<li>Adjusting mel spectrograms values to increase their size (window_size 1024, hop_size 320)</li>\n<li>Using recordings with either rating 0 or ratings &gt;= 2 and label smoothing of 0.01</li>\n<li>More epochs</li>\n</ul>",
      "rawMarkdown": "Thanks to Kaggle, competition hosts, and fellow competitors for this very interesting competition. We joined this competition in the last month and had to work hard to understand the competition as neither @neomaoro or I have done anything with audio before this. Both @neomaoro and I worked equally hard on this competition.\n\n**TLDR**\n\nOur solution is based heavily on @philippsinger @christofhenkel and @ilu000's [2nd place solution ](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463) from last year's birdclef competition. We made some modifications to this pipeline, the most significant being resizing spectrograms to 256 * 512 which gave a boost of 0.01 in public and private leaderboard.\n\n**Submission Notebook**\nhttps://www.kaggle.com/code/vexxingbanana/18th-place-solution-0-77-private-0-81-public\n\n**Code Pipeline and Data Setup**\n\nWe heavily based our notebooks and python files for training and submitting on @julian3833's [training notebook](https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-train-0-66) and [inference notebook](https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-submit-0-66). We setup training in the cloud on 2 A40 GPUs.\n\n**Bird Classifier**\n\nOur models were very similar and trained the same as last year's 2nd place solution with 30 second clips of train_audio resized into 6 x 5 second parts. We also used primary and secondary labels as targets. For inference, we fed in 6 x 5 second snippets into the model.\n\nWe used the following backbones: eca_nfnet_l0, eca_nfnet_l1, tf_efficientnetv2_m_in21k, seresnext50_32x4d, and resnest50d_4s2x40d. Eca_nfnet_l0 and eca_nfnet_l1 achieved the best private leaderboard scores with 5 fold scores of 0.76 and 0.75 respectively. \n\nDuring training, we started with the 2nd place solution's training strategy then added some slight modifications. The modifications are as follows:\n\n- Epochs: 20 -> 25\n- Optimizer: Adam -> AdamW with learning rate 1e-3 and weight decay 1e-6\n- **Resizing: After making a spectrogram, the spectrogram was resized to 256 x 512 using torchvision.**\n- Background noise, label smoothing removed.\n\n**Validation**\n\nWe couldn't find a way to make a good validation pipeline based on soundscapes since there were no soundscapes with labels given for this competition. Therefore, we simply used validation loss and public leaderboard as validation. \n\n**Ensembling**\n\nThe ensembling of our models was pretty straightforward. We trained 5 fold models of each backbone and then averaged all of the models' predictions together. We tried out a few other ensembling methods such as purely voting and a combination of voting and averaging but found averaging all models to be the best performing. \n\n**What Did Not Work**\n\nThroughout the competition, we had ups and downs, many successes and some things that didn't work well. Here are some experiments which didn't work out for us:\n\n- Applying PCEN\n- Using 2021 Data  \n- Attention head\n- Adding augmentations such as pink noise, gaussian noise, etc.\n- SED models\n- Adjusting mel spectrograms values to increase their size (window_size 1024, hop_size 320)\n- Using recordings with either rating 0 or ratings >= 2 and label smoothing of 0.01\n- More epochs",
      "votes": null
    },
    {
      "id": "1800613",
      "postDate": "05/25/2022 04:51:23",
      "content": "<p>Thanks for your write up.</p>\n<p>I'm trying to understand the inference step from yours as well as the 2nd place solution for BC2021. How do you do inference on a particular 5s segment if the model input takes in 6x5s? Do you simply repeat the single 5s chunk x6 times?</p>",
      "rawMarkdown": "Thanks for your write up.\n\nI'm trying to understand the inference step from yours as well as the 2nd place solution for BC2021. How do you do inference on a particular 5s segment if the model input takes in 6x5s? Do you simply repeat the single 5s chunk x6 times?",
      "votes": null
    },
    {
      "id": "1801267",
      "postDate": "05/25/2022 15:10:28",
      "content": "<p>So basically the model takes in 6 * batch size 5 second snippets. Then during inference, it takes duration / 5 5 second snippets. Therefore in our model, during training the model takes in a tensor shaped [96, 1, 512, 256] and during inference, it takes in a tensor shape [12, 1, 512, 256]. </p>",
      "rawMarkdown": "So basically the model takes in 6 * batch size 5 second snippets. Then during inference, it takes duration / 5 5 second snippets. Therefore in our model, during training the model takes in a tensor shaped [96, 1, 512, 256] and during inference, it takes in a tensor shape [12, 1, 512, 256].",
      "votes": null
    },
    {
      "id": "1801605",
      "postDate": "05/26/2022 00:30:24",
      "content": "<p>Congrats for becoming a competition expert!</p>",
      "rawMarkdown": "Congrats for becoming a competition expert!",
      "votes": null
    },
    {
      "id": "1801609",
      "postDate": "05/26/2022 00:36:37",
      "content": "<p>Thanks!      </p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "1801813",
      "postDate": "05/26/2022 07:02:25",
      "content": "<p>Thanks for the response 🙂 I noticed that the model takes a pool across the time segments (at least according to the diagram given by the BC21 2nd place solution). I was wondering if you added any other steps to get a (12, n_classes) output instead of (, n_classes)</p>",
      "rawMarkdown": "Thanks for the response 🙂 I noticed that the model takes a pool across the time segments (at least according to the diagram given by the BC21 2nd place solution). I was wondering if you added any other steps to get a (12, n_classes) output instead of (, n_classes)",
      "votes": null
    },
    {
      "id": "1802143",
      "postDate": "05/26/2022 13:44:24",
      "content": "<p>No we didn't change anything in terms of architecture though we did mess around a little with the type of pooling but found the best pooling to be the BC21 2nd place pooling.</p>",
      "rawMarkdown": "No we didn't change anything in terms of architecture though we did mess around a little with the type of pooling but found the best pooling to be the BC21 2nd place pooling.",
      "votes": null
    },
    {
      "id": "1802257",
      "postDate": "05/26/2022 15:43:41",
      "content": "<p>Thanks so much! I re-read your response and I think I get it now ! :) In the inference stage you don't restitch/reshape the Tensor back into it's 30 (or 60) second form. The input for inference was: [bs * parts, feats, time, freq] instead of [bs, feats, time * parts, freq] used in training.</p>",
      "rawMarkdown": "Thanks so much! I re-read your response and I think I get it now ! :) In the inference stage you don't restitch/reshape the Tensor back into it's 30 (or 60) second form. The input for inference was: [bs * parts, feats, time, freq] instead of [bs, feats, time * parts, freq] used in training.",
      "votes": null
    },
    {
      "id": "1802264",
      "postDate": "05/26/2022 15:48:03",
      "content": "<p>Congratulations!!</p>",
      "rawMarkdown": "Congratulations!!",
      "votes": null
    },
    {
      "id": "1802559",
      "postDate": "05/26/2022 22:13:25",
      "content": "<p>Thanks! Your notebook helped us a lot.</p>",
      "rawMarkdown": "Thanks! Your notebook helped us a lot.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1800613,
      "author_name": "kevenr",
      "author_url": "",
      "post_date": "05/25/2022 04:51:23",
      "content": "<p>Thanks for your write up.</p>\n<p>I'm trying to understand the inference step from yours as well as the 2nd place solution for BC2021. How do you do inference on a particular 5s segment if the model input takes in 6x5s? Do you simply repeat the single 5s chunk x6 times?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1801267,
          "author_name": "vexxingbanana",
          "author_url": "",
          "post_date": "05/25/2022 15:10:28",
          "content": "<p>So basically the model takes in 6 * batch size 5 second snippets. Then during inference, it takes duration / 5 5 second snippets. Therefore in our model, during training the model takes in a tensor shaped [96, 1, 512, 256] and during inference, it takes in a tensor shape [12, 1, 512, 256]. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1801813,
          "author_name": "kevenr",
          "author_url": "",
          "post_date": "05/26/2022 07:02:25",
          "content": "<p>Thanks for the response 🙂 I noticed that the model takes a pool across the time segments (at least according to the diagram given by the BC21 2nd place solution). I was wondering if you added any other steps to get a (12, n_classes) output instead of (, n_classes)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1802143,
          "author_name": "vexxingbanana",
          "author_url": "",
          "post_date": "05/26/2022 13:44:24",
          "content": "<p>No we didn't change anything in terms of architecture though we did mess around a little with the type of pooling but found the best pooling to be the BC21 2nd place pooling.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1802257,
          "author_name": "kevenr",
          "author_url": "",
          "post_date": "05/26/2022 15:43:41",
          "content": "<p>Thanks so much! I re-read your response and I think I get it now ! :) In the inference stage you don't restitch/reshape the Tensor back into it's 30 (or 60) second form. The input for inference was: [bs * parts, feats, time, freq] instead of [bs, feats, time * parts, freq] used in training.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1801605,
      "author_name": "neomaoro",
      "author_url": "",
      "post_date": "05/26/2022 00:30:24",
      "content": "<p>Congrats for becoming a competition expert!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1801609,
          "author_name": "vexxingbanana",
          "author_url": "",
          "post_date": "05/26/2022 00:36:37",
          "content": "<p>Thanks!      </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1802264,
      "author_name": "julian3833",
      "author_url": "",
      "post_date": "05/26/2022 15:48:03",
      "content": "<p>Congratulations!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1802559,
          "author_name": "vexxingbanana",
          "author_url": "",
          "post_date": "05/26/2022 22:13:25",
          "content": "<p>Thanks! Your notebook helped us a lot.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1800468": "Thanks to Kaggle, competition hosts, and fellow competitors for this very interesting competition. We joined this competition in the last month and had to work hard to understand the competition as neither @neomaoro or I have done anything with audio before this. Both @neomaoro and I worked equally hard on this competition.\n\n**TLDR**\n\nOur solution is based heavily on @philippsinger @christofhenkel and @ilu000's [2nd place solution ](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463) from last year's birdclef competition. We made some modifications to this pipeline, the most significant being resizing spectrograms to 256 * 512 which gave a boost of 0.01 in public and private leaderboard.\n\n**Submission Notebook**\nhttps://www.kaggle.com/code/vexxingbanana/18th-place-solution-0-77-private-0-81-public\n\n**Code Pipeline and Data Setup**\n\nWe heavily based our notebooks and python files for training and submitting on @julian3833's [training notebook](https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-train-0-66) and [inference notebook](https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-submit-0-66). We setup training in the cloud on 2 A40 GPUs.\n\n**Bird Classifier**\n\nOur models were very similar and trained the same as last year's 2nd place solution with 30 second clips of train_audio resized into 6 x 5 second parts. We also used primary and secondary labels as targets. For inference, we fed in 6 x 5 second snippets into the model.\n\nWe used the following backbones: eca_nfnet_l0, eca_nfnet_l1, tf_efficientnetv2_m_in21k, seresnext50_32x4d, and resnest50d_4s2x40d. Eca_nfnet_l0 and eca_nfnet_l1 achieved the best private leaderboard scores with 5 fold scores of 0.76 and 0.75 respectively. \n\nDuring training, we started with the 2nd place solution's training strategy then added some slight modifications. The modifications are as follows:\n\n- Epochs: 20 -> 25\n- Optimizer: Adam -> AdamW with learning rate 1e-3 and weight decay 1e-6\n- **Resizing: After making a spectrogram, the spectrogram was resized to 256 x 512 using torchvision.**\n- Background noise, label smoothing removed.\n\n**Validation**\n\nWe couldn't find a way to make a good validation pipeline based on soundscapes since there were no soundscapes with labels given for this competition. Therefore, we simply used validation loss and public leaderboard as validation. \n\n**Ensembling**\n\nThe ensembling of our models was pretty straightforward. We trained 5 fold models of each backbone and then averaged all of the models' predictions together. We tried out a few other ensembling methods such as purely voting and a combination of voting and averaging but found averaging all models to be the best performing. \n\n**What Did Not Work**\n\nThroughout the competition, we had ups and downs, many successes and some things that didn't work well. Here are some experiments which didn't work out for us:\n\n- Applying PCEN\n- Using 2021 Data  \n- Attention head\n- Adding augmentations such as pink noise, gaussian noise, etc.\n- SED models\n- Adjusting mel spectrograms values to increase their size (window_size 1024, hop_size 320)\n- Using recordings with either rating 0 or ratings >= 2 and label smoothing of 0.01\n- More epochs",
    "1800613": "Thanks for your write up.\n\nI'm trying to understand the inference step from yours as well as the 2nd place solution for BC2021. How do you do inference on a particular 5s segment if the model input takes in 6x5s? Do you simply repeat the single 5s chunk x6 times?",
    "1801267": "So basically the model takes in 6 * batch size 5 second snippets. Then during inference, it takes duration / 5 5 second snippets. Therefore in our model, during training the model takes in a tensor shaped [96, 1, 512, 256] and during inference, it takes in a tensor shape [12, 1, 512, 256].",
    "1801605": "Congrats for becoming a competition expert!",
    "1801609": "Thanks!",
    "1801813": "Thanks for the response 🙂 I noticed that the model takes a pool across the time segments (at least according to the diagram given by the BC21 2nd place solution). I was wondering if you added any other steps to get a (12, n_classes) output instead of (, n_classes)",
    "1802143": "No we didn't change anything in terms of architecture though we did mess around a little with the type of pooling but found the best pooling to be the BC21 2nd place pooling.",
    "1802257": "Thanks so much! I re-read your response and I think I get it now ! :) In the inference stage you don't restitch/reshape the Tensor back into it's 30 (or 60) second form. The input for inference was: [bs * parts, feats, time, freq] instead of [bs, feats, time * parts, freq] used in training.",
    "1802264": "Congratulations!!",
    "1802559": "Thanks! Your notebook helped us a lot."
  },
  "source": "meta"
}