{
  "id": 243756,
  "title": "8th place writeup",
  "url": "/competitions/birdclef-2021/writeups/kdl-8th-place-writeup",
  "author_name": "",
  "post_date": "2021-06-03T23:01:22.837Z",
  "votes": 43,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Thanks to Kaggle and the hosts to hold an interesting competition. This competition was quite similar to the one held last year, Cornell Birdcall Identification, but was improved also from the last one.</p>\n<p>Along with this post, we also made the <a href=\"https://www.kaggle.com/hidehisaarai1213/birdclef2021-infer-between-chunk\" target=\"_blank\">inference notebook</a> public, so please refer to it if you are interested in the details of model and/or post-processing.</p>\n<h2>Solution summary</h2>\n<p>Our solution is weighted blending of several CNNs with classification heads for SED task. All the backbone models are from EfficientNet families and EfficientNet V2 families. We train our models with either 20s chunks or 5s chunks. For inference, we predict on 40sec chunks and apply two thresholds.</p>\n<h2>Preprocess</h2>\n<p>No special things - used log-melspectrogram (We used torchaudio). Most of the models used<br>\nn_fft=2048, sr=32000, hop_length=512, n_mels=256. Only two models used n_mels=320.</p>\n<h2>Models</h2>\n<p>We used 29 models with SED heads. They are almost the same architecture as that used in <a href=\"https://www.kaggle.com/hidehisaarai1213/pytorch-inference-birdclef2021-starter\" target=\"_blank\">my public Notebook</a>.</p>\n<h2>Augmentations</h2>\n<p>We used Gaussian Noise, Pink Noise, and Volume Control for all the models. We found that PitchShift works well, however it puts a lot of load on CPU, therefore only a few models in our ensemble use PitchShift. We also found mixup works, but with the similar reason, we didn't use this for our final models.</p>\n<h2>Training</h2>\n<p>We use BCEFocal2WayLoss which was introduced in <a href=\"https://www.kaggle.com/hidehisaarai1213/pytorch-training-birdclef2021-starter\" target=\"_blank\">my training Notebook</a>. We use secondary_labels along with primary_label and weighted the same. Some of our models are used 2stage strategy - in the first stage we train our model in a usual manner. Then we make the models predict on out-of-fold clips and get soft pseudo-labels on the whole training data set. In the second stage, we use the pseudo labels to mask the loss term for some of the classes.</p>\n<h2>Inference</h2>\n<p>We make the models infer on 40s chunks. Basically, the longer the input chunk, the better. Of course, score saturate when it gets longer, so we chose 40s which was in a good balance between memory consumption and performance. As a kind of TTA, we also make the model predict for the chunks that overlaps with the previous chunk. For example, our models first predict for 0s - 40s, then we feed the models to predict for 20s - 60s, and so on. </p>\n<h2>Ensemble</h2>\n<p>After we get predictions for 40s chunk we apply weighted blending. The weights are decided using optuna. When we search for these weights, the thresholds (NOTE: we used two different thresholds. Described in the section behind) are also included as hyperparameters.</p>\n<h2>Postprocessing (Thresholds search &amp; masking)</h2>\n<p>We used two different kind of thresholds - call threshold and nocall threshold. All the predictions that exceed call threshold is labeled as positive. If all the predictions in the 5s chunk does not surpass nocall threshold, we always add <code>nocall</code> label. Therefore, for some prediction we might get prediction like \"moudov nocall\". To search for this threshold we simply used grid search.<br>\nWe then mask out some species that should not exist in that site. This is a rule based post-processing.</p>\n<h2>Validation</h2>\n<p>To match the nocall ratio of the public, we used<br>\noverall_f1 = 0.54 * nocall_f1 + 0.46 * call_f1<br>\nIt turns out that CPMP also used this validation scheme.</p>\n<h2>Things that didn't work for us</h2>\n<p>Quite a lot of things. We tried many backbones, augmentations. We also searched for better ways to use secondary_labels, but in the end, we found treat the secondary_labels as the same as primary_label is the best. We also worked for using the location information. We tried using location information by OHE encoding but it didn't help. We also tried made a table that shows which species might appear in the four locations(SNE/SSW/COR/COL) and used that to train location-wise model, but it wasn't effective also.</p>",
  "messages": [
    {
      "id": "1334943",
      "postDate": "06/03/2021 22:58:55",
      "content": "<p>Thanks to Kaggle and the hosts to hold an interesting competition. This competition was quite similar to the one held last year, Cornell Birdcall Identification, but was improved also from the last one.</p>\n<p>Along with this post, we also made the <a href=\"https://www.kaggle.com/hidehisaarai1213/birdclef2021-infer-between-chunk\" target=\"_blank\">inference notebook</a> public, so please refer to it if you are interested in the details of model and/or post-processing.</p>\n<h2>Solution summary</h2>\n<p>Our solution is weighted blending of several CNNs with classification heads for SED task. All the backbone models are from EfficientNet families and EfficientNet V2 families. We train our models with either 20s chunks or 5s chunks. For inference, we predict on 40sec chunks and apply two thresholds.</p>\n<h2>Preprocess</h2>\n<p>No special things - used log-melspectrogram (We used torchaudio). Most of the models used<br>\nn_fft=2048, sr=32000, hop_length=512, n_mels=256. Only two models used n_mels=320.</p>\n<h2>Models</h2>\n<p>We used 29 models with SED heads. They are almost the same architecture as that used in <a href=\"https://www.kaggle.com/hidehisaarai1213/pytorch-inference-birdclef2021-starter\" target=\"_blank\">my public Notebook</a>.</p>\n<h2>Augmentations</h2>\n<p>We used Gaussian Noise, Pink Noise, and Volume Control for all the models. We found that PitchShift works well, however it puts a lot of load on CPU, therefore only a few models in our ensemble use PitchShift. We also found mixup works, but with the similar reason, we didn't use this for our final models.</p>\n<h2>Training</h2>\n<p>We use BCEFocal2WayLoss which was introduced in <a href=\"https://www.kaggle.com/hidehisaarai1213/pytorch-training-birdclef2021-starter\" target=\"_blank\">my training Notebook</a>. We use secondary_labels along with primary_label and weighted the same. Some of our models are used 2stage strategy - in the first stage we train our model in a usual manner. Then we make the models predict on out-of-fold clips and get soft pseudo-labels on the whole training data set. In the second stage, we use the pseudo labels to mask the loss term for some of the classes.</p>\n<h2>Inference</h2>\n<p>We make the models infer on 40s chunks. Basically, the longer the input chunk, the better. Of course, score saturate when it gets longer, so we chose 40s which was in a good balance between memory consumption and performance. As a kind of TTA, we also make the model predict for the chunks that overlaps with the previous chunk. For example, our models first predict for 0s - 40s, then we feed the models to predict for 20s - 60s, and so on. </p>\n<h2>Ensemble</h2>\n<p>After we get predictions for 40s chunk we apply weighted blending. The weights are decided using optuna. When we search for these weights, the thresholds (NOTE: we used two different thresholds. Described in the section behind) are also included as hyperparameters.</p>\n<h2>Postprocessing (Thresholds search &amp; masking)</h2>\n<p>We used two different kind of thresholds - call threshold and nocall threshold. All the predictions that exceed call threshold is labeled as positive. If all the predictions in the 5s chunk does not surpass nocall threshold, we always add <code>nocall</code> label. Therefore, for some prediction we might get prediction like \"moudov nocall\". To search for this threshold we simply used grid search.<br>\nWe then mask out some species that should not exist in that site. This is a rule based post-processing.</p>\n<h2>Validation</h2>\n<p>To match the nocall ratio of the public, we used<br>\noverall_f1 = 0.54 * nocall_f1 + 0.46 * call_f1<br>\nIt turns out that CPMP also used this validation scheme.</p>\n<h2>Things that didn't work for us</h2>\n<p>Quite a lot of things. We tried many backbones, augmentations. We also searched for better ways to use secondary_labels, but in the end, we found treat the secondary_labels as the same as primary_label is the best. We also worked for using the location information. We tried using location information by OHE encoding but it didn't help. We also tried made a table that shows which species might appear in the four locations(SNE/SSW/COR/COL) and used that to train location-wise model, but it wasn't effective also.</p>",
      "rawMarkdown": "Thanks to Kaggle and the hosts to hold an interesting competition. This competition was quite similar to the one held last year, Cornell Birdcall Identification, but was improved also from the last one.\n\nAlong with this post, we also made the [inference notebook](https://www.kaggle.com/hidehisaarai1213/birdclef2021-infer-between-chunk) public, so please refer to it if you are interested in the details of model and/or post-processing.\n\n## Solution summary\n\nOur solution is weighted blending of several CNNs with classification heads for SED task. All the backbone models are from EfficientNet families and EfficientNet V2 families. We train our models with either 20s chunks or 5s chunks. For inference, we predict on 40sec chunks and apply two thresholds.\n\n## Preprocess\n\nNo special things - used log-melspectrogram (We used torchaudio). Most of the models used\nn_fft=2048, sr=32000, hop_length=512, n_mels=256. Only two models used n_mels=320.\n\n## Models\n\nWe used 29 models with SED heads. They are almost the same architecture as that used in [my public Notebook](https://www.kaggle.com/hidehisaarai1213/pytorch-inference-birdclef2021-starter).\n\n## Augmentations\n\nWe used Gaussian Noise, Pink Noise, and Volume Control for all the models. We found that PitchShift works well, however it puts a lot of load on CPU, therefore only a few models in our ensemble use PitchShift. We also found mixup works, but with the similar reason, we didn't use this for our final models.\n\n## Training\n\nWe use BCEFocal2WayLoss which was introduced in [my training Notebook](https://www.kaggle.com/hidehisaarai1213/pytorch-training-birdclef2021-starter). We use secondary_labels along with primary_label and weighted the same. Some of our models are used 2stage strategy - in the first stage we train our model in a usual manner. Then we make the models predict on out-of-fold clips and get soft pseudo-labels on the whole training data set. In the second stage, we use the pseudo labels to mask the loss term for some of the classes.\n\n## Inference\n\nWe make the models infer on 40s chunks. Basically, the longer the input chunk, the better. Of course, score saturate when it gets longer, so we chose 40s which was in a good balance between memory consumption and performance. As a kind of TTA, we also make the model predict for the chunks that overlaps with the previous chunk. For example, our models first predict for 0s - 40s, then we feed the models to predict for 20s - 60s, and so on. \n\n## Ensemble\n\nAfter we get predictions for 40s chunk we apply weighted blending. The weights are decided using optuna. When we search for these weights, the thresholds (NOTE: we used two different thresholds. Described in the section behind) are also included as hyperparameters.\n\n## Postprocessing (Thresholds search & masking)\n\nWe used two different kind of thresholds - call threshold and nocall threshold. All the predictions that exceed call threshold is labeled as positive. If all the predictions in the 5s chunk does not surpass nocall threshold, we always add `nocall` label. Therefore, for some prediction we might get prediction like \"moudov nocall\". To search for this threshold we simply used grid search.\nWe then mask out some species that should not exist in that site. This is a rule based post-processing.\n\n## Validation\n\nTo match the nocall ratio of the public, we used\noverall_f1 = 0.54 * nocall_f1 + 0.46 * call_f1\nIt turns out that CPMP also used this validation scheme.\n\n## Things that didn't work for us\n\nQuite a lot of things. We tried many backbones, augmentations. We also searched for better ways to use secondary_labels, but in the end, we found treat the secondary_labels as the same as primary_label is the best. We also worked for using the location information. We tried using location information by OHE encoding but it didn't help. We also tried made a table that shows which species might appear in the four locations(SNE/SSW/COR/COL) and used that to train location-wise model, but it wasn't effective also.",
      "votes": null
    },
    {
      "id": "1334983",
      "postDate": "06/04/2021 00:13:38",
      "content": "<p>Thank you for sharing.</p>\n<p><a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> Congrats on Grand Master!!</p>",
      "rawMarkdown": "Thank you for sharing.\n\n@hidehisaarai1213 Congrats on Grand Master!!",
      "votes": null
    },
    {
      "id": "1335375",
      "postDate": "06/04/2021 07:18:46",
      "content": "<p>Thank you for sharing and <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> to becoming GM !! That's a huge milestone!</p>\n<p>Would you mind open-source your checkpoint dataset so that we can play around with it? Thanks in advance!</p>",
      "rawMarkdown": "Thank you for sharing and @hidehisaarai1213 to becoming GM !! That's a huge milestone!\n\nWould you mind open-source your checkpoint dataset so that we can play around with it? Thanks in advance!",
      "votes": null
    },
    {
      "id": "1336033",
      "postDate": "06/04/2021 15:39:46",
      "content": "<p><a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> Congrats on getting GM!</p>",
      "rawMarkdown": "hidehisaarai1213 Congrats on getting GM!",
      "votes": null
    },
    {
      "id": "1336123",
      "postDate": "06/04/2021 16:26:20",
      "content": "<p>Interesting, you had entries with both nocall and some birds? </p>\n<p>It makes sense when the model is unsure.</p>\n<p>Congrats on the result and your GM title.</p>",
      "rawMarkdown": "Interesting, you had entries with both nocall and some birds? \n\nIt makes sense when the model is unsure.\n\nCongrats on the result and your GM title.",
      "votes": null
    },
    {
      "id": "1336469",
      "postDate": "06/05/2021 00:24:57",
      "content": "<p>Thank you!!!</p>",
      "rawMarkdown": "Thank you!!!",
      "votes": null
    },
    {
      "id": "1336471",
      "postDate": "06/05/2021 00:26:58",
      "content": "<p>Thank you!</p>\n<blockquote>\n  <p>Would you mind open-source your checkpoint dataset so that we can play around with it?</p>\n</blockquote>\n<p>I'll ask for this to my teammate. It's a dataset of him.</p>",
      "rawMarkdown": "Thank you!\n> Would you mind open-source your checkpoint dataset so that we can play around with it?\n\nI'll ask for this to my teammate. It's a dataset of him.",
      "votes": null
    },
    {
      "id": "1336472",
      "postDate": "06/05/2021 00:27:17",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "1336474",
      "postDate": "06/05/2021 00:30:02",
      "content": "<blockquote>\n  <p>you had entries with both nocall and some birds?</p>\n</blockquote>\n<p>Yes. I learned this technique from one of the participant of Cornell Birdcall Identification.</p>\n<blockquote>\n  <p>Congrats on the result and your GM title.</p>\n</blockquote>\n<p>Thank you!</p>",
      "rawMarkdown": "> you had entries with both nocall and some birds?\n\nYes. I learned this technique from one of the participant of Cornell Birdcall Identification.\n\n> Congrats on the result and your GM title.\n\nThank you!",
      "votes": null
    },
    {
      "id": "1345515",
      "postDate": "06/11/2021 16:17:02",
      "content": "<p>Just wanted to say thanks again to Hidehisa for sharing the open notebooks in the last competition; I think it really raised the bar in that competition, which has led to an even stronger competition this time around. <br>\nCheers!</p>",
      "rawMarkdown": "Just wanted to say thanks again to Hidehisa for sharing the open notebooks in the last competition; I think it really raised the bar in that competition, which has led to an even stronger competition this time around. \nCheers!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1334983,
      "author_name": "ttahara",
      "author_url": "",
      "post_date": "06/04/2021 00:13:38",
      "content": "<p>Thank you for sharing.</p>\n<p><a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> Congrats on Grand Master!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1336469,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "06/05/2021 00:24:57",
          "content": "<p>Thank you!!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335375,
      "author_name": "jonathanbesomi",
      "author_url": "",
      "post_date": "06/04/2021 07:18:46",
      "content": "<p>Thank you for sharing and <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> to becoming GM !! That's a huge milestone!</p>\n<p>Would you mind open-source your checkpoint dataset so that we can play around with it? Thanks in advance!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1336471,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "06/05/2021 00:26:58",
          "content": "<p>Thank you!</p>\n<blockquote>\n  <p>Would you mind open-source your checkpoint dataset so that we can play around with it?</p>\n</blockquote>\n<p>I'll ask for this to my teammate. It's a dataset of him.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1336033,
      "author_name": "taggatle",
      "author_url": "",
      "post_date": "06/04/2021 15:39:46",
      "content": "<p><a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> Congrats on getting GM!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1336472,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "06/05/2021 00:27:17",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1336123,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/04/2021 16:26:20",
      "content": "<p>Interesting, you had entries with both nocall and some birds? </p>\n<p>It makes sense when the model is unsure.</p>\n<p>Congrats on the result and your GM title.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1336474,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "06/05/2021 00:30:02",
          "content": "<blockquote>\n  <p>you had entries with both nocall and some birds?</p>\n</blockquote>\n<p>Yes. I learned this technique from one of the participant of Cornell Birdcall Identification.</p>\n<blockquote>\n  <p>Congrats on the result and your GM title.</p>\n</blockquote>\n<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1345515,
      "author_name": "tomdenton",
      "author_url": "",
      "post_date": "06/11/2021 16:17:02",
      "content": "<p>Just wanted to say thanks again to Hidehisa for sharing the open notebooks in the last competition; I think it really raised the bar in that competition, which has led to an even stronger competition this time around. <br>\nCheers!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1334943": "Thanks to Kaggle and the hosts to hold an interesting competition. This competition was quite similar to the one held last year, Cornell Birdcall Identification, but was improved also from the last one.\n\nAlong with this post, we also made the [inference notebook](https://www.kaggle.com/hidehisaarai1213/birdclef2021-infer-between-chunk) public, so please refer to it if you are interested in the details of model and/or post-processing.\n\n## Solution summary\n\nOur solution is weighted blending of several CNNs with classification heads for SED task. All the backbone models are from EfficientNet families and EfficientNet V2 families. We train our models with either 20s chunks or 5s chunks. For inference, we predict on 40sec chunks and apply two thresholds.\n\n## Preprocess\n\nNo special things - used log-melspectrogram (We used torchaudio). Most of the models used\nn_fft=2048, sr=32000, hop_length=512, n_mels=256. Only two models used n_mels=320.\n\n## Models\n\nWe used 29 models with SED heads. They are almost the same architecture as that used in [my public Notebook](https://www.kaggle.com/hidehisaarai1213/pytorch-inference-birdclef2021-starter).\n\n## Augmentations\n\nWe used Gaussian Noise, Pink Noise, and Volume Control for all the models. We found that PitchShift works well, however it puts a lot of load on CPU, therefore only a few models in our ensemble use PitchShift. We also found mixup works, but with the similar reason, we didn't use this for our final models.\n\n## Training\n\nWe use BCEFocal2WayLoss which was introduced in [my training Notebook](https://www.kaggle.com/hidehisaarai1213/pytorch-training-birdclef2021-starter). We use secondary_labels along with primary_label and weighted the same. Some of our models are used 2stage strategy - in the first stage we train our model in a usual manner. Then we make the models predict on out-of-fold clips and get soft pseudo-labels on the whole training data set. In the second stage, we use the pseudo labels to mask the loss term for some of the classes.\n\n## Inference\n\nWe make the models infer on 40s chunks. Basically, the longer the input chunk, the better. Of course, score saturate when it gets longer, so we chose 40s which was in a good balance between memory consumption and performance. As a kind of TTA, we also make the model predict for the chunks that overlaps with the previous chunk. For example, our models first predict for 0s - 40s, then we feed the models to predict for 20s - 60s, and so on. \n\n## Ensemble\n\nAfter we get predictions for 40s chunk we apply weighted blending. The weights are decided using optuna. When we search for these weights, the thresholds (NOTE: we used two different thresholds. Described in the section behind) are also included as hyperparameters.\n\n## Postprocessing (Thresholds search & masking)\n\nWe used two different kind of thresholds - call threshold and nocall threshold. All the predictions that exceed call threshold is labeled as positive. If all the predictions in the 5s chunk does not surpass nocall threshold, we always add `nocall` label. Therefore, for some prediction we might get prediction like \"moudov nocall\". To search for this threshold we simply used grid search.\nWe then mask out some species that should not exist in that site. This is a rule based post-processing.\n\n## Validation\n\nTo match the nocall ratio of the public, we used\noverall_f1 = 0.54 * nocall_f1 + 0.46 * call_f1\nIt turns out that CPMP also used this validation scheme.\n\n## Things that didn't work for us\n\nQuite a lot of things. We tried many backbones, augmentations. We also searched for better ways to use secondary_labels, but in the end, we found treat the secondary_labels as the same as primary_label is the best. We also worked for using the location information. We tried using location information by OHE encoding but it didn't help. We also tried made a table that shows which species might appear in the four locations(SNE/SSW/COR/COL) and used that to train location-wise model, but it wasn't effective also.",
    "1334983": "Thank you for sharing.\n\n@hidehisaarai1213 Congrats on Grand Master!!",
    "1335375": "Thank you for sharing and @hidehisaarai1213 to becoming GM !! That's a huge milestone!\n\nWould you mind open-source your checkpoint dataset so that we can play around with it? Thanks in advance!",
    "1336033": "hidehisaarai1213 Congrats on getting GM!",
    "1336123": "Interesting, you had entries with both nocall and some birds? \n\nIt makes sense when the model is unsure.\n\nCongrats on the result and your GM title.",
    "1336469": "Thank you!!!",
    "1336471": "Thank you!\n> Would you mind open-source your checkpoint dataset so that we can play around with it?\n\nI'll ask for this to my teammate. It's a dataset of him.",
    "1336472": "Thank you!",
    "1336474": "> you had entries with both nocall and some birds?\n\nYes. I learned this technique from one of the participant of Cornell Birdcall Identification.\n\n> Congrats on the result and your GM title.\n\nThank you!",
    "1345515": "Just wanted to say thanks again to Hidehisa for sharing the open notebooks in the last competition; I think it really raised the bar in that competition, which has led to an even stronger competition this time around. \nCheers!"
  },
  "source": "meta"
}