{
  "id": 183315,
  "title": "9th place solution",
  "url": "/competitions/birdsong-recognition/writeups/dee-is-a-bird-9th-place-solution",
  "author_name": "",
  "post_date": "2020-09-16T19:54:04.973Z",
  "votes": 8,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I would like to thank the hosts for this unique challenge. Congratulations to my teammate <a href=\"https://www.kaggle.com/canalici\" target=\"_blank\">@canalici</a> and to all competitors. It was indeed a very educative competition in the audio domain. Many thanks to <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> guidance through competition, and <a href=\"https://www.kaggle.com/doanquanvietnamca\" target=\"_blank\">@doanquanvietnamca</a> for GRU-SED and Dmytro Karabash for ideas.</p>\n<h2>Data Augmentation</h2>\n<p>Models have trained in both original and extended datasets. Augmentations applied in both waveform and mel-spectrogram level.</p>\n<ul>\n<li>Gaussian Noise</li>\n<li>SpecAug</li>\n</ul>\n<h2>Modeling</h2>\n<p>I have modified the SED model (PANN’s) and replace its inefficient backbone with a noisy-Efficientnet and further experimented with GRU’s, LSTM’s, and with Transformers for temporal modeling. We had a CNN backbone -&gt; a GRU layer -&gt; and attention layer in the final model. </p>\n<ul>\n<li>EfficientNet-B4 (Noisy Student) </li>\n<li>EfficientNet-B7 (Noisy Student)</li>\n<li>EfficientNet-B7 (Noisy Student)</li>\n</ul>\n<p>To ensemble different solutions, I have removed the attention layer for each model, kept pre-trained weights of the extracted part, and re-trained an attention layer from features extracted from 3 different models listed above.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4470174%2F00b6195649463cefa402d7d0097e17d1%2FUntitled-3.png?generation=1600242638521318&amp;alt=media\" alt=\"\"></p>\n<h2>Training</h2>\n<ul>\n<li>Batch size of 32 for B4, and 8 for B7 models (single GPU)</li>\n<li>No mixup :( </li>\n<li>BCELoss </li>\n<li>AdamW with Cosine Anneal</li>\n<li>5 seconds of audio clips (501, 64) (scaling mel_bins into 224, 300 would really have helped, but very expensive)</li>\n<li>Pre-training on primary labels, fine tuning with secondary labels</li>\n</ul>\n<p>For validation, we have hand-labeled no-calls into gt_birdclef2020_validation_data and excluded irrelevant species, which provided a chance to test the algorithm in the wild. </p>\n<h3>Possible further work</h3>\n<p>Pre-trained models proven self to be leverage in many knowledge transfer tasks. It is tough not to over-fit our classifier, especially in this competition, where we had a few audio clips with very noisy labels. Hidehisa Arai pointed out PANN’s(one of the largest pre-trained models in the audio domain) for this issue, their CNN backbone was less potent than lighter alternatives. We have used a firm CNN backbone to overcome this issue, pre-trained on a large corpus of images (Noisy Student, Efficientnet). However, it is possible to extract mel-spectrogram encoder/decoder parts from very famous text-to-speech, speech conversation (Tacotron, Glow TTS…) algorithms that trained on a relatively larger corpus. It could be beneficial to adapt successfully pre-trained models from the Audio domain, fine-tuning it with all the bird data we have, then applying a noisy-student training scheme. </p>",
  "messages": [
    {
      "id": "1012650",
      "postDate": "09/16/2020 07:54:28",
      "content": "<p>I would like to thank the hosts for this unique challenge. Congratulations to my teammate <a href=\"https://www.kaggle.com/canalici\" target=\"_blank\">@canalici</a> and to all competitors. It was indeed a very educative competition in the audio domain. Many thanks to <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> guidance through competition, and <a href=\"https://www.kaggle.com/doanquanvietnamca\" target=\"_blank\">@doanquanvietnamca</a> for GRU-SED and Dmytro Karabash for ideas.</p>\n<h2>Data Augmentation</h2>\n<p>Models have trained in both original and extended datasets. Augmentations applied in both waveform and mel-spectrogram level.</p>\n<ul>\n<li>Gaussian Noise</li>\n<li>SpecAug</li>\n</ul>\n<h2>Modeling</h2>\n<p>I have modified the SED model (PANN’s) and replace its inefficient backbone with a noisy-Efficientnet and further experimented with GRU’s, LSTM’s, and with Transformers for temporal modeling. We had a CNN backbone -&gt; a GRU layer -&gt; and attention layer in the final model. </p>\n<ul>\n<li>EfficientNet-B4 (Noisy Student) </li>\n<li>EfficientNet-B7 (Noisy Student)</li>\n<li>EfficientNet-B7 (Noisy Student)</li>\n</ul>\n<p>To ensemble different solutions, I have removed the attention layer for each model, kept pre-trained weights of the extracted part, and re-trained an attention layer from features extracted from 3 different models listed above.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4470174%2F00b6195649463cefa402d7d0097e17d1%2FUntitled-3.png?generation=1600242638521318&amp;alt=media\" alt=\"\"></p>\n<h2>Training</h2>\n<ul>\n<li>Batch size of 32 for B4, and 8 for B7 models (single GPU)</li>\n<li>No mixup :( </li>\n<li>BCELoss </li>\n<li>AdamW with Cosine Anneal</li>\n<li>5 seconds of audio clips (501, 64) (scaling mel_bins into 224, 300 would really have helped, but very expensive)</li>\n<li>Pre-training on primary labels, fine tuning with secondary labels</li>\n</ul>\n<p>For validation, we have hand-labeled no-calls into gt_birdclef2020_validation_data and excluded irrelevant species, which provided a chance to test the algorithm in the wild. </p>\n<h3>Possible further work</h3>\n<p>Pre-trained models proven self to be leverage in many knowledge transfer tasks. It is tough not to over-fit our classifier, especially in this competition, where we had a few audio clips with very noisy labels. Hidehisa Arai pointed out PANN’s(one of the largest pre-trained models in the audio domain) for this issue, their CNN backbone was less potent than lighter alternatives. We have used a firm CNN backbone to overcome this issue, pre-trained on a large corpus of images (Noisy Student, Efficientnet). However, it is possible to extract mel-spectrogram encoder/decoder parts from very famous text-to-speech, speech conversation (Tacotron, Glow TTS…) algorithms that trained on a relatively larger corpus. It could be beneficial to adapt successfully pre-trained models from the Audio domain, fine-tuning it with all the bird data we have, then applying a noisy-student training scheme. </p>",
      "rawMarkdown": "I would like to thank the hosts for this unique challenge. Congratulations to my teammate @canalici and to all competitors. It was indeed a very educative competition in the audio domain. Many thanks to @hidehisaarai1213 guidance through competition, and @doanquanvietnamca for GRU-SED and Dmytro Karabash for ideas.\n\n##Data Augmentation\nModels have trained in both original and extended datasets. Augmentations applied in both waveform and mel-spectrogram level.\n\n- Gaussian Noise\n- SpecAug\n\n##Modeling\nI have modified the SED model (PANN’s) and replace its inefficient backbone with a noisy-Efficientnet and further experimented with GRU’s, LSTM’s, and with Transformers for temporal modeling. We had a CNN backbone -> a GRU layer -> and attention layer in the final model. \n\n- EfficientNet-B4 (Noisy Student) \n- EfficientNet-B7 (Noisy Student)\n- EfficientNet-B7 (Noisy Student)\n\nTo ensemble different solutions, I have removed the attention layer for each model, kept pre-trained weights of the extracted part, and re-trained an attention layer from features extracted from 3 different models listed above.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4470174%2F00b6195649463cefa402d7d0097e17d1%2FUntitled-3.png?generation=1600242638521318&alt=media)\n\n##Training\n- Batch size of 32 for B4, and 8 for B7 models (single GPU)\n- No mixup :( \n- BCELoss \n- AdamW with Cosine Anneal\n- 5 seconds of audio clips (501, 64) (scaling mel_bins into 224, 300 would really have helped, but very expensive)\n- Pre-training on primary labels, fine tuning with secondary labels\n\nFor validation, we have hand-labeled no-calls into gt_birdclef2020_validation_data and excluded irrelevant species, which provided a chance to test the algorithm in the wild. \n \n \n###Possible further work\nPre-trained models proven self to be leverage in many knowledge transfer tasks. It is tough not to over-fit our classifier, especially in this competition, where we had a few audio clips with very noisy labels. Hidehisa Arai pointed out PANN’s(one of the largest pre-trained models in the audio domain) for this issue, their CNN backbone was less potent than lighter alternatives. We have used a firm CNN backbone to overcome this issue, pre-trained on a large corpus of images (Noisy Student, Efficientnet). However, it is possible to extract mel-spectrogram encoder/decoder parts from very famous text-to-speech, speech conversation (Tacotron, Glow TTS...) algorithms that trained on a relatively larger corpus. It could be beneficial to adapt successfully pre-trained models from the Audio domain, fine-tuning it with all the bird data we have, then applying a noisy-student training scheme.",
      "votes": null
    },
    {
      "id": "1012677",
      "postDate": "09/16/2020 08:11:03",
      "content": "<p>thanks so much! any reason for choosing Noisy Student as the pre-trained weight instead of the original one?<br>\nIs it by experiment CV/LB result?</p>",
      "rawMarkdown": "thanks so much! any reason for choosing Noisy Student as the pre-trained weight instead of the original one?\nIs it by experiment CV/LB result?",
      "votes": null
    },
    {
      "id": "1012690",
      "postDate": "09/16/2020 08:16:50",
      "content": "<p>The reason I would say is extended pre-training corpus and smarter optimization. You can review this table for comparison. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4470174%2F58cbe63ae9d32ba8334831b8d4007969%2F11.png?generation=1600244025433433&amp;alt=media\" alt=\"\"></p>\n<p>My original plan was to stack noisy effnet-l2 models, though I couldn't do it on single GPU 😅</p>",
      "rawMarkdown": "The reason I would say is extended pre-training corpus and smarter optimization. You can review this table for comparison. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4470174%2F58cbe63ae9d32ba8334831b8d4007969%2F11.png?generation=1600244025433433&alt=media)\n\nMy original plan was to stack noisy effnet-l2 models, though I couldn't do it on single GPU 😅",
      "votes": null
    },
    {
      "id": "1012734",
      "postDate": "09/16/2020 08:59:57",
      "content": "<p>Thank you for sharing your solution, congratulation👍</p>",
      "rawMarkdown": "Thank you for sharing your solution, congratulation👍",
      "votes": null
    },
    {
      "id": "1012751",
      "postDate": "09/16/2020 09:14:32",
      "content": "<p>Congrats on the solution and result!  From your picture it looks like GRU was better than LSTM or Transformer head.  Do you confirm?</p>",
      "rawMarkdown": "Congrats on the solution and result!  From your picture it looks like GRU was better than LSTM or Transformer head.  Do you confirm?",
      "votes": null
    },
    {
      "id": "1012756",
      "postDate": "09/16/2020 09:17:12",
      "content": "<p>Congrats on result and thanks for sharing summary of solution <a href=\"https://www.kaggle.com/realsleim\" target=\"_blank\">@realsleim</a> </p>",
      "rawMarkdown": "Congrats on result and thanks for sharing summary of solution @realsleim",
      "votes": null
    },
    {
      "id": "1012762",
      "postDate": "09/16/2020 09:21:56",
      "content": "<p>Although I had comparative scores on local validation, GRU outperformed it on LB. My idea was to handling longer sequences with Linformer (self-attention with linear complexity). I think it's my bad and with better initilization and better optimization, it could outperform GRU or LSTM with a large margin, yet sadly I didn't have time to test all ideas :( Thank you very much for the kind comment.</p>",
      "rawMarkdown": "Although I had comparative scores on local validation, GRU outperformed it on LB. My idea was to handling longer sequences with Linformer (self-attention with linear complexity). I think it's my bad and with better initilization and better optimization, it could outperform GRU or LSTM with a large margin, yet sadly I didn't have time to test all ideas :( Thank you very much for the kind comment.",
      "votes": null
    },
    {
      "id": "1012957",
      "postDate": "09/16/2020 12:20:30",
      "content": "<p><a href=\"https://www.kaggle.com/realsleim\" target=\"_blank\">@realsleim</a> congratulations for your 1st gold medal and well organised models report, is it a model dashboard. Curious to see your code organisation. is any plans to share your code ? </p>",
      "rawMarkdown": "realsleim congratulations for your 1st gold medal and well organised models report, is it a model dashboard. Curious to see your code organisation. is any plans to share your code ?",
      "votes": null
    },
    {
      "id": "1012968",
      "postDate": "09/16/2020 12:31:01",
      "content": "<p>Thank you very much. Unfortunately, I don't know what is \"model dashboard\" 😅 I am planning to share the code.</p>",
      "rawMarkdown": "Thank you very much. Unfortunately, I don't know what is \"model dashboard\" 😅 I am planning to share the code.",
      "votes": null
    },
    {
      "id": "1012981",
      "postDate": "09/16/2020 12:41:42",
      "content": "<p>the way you organised your results is a model dashboard i mean. Will wait for your code :)</p>",
      "rawMarkdown": "the way you organised your results is a model dashboard i mean. Will wait for your code :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1012677,
      "author_name": "fiyeroleung",
      "author_url": "",
      "post_date": "09/16/2020 08:11:03",
      "content": "<p>thanks so much! any reason for choosing Noisy Student as the pre-trained weight instead of the original one?<br>\nIs it by experiment CV/LB result?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1012690,
          "author_name": "realsleim",
          "author_url": "",
          "post_date": "09/16/2020 08:16:50",
          "content": "<p>The reason I would say is extended pre-training corpus and smarter optimization. You can review this table for comparison. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4470174%2F58cbe63ae9d32ba8334831b8d4007969%2F11.png?generation=1600244025433433&amp;alt=media\" alt=\"\"></p>\n<p>My original plan was to stack noisy effnet-l2 models, though I couldn't do it on single GPU 😅</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1012957,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "09/16/2020 12:20:30",
          "content": "<p><a href=\"https://www.kaggle.com/realsleim\" target=\"_blank\">@realsleim</a> congratulations for your 1st gold medal and well organised models report, is it a model dashboard. Curious to see your code organisation. is any plans to share your code ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1012968,
          "author_name": "realsleim",
          "author_url": "",
          "post_date": "09/16/2020 12:31:01",
          "content": "<p>Thank you very much. Unfortunately, I don't know what is \"model dashboard\" 😅 I am planning to share the code.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1012981,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "09/16/2020 12:41:42",
          "content": "<p>the way you organised your results is a model dashboard i mean. Will wait for your code :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1012734,
      "author_name": "",
      "author_url": "",
      "post_date": "09/16/2020 08:59:57",
      "content": "<p>Thank you for sharing your solution, congratulation👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1012751,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "09/16/2020 09:14:32",
      "content": "<p>Congrats on the solution and result!  From your picture it looks like GRU was better than LSTM or Transformer head.  Do you confirm?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1012762,
          "author_name": "realsleim",
          "author_url": "",
          "post_date": "09/16/2020 09:21:56",
          "content": "<p>Although I had comparative scores on local validation, GRU outperformed it on LB. My idea was to handling longer sequences with Linformer (self-attention with linear complexity). I think it's my bad and with better initilization and better optimization, it could outperform GRU or LSTM with a large margin, yet sadly I didn't have time to test all ideas :( Thank you very much for the kind comment.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1012756,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "09/16/2020 09:17:12",
      "content": "<p>Congrats on result and thanks for sharing summary of solution <a href=\"https://www.kaggle.com/realsleim\" target=\"_blank\">@realsleim</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1012650": "I would like to thank the hosts for this unique challenge. Congratulations to my teammate @canalici and to all competitors. It was indeed a very educative competition in the audio domain. Many thanks to @hidehisaarai1213 guidance through competition, and @doanquanvietnamca for GRU-SED and Dmytro Karabash for ideas.\n\n##Data Augmentation\nModels have trained in both original and extended datasets. Augmentations applied in both waveform and mel-spectrogram level.\n\n- Gaussian Noise\n- SpecAug\n\n##Modeling\nI have modified the SED model (PANN’s) and replace its inefficient backbone with a noisy-Efficientnet and further experimented with GRU’s, LSTM’s, and with Transformers for temporal modeling. We had a CNN backbone -> a GRU layer -> and attention layer in the final model. \n\n- EfficientNet-B4 (Noisy Student) \n- EfficientNet-B7 (Noisy Student)\n- EfficientNet-B7 (Noisy Student)\n\nTo ensemble different solutions, I have removed the attention layer for each model, kept pre-trained weights of the extracted part, and re-trained an attention layer from features extracted from 3 different models listed above.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4470174%2F00b6195649463cefa402d7d0097e17d1%2FUntitled-3.png?generation=1600242638521318&alt=media)\n\n##Training\n- Batch size of 32 for B4, and 8 for B7 models (single GPU)\n- No mixup :( \n- BCELoss \n- AdamW with Cosine Anneal\n- 5 seconds of audio clips (501, 64) (scaling mel_bins into 224, 300 would really have helped, but very expensive)\n- Pre-training on primary labels, fine tuning with secondary labels\n\nFor validation, we have hand-labeled no-calls into gt_birdclef2020_validation_data and excluded irrelevant species, which provided a chance to test the algorithm in the wild. \n \n \n###Possible further work\nPre-trained models proven self to be leverage in many knowledge transfer tasks. It is tough not to over-fit our classifier, especially in this competition, where we had a few audio clips with very noisy labels. Hidehisa Arai pointed out PANN’s(one of the largest pre-trained models in the audio domain) for this issue, their CNN backbone was less potent than lighter alternatives. We have used a firm CNN backbone to overcome this issue, pre-trained on a large corpus of images (Noisy Student, Efficientnet). However, it is possible to extract mel-spectrogram encoder/decoder parts from very famous text-to-speech, speech conversation (Tacotron, Glow TTS...) algorithms that trained on a relatively larger corpus. It could be beneficial to adapt successfully pre-trained models from the Audio domain, fine-tuning it with all the bird data we have, then applying a noisy-student training scheme.",
    "1012677": "thanks so much! any reason for choosing Noisy Student as the pre-trained weight instead of the original one?\nIs it by experiment CV/LB result?",
    "1012690": "The reason I would say is extended pre-training corpus and smarter optimization. You can review this table for comparison. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4470174%2F58cbe63ae9d32ba8334831b8d4007969%2F11.png?generation=1600244025433433&alt=media)\n\nMy original plan was to stack noisy effnet-l2 models, though I couldn't do it on single GPU 😅",
    "1012734": "Thank you for sharing your solution, congratulation👍",
    "1012751": "Congrats on the solution and result!  From your picture it looks like GRU was better than LSTM or Transformer head.  Do you confirm?",
    "1012756": "Congrats on result and thanks for sharing summary of solution @realsleim",
    "1012762": "Although I had comparative scores on local validation, GRU outperformed it on LB. My idea was to handling longer sequences with Linformer (self-attention with linear complexity). I think it's my bad and with better initilization and better optimization, it could outperform GRU or LSTM with a large margin, yet sadly I didn't have time to test all ideas :( Thank you very much for the kind comment.",
    "1012957": "realsleim congratulations for your 1st gold medal and well organised models report, is it a model dashboard. Curious to see your code organisation. is any plans to share your code ?",
    "1012968": "Thank you very much. Unfortunately, I don't know what is \"model dashboard\" 😅 I am planning to share the code.",
    "1012981": "the way you organised your results is a model dashboard i mean. Will wait for your code :)"
  },
  "source": "meta"
}