{
  "id": 327118,
  "title": "44th place solution",
  "url": "/competitions/birdclef-2022/writeups/kinosuke-44th-place-solution",
  "author_name": "",
  "post_date": "2022-05-26T03:36:06.063Z",
  "votes": 14,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Thank you for organizing such an interesting competition. It was a great learning experience for me.</p>\n<h2>Short Summary</h2>\n<p>My solution is SED, based on the one published by <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> in BirdCLEF2021.<br>\nThis ranking was achieved without modifying the model, but by devising a new data set and changing the backborn.<br>\nA big thank you to hidehisaarai1213 for publishing this wonderful notebook!</p>\n<p>train : <a href=\"https://www.kaggle.com/code/hidehisaarai1213/pytorch-training-birdclef2021-starter\" target=\"_blank\">https://www.kaggle.com/code/hidehisaarai1213/pytorch-training-birdclef2021-starter</a><br>\nsimple inference : <a href=\"https://www.kaggle.com/code/hidehisaarai1213/pytorch-inference-birdclef2021-starter\" target=\"_blank\">https://www.kaggle.com/code/hidehisaarai1213/pytorch-inference-birdclef2021-starter</a><br>\ninfer between chunk : <a href=\"https://www.kaggle.com/code/hidehisaarai1213/birdclef2021-infer-between-chunk\" target=\"_blank\">https://www.kaggle.com/code/hidehisaarai1213/birdclef2021-infer-between-chunk</a></p>\n<h2>Preprocessing</h2>\n<p>As you all know, the lable this time was Imbalanced, so I adopted the approach of acquiring the sound source multiple times if the length of the sound source was short.  <br>\nI have expressed the number of labels on three levels as follows.  <br>\nFew tripled and VERY FEW quadrupled their data and studied.  <br>\nAs I will explain later, the data used for training is a section of the sound source(20sec), so I hypothesized that it could be used multiple times by changing the starting point.  </p>\n<p><img src=\"https://user-images.githubusercontent.com/55369709/170328390-b056fc7e-a580-4497-9b77-71b6334f90b3.png\" alt=\"\"></p>\n<p>MANY : 'skylar', 'houfin', 'jabwar', 'warwhe1', 'yefcan', 'apapan','iiwi', 'hawcre'  <br>\nFEW : 'hawama', 'omao', 'barpet', 'akiapo', 'elepai', 'aniani'  <br>\nVERY FEW : 'hawgoo', 'ercfra', 'hawhaw', 'hawpet1', 'puaioh', 'crehon','maupar'  </p>\n<h2>Augmentation</h2>\n<p>Weak augmentation for MANY and strong augmetation for FEW. Strong means that it transforms with a high probability.  </p>\n<pre><code>augment_strong = Compose([\n    AddGaussianNoise(min_amplitude=0.001, max_amplitude=0.015, p=0.5),\n    AddGaussianSNR(p=0.5),\n    Gain(min_gain_in_db=-12, max_gain_in_db=12, p=0.5),\n    TimeStretch(min_rate=0.8, max_rate=1.25, p=0.5),\n    AddBackgroundNoise(sounds_path=noise_dir, min_snr_in_db=3, max_snr_in_db=30, p=0.5),\n    AddShortNoises(noise_dir, p=0.5),\n    PitchShift(min_semitones=-4, max_semitones=4, p=0.5),\n    Shift(min_fraction=-0.5, max_fraction=0.5, p=0.5),\n])\n\naugment_weak = Compose([\n    AddGaussianNoise(min_amplitude=0.001, max_amplitude=0.015, p=0.2),\n    AddGaussianSNR(p=0.2),\n    Gain(min_gain_in_db=-12, max_gain_in_db=12, p=0.2),\n    TimeStretch(min_rate=0.8, max_rate=1.25, p=0.2),\n    AddBackgroundNoise(sounds_path=noise_dir, min_snr_in_db=3, max_snr_in_db=30, p=0.2),\n    AddShortNoises(noise_dir, p=0.2),\n    PitchShift(min_semitones=-4, max_semitones=4, p=0.2),\n    Shift(min_fraction=-0.5, max_fraction=0.5, p=0.2),\n])\n</code></pre>\n<h2>Training</h2>\n<p>First of all, if I run the above hidehisaarai1213's notebook without preprocess with backbone as tf_efficientnet_b0_ns, I get 0.6406 for public and 0.6843 for private.    </p>\n<ul>\n<li>epoch -&gt; 30</li>\n<li>loss -&gt; BCEFocal2WayLoss</li>\n<li>backbone -&gt; tf_efficientnet_b0_ns</li>\n<li>inf threshold -&gt; 0.02</li>\n</ul>\n<p>From here, adding the above preprocess and running it, public was 0.7025 and private was 0.6916. Since public went up significantly, I determined that the strategy of increasing FEW and VERY FEW was so effective.  </p>\n<p>Adding a second label further increased the score from this point. 0.7053 for public, 0.6871 for private.</p>\n<p>That I tried but it did not lead to an increase in PUBLIC scores. but I did use it in an ensemble.</p>\n<ul>\n<li>more increase FEW and VERY FEW. Increased FEW by 4x and VERY FEW by 7x.  </li>\n<li>Increased train sound source from 20 sec to 40 sec.<ul>\n<li>I thought the longer section would contain more call, but it didn't WORK.</li></ul></li>\n<li>Change backbone to tf_efficientnet_b3_ns.</li>\n</ul>\n<p>What worked best.  </p>\n<ul>\n<li>\"nocall\" was added to label and increased from 152 to 153 classes.<ul>\n<li>nocall data found from here <a href=\"https://www.kaggle.com/datasets/kami634/ff1010bird-duration10\" target=\"_blank\">https://www.kaggle.com/datasets/kami634/ff1010bird-duration10</a></li>\n<li>This brings the score up to private 0.7473 public 0.7572</li></ul></li>\n</ul>\n<h2>Choosing subs</h2>\n<p>All scores above are obtained from fold0 only. The combination of fold and weight, I reached public 0.7730 private 0.7600, but I could not choose this sub because I considered that this was overfitting to public, using only fold0 and fold2 out of 5 unfolds.  <br>\nI chose a sub that mixed all folds equally to try to reduce overfitting to the public.  </p>",
  "messages": [
    {
      "id": "1801404",
      "postDate": "05/25/2022 17:41:43",
      "content": "<p>Thank you for organizing such an interesting competition. It was a great learning experience for me.</p>\n<h2>Short Summary</h2>\n<p>My solution is SED, based on the one published by <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> in BirdCLEF2021.<br>\nThis ranking was achieved without modifying the model, but by devising a new data set and changing the backborn.<br>\nA big thank you to hidehisaarai1213 for publishing this wonderful notebook!</p>\n<p>train : <a href=\"https://www.kaggle.com/code/hidehisaarai1213/pytorch-training-birdclef2021-starter\" target=\"_blank\">https://www.kaggle.com/code/hidehisaarai1213/pytorch-training-birdclef2021-starter</a><br>\nsimple inference : <a href=\"https://www.kaggle.com/code/hidehisaarai1213/pytorch-inference-birdclef2021-starter\" target=\"_blank\">https://www.kaggle.com/code/hidehisaarai1213/pytorch-inference-birdclef2021-starter</a><br>\ninfer between chunk : <a href=\"https://www.kaggle.com/code/hidehisaarai1213/birdclef2021-infer-between-chunk\" target=\"_blank\">https://www.kaggle.com/code/hidehisaarai1213/birdclef2021-infer-between-chunk</a></p>\n<h2>Preprocessing</h2>\n<p>As you all know, the lable this time was Imbalanced, so I adopted the approach of acquiring the sound source multiple times if the length of the sound source was short.  <br>\nI have expressed the number of labels on three levels as follows.  <br>\nFew tripled and VERY FEW quadrupled their data and studied.  <br>\nAs I will explain later, the data used for training is a section of the sound source(20sec), so I hypothesized that it could be used multiple times by changing the starting point.  </p>\n<p><img src=\"https://user-images.githubusercontent.com/55369709/170328390-b056fc7e-a580-4497-9b77-71b6334f90b3.png\" alt=\"\"></p>\n<p>MANY : 'skylar', 'houfin', 'jabwar', 'warwhe1', 'yefcan', 'apapan','iiwi', 'hawcre'  <br>\nFEW : 'hawama', 'omao', 'barpet', 'akiapo', 'elepai', 'aniani'  <br>\nVERY FEW : 'hawgoo', 'ercfra', 'hawhaw', 'hawpet1', 'puaioh', 'crehon','maupar'  </p>\n<h2>Augmentation</h2>\n<p>Weak augmentation for MANY and strong augmetation for FEW. Strong means that it transforms with a high probability.  </p>\n<pre><code>augment_strong = Compose([\n    AddGaussianNoise(min_amplitude=0.001, max_amplitude=0.015, p=0.5),\n    AddGaussianSNR(p=0.5),\n    Gain(min_gain_in_db=-12, max_gain_in_db=12, p=0.5),\n    TimeStretch(min_rate=0.8, max_rate=1.25, p=0.5),\n    AddBackgroundNoise(sounds_path=noise_dir, min_snr_in_db=3, max_snr_in_db=30, p=0.5),\n    AddShortNoises(noise_dir, p=0.5),\n    PitchShift(min_semitones=-4, max_semitones=4, p=0.5),\n    Shift(min_fraction=-0.5, max_fraction=0.5, p=0.5),\n])\n\naugment_weak = Compose([\n    AddGaussianNoise(min_amplitude=0.001, max_amplitude=0.015, p=0.2),\n    AddGaussianSNR(p=0.2),\n    Gain(min_gain_in_db=-12, max_gain_in_db=12, p=0.2),\n    TimeStretch(min_rate=0.8, max_rate=1.25, p=0.2),\n    AddBackgroundNoise(sounds_path=noise_dir, min_snr_in_db=3, max_snr_in_db=30, p=0.2),\n    AddShortNoises(noise_dir, p=0.2),\n    PitchShift(min_semitones=-4, max_semitones=4, p=0.2),\n    Shift(min_fraction=-0.5, max_fraction=0.5, p=0.2),\n])\n</code></pre>\n<h2>Training</h2>\n<p>First of all, if I run the above hidehisaarai1213's notebook without preprocess with backbone as tf_efficientnet_b0_ns, I get 0.6406 for public and 0.6843 for private.    </p>\n<ul>\n<li>epoch -&gt; 30</li>\n<li>loss -&gt; BCEFocal2WayLoss</li>\n<li>backbone -&gt; tf_efficientnet_b0_ns</li>\n<li>inf threshold -&gt; 0.02</li>\n</ul>\n<p>From here, adding the above preprocess and running it, public was 0.7025 and private was 0.6916. Since public went up significantly, I determined that the strategy of increasing FEW and VERY FEW was so effective.  </p>\n<p>Adding a second label further increased the score from this point. 0.7053 for public, 0.6871 for private.</p>\n<p>That I tried but it did not lead to an increase in PUBLIC scores. but I did use it in an ensemble.</p>\n<ul>\n<li>more increase FEW and VERY FEW. Increased FEW by 4x and VERY FEW by 7x.  </li>\n<li>Increased train sound source from 20 sec to 40 sec.<ul>\n<li>I thought the longer section would contain more call, but it didn't WORK.</li></ul></li>\n<li>Change backbone to tf_efficientnet_b3_ns.</li>\n</ul>\n<p>What worked best.  </p>\n<ul>\n<li>\"nocall\" was added to label and increased from 152 to 153 classes.<ul>\n<li>nocall data found from here <a href=\"https://www.kaggle.com/datasets/kami634/ff1010bird-duration10\" target=\"_blank\">https://www.kaggle.com/datasets/kami634/ff1010bird-duration10</a></li>\n<li>This brings the score up to private 0.7473 public 0.7572</li></ul></li>\n</ul>\n<h2>Choosing subs</h2>\n<p>All scores above are obtained from fold0 only. The combination of fold and weight, I reached public 0.7730 private 0.7600, but I could not choose this sub because I considered that this was overfitting to public, using only fold0 and fold2 out of 5 unfolds.  <br>\nI chose a sub that mixed all folds equally to try to reduce overfitting to the public.  </p>",
      "rawMarkdown": "Thank you for organizing such an interesting competition. It was a great learning experience for me.\n\n## Short Summary\nMy solution is SED, based on the one published by @hidehisaarai1213 in BirdCLEF2021.\nThis ranking was achieved without modifying the model, but by devising a new data set and changing the backborn.\nA big thank you to hidehisaarai1213 for publishing this wonderful notebook!\n\ntrain : https://www.kaggle.com/code/hidehisaarai1213/pytorch-training-birdclef2021-starter\nsimple inference : https://www.kaggle.com/code/hidehisaarai1213/pytorch-inference-birdclef2021-starter\ninfer between chunk : https://www.kaggle.com/code/hidehisaarai1213/birdclef2021-infer-between-chunk\n\n## Preprocessing\nAs you all know, the lable this time was Imbalanced, so I adopted the approach of acquiring the sound source multiple times if the length of the sound source was short.  \nI have expressed the number of labels on three levels as follows.  \nFew tripled and VERY FEW quadrupled their data and studied.  \nAs I will explain later, the data used for training is a section of the sound source(20sec), so I hypothesized that it could be used multiple times by changing the starting point.  \n\n![](https://user-images.githubusercontent.com/55369709/170328390-b056fc7e-a580-4497-9b77-71b6334f90b3.png)\n\nMANY : 'skylar', 'houfin', 'jabwar', 'warwhe1', 'yefcan', 'apapan','iiwi', 'hawcre'  \nFEW : 'hawama', 'omao', 'barpet', 'akiapo', 'elepai', 'aniani'  \nVERY FEW : 'hawgoo', 'ercfra', 'hawhaw', 'hawpet1', 'puaioh', 'crehon','maupar'  \n\n## Augmentation\nWeak augmentation for MANY and strong augmetation for FEW. Strong means that it transforms with a high probability.  \n\n```\naugment_strong = Compose([\n    AddGaussianNoise(min_amplitude=0.001, max_amplitude=0.015, p=0.5),\n    AddGaussianSNR(p=0.5),\n    Gain(min_gain_in_db=-12, max_gain_in_db=12, p=0.5),\n    TimeStretch(min_rate=0.8, max_rate=1.25, p=0.5),\n    AddBackgroundNoise(sounds_path=noise_dir, min_snr_in_db=3, max_snr_in_db=30, p=0.5),\n    AddShortNoises(noise_dir, p=0.5),\n    PitchShift(min_semitones=-4, max_semitones=4, p=0.5),\n    Shift(min_fraction=-0.5, max_fraction=0.5, p=0.5),\n])\n\naugment_weak = Compose([\n    AddGaussianNoise(min_amplitude=0.001, max_amplitude=0.015, p=0.2),\n    AddGaussianSNR(p=0.2),\n    Gain(min_gain_in_db=-12, max_gain_in_db=12, p=0.2),\n    TimeStretch(min_rate=0.8, max_rate=1.25, p=0.2),\n    AddBackgroundNoise(sounds_path=noise_dir, min_snr_in_db=3, max_snr_in_db=30, p=0.2),\n    AddShortNoises(noise_dir, p=0.2),\n    PitchShift(min_semitones=-4, max_semitones=4, p=0.2),\n    Shift(min_fraction=-0.5, max_fraction=0.5, p=0.2),\n])\n```\n\n## Training\nFirst of all, if I run the above hidehisaarai1213's notebook without preprocess with backbone as tf_efficientnet_b0_ns, I get 0.6406 for public and 0.6843 for private.    \n* epoch -> 30\n* loss -> BCEFocal2WayLoss\n* backbone -> tf_efficientnet_b0_ns\n* inf threshold -> 0.02\n\nFrom here, adding the above preprocess and running it, public was 0.7025 and private was 0.6916. Since public went up significantly, I determined that the strategy of increasing FEW and VERY FEW was so effective.  \n\nAdding a second label further increased the score from this point. 0.7053 for public, 0.6871 for private.\n\nThat I tried but it did not lead to an increase in PUBLIC scores. but I did use it in an ensemble.\n* more increase FEW and VERY FEW. Increased FEW by 4x and VERY FEW by 7x.  \n* Increased train sound source from 20 sec to 40 sec.\n    * I thought the longer section would contain more call, but it didn't WORK.\n* Change backbone to tf_efficientnet_b3_ns.\n\nWhat worked best.  \n* \"nocall\" was added to label and increased from 152 to 153 classes.\n    * nocall data found from here https://www.kaggle.com/datasets/kami634/ff1010bird-duration10\n    * This brings the score up to private 0.7473 public 0.7572\n\n\n## Choosing subs\nAll scores above are obtained from fold0 only. The combination of fold and weight, I reached public 0.7730 private 0.7600, but I could not choose this sub because I considered that this was overfitting to public, using only fold0 and fold2 out of 5 unfolds.  \nI chose a sub that mixed all folds equally to try to reduce overfitting to the public.",
      "votes": null
    },
    {
      "id": "1801780",
      "postDate": "05/26/2022 06:31:47",
      "content": "<p>Hello! congratulations on the silver. I have a question: under what grounds did you determine that your initial ensemble(fold0+fold2) was overfitting?</p>",
      "rawMarkdown": "Hello! congratulations on the silver. I have a question: under what grounds did you determine that your initial ensemble(fold0+fold2) was overfitting?",
      "votes": null
    },
    {
      "id": "1802180",
      "postDate": "05/26/2022 14:28:57",
      "content": "<p>Thanks for your question. To be honest I was not able to create a trusted cv this time, so it was just a gut feeling. However, it is generally believed that bagging improves generalization performance, and I trusted it.</p>",
      "rawMarkdown": "Thanks for your question. To be honest I was not able to create a trusted cv this time, so it was just a gut feeling. However, it is generally believed that bagging improves generalization performance, and I trusted it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1801780,
      "author_name": "jayeonyi",
      "author_url": "",
      "post_date": "05/26/2022 06:31:47",
      "content": "<p>Hello! congratulations on the silver. I have a question: under what grounds did you determine that your initial ensemble(fold0+fold2) was overfitting?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1802180,
      "author_name": "chiman3se",
      "author_url": "",
      "post_date": "05/26/2022 14:28:57",
      "content": "<p>Thanks for your question. To be honest I was not able to create a trusted cv this time, so it was just a gut feeling. However, it is generally believed that bagging improves generalization performance, and I trusted it.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1801404": "Thank you for organizing such an interesting competition. It was a great learning experience for me.\n\n## Short Summary\nMy solution is SED, based on the one published by @hidehisaarai1213 in BirdCLEF2021.\nThis ranking was achieved without modifying the model, but by devising a new data set and changing the backborn.\nA big thank you to hidehisaarai1213 for publishing this wonderful notebook!\n\ntrain : https://www.kaggle.com/code/hidehisaarai1213/pytorch-training-birdclef2021-starter\nsimple inference : https://www.kaggle.com/code/hidehisaarai1213/pytorch-inference-birdclef2021-starter\ninfer between chunk : https://www.kaggle.com/code/hidehisaarai1213/birdclef2021-infer-between-chunk\n\n## Preprocessing\nAs you all know, the lable this time was Imbalanced, so I adopted the approach of acquiring the sound source multiple times if the length of the sound source was short.  \nI have expressed the number of labels on three levels as follows.  \nFew tripled and VERY FEW quadrupled their data and studied.  \nAs I will explain later, the data used for training is a section of the sound source(20sec), so I hypothesized that it could be used multiple times by changing the starting point.  \n\n![](https://user-images.githubusercontent.com/55369709/170328390-b056fc7e-a580-4497-9b77-71b6334f90b3.png)\n\nMANY : 'skylar', 'houfin', 'jabwar', 'warwhe1', 'yefcan', 'apapan','iiwi', 'hawcre'  \nFEW : 'hawama', 'omao', 'barpet', 'akiapo', 'elepai', 'aniani'  \nVERY FEW : 'hawgoo', 'ercfra', 'hawhaw', 'hawpet1', 'puaioh', 'crehon','maupar'  \n\n## Augmentation\nWeak augmentation for MANY and strong augmetation for FEW. Strong means that it transforms with a high probability.  \n\n```\naugment_strong = Compose([\n    AddGaussianNoise(min_amplitude=0.001, max_amplitude=0.015, p=0.5),\n    AddGaussianSNR(p=0.5),\n    Gain(min_gain_in_db=-12, max_gain_in_db=12, p=0.5),\n    TimeStretch(min_rate=0.8, max_rate=1.25, p=0.5),\n    AddBackgroundNoise(sounds_path=noise_dir, min_snr_in_db=3, max_snr_in_db=30, p=0.5),\n    AddShortNoises(noise_dir, p=0.5),\n    PitchShift(min_semitones=-4, max_semitones=4, p=0.5),\n    Shift(min_fraction=-0.5, max_fraction=0.5, p=0.5),\n])\n\naugment_weak = Compose([\n    AddGaussianNoise(min_amplitude=0.001, max_amplitude=0.015, p=0.2),\n    AddGaussianSNR(p=0.2),\n    Gain(min_gain_in_db=-12, max_gain_in_db=12, p=0.2),\n    TimeStretch(min_rate=0.8, max_rate=1.25, p=0.2),\n    AddBackgroundNoise(sounds_path=noise_dir, min_snr_in_db=3, max_snr_in_db=30, p=0.2),\n    AddShortNoises(noise_dir, p=0.2),\n    PitchShift(min_semitones=-4, max_semitones=4, p=0.2),\n    Shift(min_fraction=-0.5, max_fraction=0.5, p=0.2),\n])\n```\n\n## Training\nFirst of all, if I run the above hidehisaarai1213's notebook without preprocess with backbone as tf_efficientnet_b0_ns, I get 0.6406 for public and 0.6843 for private.    \n* epoch -> 30\n* loss -> BCEFocal2WayLoss\n* backbone -> tf_efficientnet_b0_ns\n* inf threshold -> 0.02\n\nFrom here, adding the above preprocess and running it, public was 0.7025 and private was 0.6916. Since public went up significantly, I determined that the strategy of increasing FEW and VERY FEW was so effective.  \n\nAdding a second label further increased the score from this point. 0.7053 for public, 0.6871 for private.\n\nThat I tried but it did not lead to an increase in PUBLIC scores. but I did use it in an ensemble.\n* more increase FEW and VERY FEW. Increased FEW by 4x and VERY FEW by 7x.  \n* Increased train sound source from 20 sec to 40 sec.\n    * I thought the longer section would contain more call, but it didn't WORK.\n* Change backbone to tf_efficientnet_b3_ns.\n\nWhat worked best.  \n* \"nocall\" was added to label and increased from 152 to 153 classes.\n    * nocall data found from here https://www.kaggle.com/datasets/kami634/ff1010bird-duration10\n    * This brings the score up to private 0.7473 public 0.7572\n\n\n## Choosing subs\nAll scores above are obtained from fold0 only. The combination of fold and weight, I reached public 0.7730 private 0.7600, but I could not choose this sub because I considered that this was overfitting to public, using only fold0 and fold2 out of 5 unfolds.  \nI chose a sub that mixed all folds equally to try to reduce overfitting to the public.",
    "1801780": "Hello! congratulations on the silver. I have a question: under what grounds did you determine that your initial ensemble(fold0+fold2) was overfitting?",
    "1802180": "Thanks for your question. To be honest I was not able to create a trusted cv this time, so it was just a gut feeling. However, it is generally believed that bagging improves generalization performance, and I trusted it."
  },
  "source": "meta"
}