{
  "id": 511596,
  "title": "10th Solution",
  "url": "/competitions/birdclef-2024/writeups/tamo-10th-solution",
  "author_name": "",
  "post_date": "2024-06-11T11:31:52.305489400Z",
  "votes": 26,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Congrats to all the winners, and thank you to Kaggle and Cornell Lab of Ornithology for organizing such an interesting competition. I will mainly explain the method from my solution that I believe helped improve my LB position.</p>\n<h3>Summary</h3>\n<p>I trained the model in three stages. First, I used all the data from 2021 to 2024. Next, I trained on either just the 2024 data or on the data from 2021 to 2024, focusing on the target of this competition. Finally, I trained on the second set of data, adding the unlabeled data from 2024. The final stage of training improved the model's performance on both the public and private LB.</p>\n<h3>Data</h3>\n<p>2021,2022,2023,2024</p>\n<h3>Model</h3>\n<p>tf_efficientnetv2_b0, b3 (timm)<br>\ninput: single-channel spectrogram</p>\n<h3>Use of Unlabeled Data</h3>\n<p>I created pseudo-labels for the unlabeled data and used them by adding them to the labeled data with a 50% probability. This method was inspired by the <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412922\" target=\"_blank\">7th solution in 2023</a> and implemented with reference to <a href=\"https://github.com/Sato-Kunihiko/audio-SNR\" target=\"_blank\">GitHub</a>. tgt_wav was the unlabeled data, add_wav was the labeled data, and db was a random positive value(~50). The labels were simply combined.</p>\n<pre><code>def __add_noise(self, tgt_wav, add_wav, db):\n    tgt_rms = .(.(.square(tgt_wav), axis=-))\n    add_rms = .(.(.square(add_wav), axis=-))\n\n    noise_rms = tgt_rms / (**((db) / )) \n    new_wav = tgt_wav + add_wav * (noise_rms / (add_rms + ))\n    new_wav = .clip(new_wav, .(new_wav)*, .(new_wav)*)\n     new_wav\n</code></pre>\n<p>I divided each 4-minute unlabeled data file into 5-second segments, creating 48 files from one unlabeled file.　After generating pseudo-labels for each segmented file, I used the top 10% of those with the lowest class label entropy for training. Labels below the top 90% for each class's pseudo-label were set to 0. The labels were used as soft labels. This method improved the LB by ~0.04.</p>",
  "messages": [
    {
      "id": "2866527",
      "postDate": "06/11/2024 11:31:52",
      "content": "<p>Congrats to all the winners, and thank you to Kaggle and Cornell Lab of Ornithology for organizing such an interesting competition. I will mainly explain the method from my solution that I believe helped improve my LB position.</p>\n<h3>Summary</h3>\n<p>I trained the model in three stages. First, I used all the data from 2021 to 2024. Next, I trained on either just the 2024 data or on the data from 2021 to 2024, focusing on the target of this competition. Finally, I trained on the second set of data, adding the unlabeled data from 2024. The final stage of training improved the model's performance on both the public and private LB.</p>\n<h3>Data</h3>\n<p>2021,2022,2023,2024</p>\n<h3>Model</h3>\n<p>tf_efficientnetv2_b0, b3 (timm)<br>\ninput: single-channel spectrogram</p>\n<h3>Use of Unlabeled Data</h3>\n<p>I created pseudo-labels for the unlabeled data and used them by adding them to the labeled data with a 50% probability. This method was inspired by the <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412922\" target=\"_blank\">7th solution in 2023</a> and implemented with reference to <a href=\"https://github.com/Sato-Kunihiko/audio-SNR\" target=\"_blank\">GitHub</a>. tgt_wav was the unlabeled data, add_wav was the labeled data, and db was a random positive value(~50). The labels were simply combined.</p>\n<pre><code>def __add_noise(self, tgt_wav, add_wav, db):\n    tgt_rms = .(.(.square(tgt_wav), axis=-))\n    add_rms = .(.(.square(add_wav), axis=-))\n\n    noise_rms = tgt_rms / (**((db) / )) \n    new_wav = tgt_wav + add_wav * (noise_rms / (add_rms + ))\n    new_wav = .clip(new_wav, .(new_wav)*, .(new_wav)*)\n     new_wav\n</code></pre>\n<p>I divided each 4-minute unlabeled data file into 5-second segments, creating 48 files from one unlabeled file.　After generating pseudo-labels for each segmented file, I used the top 10% of those with the lowest class label entropy for training. Labels below the top 90% for each class's pseudo-label were set to 0. The labels were used as soft labels. This method improved the LB by ~0.04.</p>",
      "rawMarkdown": "Congrats to all the winners, and thank you to Kaggle and Cornell Lab of Ornithology for organizing such an interesting competition. I will mainly explain the method from my solution that I believe helped improve my LB position.\n\n### Summary\nI trained the model in three stages. First, I used all the data from 2021 to 2024. Next, I trained on either just the 2024 data or on the data from 2021 to 2024, focusing on the target of this competition. Finally, I trained on the second set of data, adding the unlabeled data from 2024. The final stage of training improved the model's performance on both the public and private LB.\n\n### Data\n2021,2022,2023,2024\n\n### Model\ntf_efficientnetv2_b0, b3 (timm)\ninput: single-channel spectrogram\n\n### Use of Unlabeled Data\nI created pseudo-labels for the unlabeled data and used them by adding them to the labeled data with a 50% probability. This method was inspired by the [7th solution in 2023](https://www.kaggle.com/competitions/birdclef-2023/discussion/412922) and implemented with reference to [GitHub](https://github.com/Sato-Kunihiko/audio-SNR). tgt_wav was the unlabeled data, add_wav was the labeled data, and db was a random positive value(~50). The labels were simply combined.\n\n```\ndef __add_noise(self, tgt_wav, add_wav, db):\n    tgt_rms = np.sqrt(np.mean(np.square(tgt_wav), axis=-1))\n    add_rms = np.sqrt(np.mean(np.square(add_wav), axis=-1))\n    \n    noise_rms = tgt_rms / (10**(float(db) / 20)) \n    new_wav = tgt_wav + add_wav * (noise_rms / (add_rms + 1e-6))\n    new_wav = np.clip(new_wav, np.min(new_wav)*2, np.max(new_wav)*2)\n    return new_wav\n```\n\nI divided each 4-minute unlabeled data file into 5-second segments, creating 48 files from one unlabeled file.　After generating pseudo-labels for each segmented file, I used the top 10% of those with the lowest class label entropy for training. Labels below the top 90% for each class's pseudo-label were set to 0. The labels were used as soft labels. This method improved the LB by ~0.04.",
      "votes": null
    },
    {
      "id": "2866614",
      "postDate": "06/11/2024 12:04:37",
      "content": "<p>Congratulations solo gold medal.</p>",
      "rawMarkdown": "Congratulations solo gold medal.",
      "votes": null
    },
    {
      "id": "2867834",
      "postDate": "06/12/2024 05:24:52",
      "content": "<p>Congratulations on  securing 10th Place in this competition. Thanks for sharing details.</p>",
      "rawMarkdown": "Congratulations on  securing 10th Place in this competition. Thanks for sharing details.",
      "votes": null
    },
    {
      "id": "2868746",
      "postDate": "06/12/2024 15:40:04",
      "content": "<p>Congratulations on securing 10th place in this competition, <a href=\"https://www.kaggle.com/yuyuki11235\" target=\"_blank\">@yuyuki11235</a> </p>",
      "rawMarkdown": "Congratulations on securing 10th place in this competition, @yuyuki11235",
      "votes": null
    },
    {
      "id": "2870366",
      "postDate": "06/13/2024 14:45:24",
      "content": "<p>Congratulations on securing 10th Place in this competition. Thanks for sharing details.</p>",
      "rawMarkdown": "Congratulations on securing 10th Place in this competition. Thanks for sharing details.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2866614,
      "author_name": "clora16",
      "author_url": "",
      "post_date": "06/11/2024 12:04:37",
      "content": "<p>Congratulations solo gold medal.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2867834,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "06/12/2024 05:24:52",
      "content": "<p>Congratulations on  securing 10th Place in this competition. Thanks for sharing details.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2868746,
      "author_name": "mubashirsidiki",
      "author_url": "",
      "post_date": "06/12/2024 15:40:04",
      "content": "<p>Congratulations on securing 10th place in this competition, <a href=\"https://www.kaggle.com/yuyuki11235\" target=\"_blank\">@yuyuki11235</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2870366,
      "author_name": "germanarley",
      "author_url": "",
      "post_date": "06/13/2024 14:45:24",
      "content": "<p>Congratulations on securing 10th Place in this competition. Thanks for sharing details.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2866527": "Congrats to all the winners, and thank you to Kaggle and Cornell Lab of Ornithology for organizing such an interesting competition. I will mainly explain the method from my solution that I believe helped improve my LB position.\n\n### Summary\nI trained the model in three stages. First, I used all the data from 2021 to 2024. Next, I trained on either just the 2024 data or on the data from 2021 to 2024, focusing on the target of this competition. Finally, I trained on the second set of data, adding the unlabeled data from 2024. The final stage of training improved the model's performance on both the public and private LB.\n\n### Data\n2021,2022,2023,2024\n\n### Model\ntf_efficientnetv2_b0, b3 (timm)\ninput: single-channel spectrogram\n\n### Use of Unlabeled Data\nI created pseudo-labels for the unlabeled data and used them by adding them to the labeled data with a 50% probability. This method was inspired by the [7th solution in 2023](https://www.kaggle.com/competitions/birdclef-2023/discussion/412922) and implemented with reference to [GitHub](https://github.com/Sato-Kunihiko/audio-SNR). tgt_wav was the unlabeled data, add_wav was the labeled data, and db was a random positive value(~50). The labels were simply combined.\n\n```\ndef __add_noise(self, tgt_wav, add_wav, db):\n    tgt_rms = np.sqrt(np.mean(np.square(tgt_wav), axis=-1))\n    add_rms = np.sqrt(np.mean(np.square(add_wav), axis=-1))\n    \n    noise_rms = tgt_rms / (10**(float(db) / 20)) \n    new_wav = tgt_wav + add_wav * (noise_rms / (add_rms + 1e-6))\n    new_wav = np.clip(new_wav, np.min(new_wav)*2, np.max(new_wav)*2)\n    return new_wav\n```\n\nI divided each 4-minute unlabeled data file into 5-second segments, creating 48 files from one unlabeled file.　After generating pseudo-labels for each segmented file, I used the top 10% of those with the lowest class label entropy for training. Labels below the top 90% for each class's pseudo-label were set to 0. The labels were used as soft labels. This method improved the LB by ~0.04.",
    "2866614": "Congratulations solo gold medal.",
    "2867834": "Congratulations on  securing 10th Place in this competition. Thanks for sharing details.",
    "2868746": "Congratulations on securing 10th place in this competition, @yuyuki11235",
    "2870366": "Congratulations on securing 10th Place in this competition. Thanks for sharing details."
  },
  "source": "meta"
}