{
  "id": 511540,
  "title": "7th Place Solution for the BirdCLEF 2024 Competition",
  "url": "/competitions/birdclef-2024/discussion/511540",
  "author_name": "RihanPiggy",
  "post_date": "2024-06-11T06:14:19.432000",
  "votes": 38,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Congratulations to all the winners! Thanks to Kaggle and Cornell Lab of Ornithology for hosting this interesting competition.</p>\n<p>This year's birdclef is really a hard one. Unlike padded CMAP which suppress the impact of species with very few positive labels, AUC treat every species equally and is very sensitive.</p>\n<p>I am happy to survive the final shake. Let me introduce my solution.</p>\n<p>Thanks to every competitor who gave me inspiration. Special thanks to atsunorifujita, martynoveduard, loonypenguin, vladimirsydor, anonamename</p>\n<h1>Context</h1>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/birdclef-2024/overview\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/birdclef-2024/data\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/data</a></li>\n</ul>\n<h1>Overview of the approach</h1>\n<p>My final submission is a combination of 2 SED model and 1 CNN model. The training scheme is almost the same as my <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412707\" target=\"_blank\">last year's 2nd place solution</a>.<br>\nThe new stuffs I added include</p>\n<ul>\n<li>Extract training sample from soundscape using Birdnet and Bird-vocalization-classifier</li>\n<li>Knowledge distillation proposed by <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412753\" target=\"_blank\">last year's 4th place</a></li>\n<li>Sumixup proposed by <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412922\" target=\"_blank\">last year's 7th place</a></li>\n<li>Int8 quantize</li>\n<li>Split species to serveral subset (most important for me)</li>\n</ul>\n<p>The key to win this competition is to split species to several subset and train the subset separately, which enables model to focus more on each species compared to train them all.</p>\n<p>My final submission is an ensemble of 2 models with all species and 1 model with 66 species which have relatively small sample size.<br>\nActually I had an experiment which split species to 3 subset, which gave me 0.690440(1st place) in private LB, but it was too risky to choose because I was not sure whether I was just overfitting the public LB.</p>\n<h1>Details of the submission</h1>\n<h2>Models</h2>\n<ul>\n<li>SED with tf_efficientnetv2_s_in21k (all species)</li>\n<li>SED with seresnext26t_32x4d (66 rare species)</li>\n<li>CNN with resnet34d (all species)</li>\n</ul>\n<h2>Extra data on xeno-canto</h2>\n<p>Like last year's situation, I collected extra audio from xeno-canto and trained a baseline SED v2s model. The baseline score was 0.65.</p>\n<h2>Extracting train sample from soundscape using Birdnet and Bird-vocalization-classifier</h2>\n<p>Birdnet covers 181 of 182 species and Bird-vocalization-classifier covers 180 of 182 species. I mainly used birdnet and extracted niwpig1 with bird-vocalization-classifier. I extracted the 15 second audio clip with threshold 0.3.</p>\n<p>This boosted my baseline score to 0.68.</p>\n<h2>Knowledge Distillation</h2>\n<p>I implemented the same knowledge distillation scheme as last year's 4th place. This further boosted my baseline score to 0.70.</p>\n<h2>Further Extracting train sample from soundscape using trained models</h2>\n<p>I used my own trained models to further extract train samples from soundscape, this gave me a little boost in LB.</p>\n<h2>Species subset</h2>\n<p>Inspired by <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/327193\" target=\"_blank\">3rd place in birdclef 2022</a>, I tried to train a model with 66 species which have relatively small sample size, and I found that this model outperformed model trained with all species on 43 of 66 species.</p>\n<p>I was not sure whether this will also happen in private LB, but I decided to ensemble 1 model trained on 66 species into my final submission.</p>\n<p>I also tried to split species into 3 subsets with some overlap. I trained them separately and then created a submission with 3 subset model ensemble. This gave me 0.690440(1st place) in private LB, thus maybe splitting birds to subset is the key to win. But it was too risky to choose because I was not sure whether I was just overfitting the public LB.</p>\n<h2>Int8 quantize</h2>\n<p>Thanks to the unlabeled soundscape, we can perform quantization this year.</p>\n<p>I used nncf to quantize encoder part of sed model. Quantizing CNN model led to 0.01 decrease in LB score, so I didn't perform quantization for CNN.</p>\n<p>Quantization reduced inference time by 20 minutes for sed v2s model. (100 minutes to 80 minutes)</p>\n<pre><code> nncf\n openvino  ov\n\nnncf_dataset = nncf.Dataset(pytorch_dataset, transform_fn)\nmodel = ov.Core().read_model(path_to_model)\nquantized_model = nncf.quantize(\n    model, nncf_dataset,\n    target_device=nncf.TargetDevice.CPU,\n    subset_size=,\n    fast_bias_correction=,\n    preset=nncf.QuantizationPreset.MIXED,\n)\n\nov.save_model(quantized_model, quantized_model_path)\n</code></pre>\n<h2>ensemble strategy</h2>\n<ul>\n<li>logit average of 3 models</li>\n<li>rank average of 3 models</li>\n</ul>\n<h2>Things didn't work for me.</h2>\n<p>Too many…</p>\n<ul>\n<li>Adding validation data to training led to 0.01 LB decrease. It is very weird, but I have to accept it and be faithful…</li>\n<li>Calculating species weight to deal with label shift and covariance shift</li>\n<li>Giving more weight to audios recorded in india according to geometric information (latitude 8~21, longitude 72.5 ~ 79)</li>\n<li>changing weight of the sampler significantly decrease the LB. Sampling every species equally performed best</li>\n<li>Filtering out 2024 train data duplicates not only by id but also by 'author + primary_label'(proposed <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412808\" target=\"_blank\">here</a>)</li>\n<li>using embedding extracted from birdnet and bird-vocalization-classifier proposed by <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412708\" target=\"_blank\">last year's 6th place</a></li>\n</ul>\n<h1>Sources</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412707\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/discussion/412707</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412753\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/discussion/412753</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412922\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/discussion/412922</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412808\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/discussion/412808</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412708\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/discussion/412708</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/327193\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/327193</a></li>\n<li><a href=\"https://docs.openvino.ai/2023.3/basic_quantization_flow.html\" target=\"_blank\">https://docs.openvino.ai/2023.3/basic_quantization_flow.html</a></li>\n</ul>",
  "messages": [
    {
      "id": 2866078,
      "postDate": "2024-06-11T06:14:19.433Z",
      "content": "<p>Congratulations to all the winners! Thanks to Kaggle and Cornell Lab of Ornithology for hosting this interesting competition.</p>\n<p>This year's birdclef is really a hard one. Unlike padded CMAP which suppress the impact of species with very few positive labels, AUC treat every species equally and is very sensitive.</p>\n<p>I am happy to survive the final shake. Let me introduce my solution.</p>\n<p>Thanks to every competitor who gave me inspiration. Special thanks to atsunorifujita, martynoveduard, loonypenguin, vladimirsydor, anonamename</p>\n<h1>Context</h1>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/birdclef-2024/overview\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/birdclef-2024/data\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/data</a></li>\n</ul>\n<h1>Overview of the approach</h1>\n<p>My final submission is a combination of 2 SED model and 1 CNN model. The training scheme is almost the same as my <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412707\" target=\"_blank\">last year's 2nd place solution</a>.<br>\nThe new stuffs I added include</p>\n<ul>\n<li>Extract training sample from soundscape using Birdnet and Bird-vocalization-classifier</li>\n<li>Knowledge distillation proposed by <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412753\" target=\"_blank\">last year's 4th place</a></li>\n<li>Sumixup proposed by <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412922\" target=\"_blank\">last year's 7th place</a></li>\n<li>Int8 quantize</li>\n<li>Split species to serveral subset (most important for me)</li>\n</ul>\n<p>The key to win this competition is to split species to several subset and train the subset separately, which enables model to focus more on each species compared to train them all.</p>\n<p>My final submission is an ensemble of 2 models with all species and 1 model with 66 species which have relatively small sample size.<br>\nActually I had an experiment which split species to 3 subset, which gave me 0.690440(1st place) in private LB, but it was too risky to choose because I was not sure whether I was just overfitting the public LB.</p>\n<h1>Details of the submission</h1>\n<h2>Models</h2>\n<ul>\n<li>SED with tf_efficientnetv2_s_in21k (all species)</li>\n<li>SED with seresnext26t_32x4d (66 rare species)</li>\n<li>CNN with resnet34d (all species)</li>\n</ul>\n<h2>Extra data on xeno-canto</h2>\n<p>Like last year's situation, I collected extra audio from xeno-canto and trained a baseline SED v2s model. The baseline score was 0.65.</p>\n<h2>Extracting train sample from soundscape using Birdnet and Bird-vocalization-classifier</h2>\n<p>Birdnet covers 181 of 182 species and Bird-vocalization-classifier covers 180 of 182 species. I mainly used birdnet and extracted niwpig1 with bird-vocalization-classifier. I extracted the 15 second audio clip with threshold 0.3.</p>\n<p>This boosted my baseline score to 0.68.</p>\n<h2>Knowledge Distillation</h2>\n<p>I implemented the same knowledge distillation scheme as last year's 4th place. This further boosted my baseline score to 0.70.</p>\n<h2>Further Extracting train sample from soundscape using trained models</h2>\n<p>I used my own trained models to further extract train samples from soundscape, this gave me a little boost in LB.</p>\n<h2>Species subset</h2>\n<p>Inspired by <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/327193\" target=\"_blank\">3rd place in birdclef 2022</a>, I tried to train a model with 66 species which have relatively small sample size, and I found that this model outperformed model trained with all species on 43 of 66 species.</p>\n<p>I was not sure whether this will also happen in private LB, but I decided to ensemble 1 model trained on 66 species into my final submission.</p>\n<p>I also tried to split species into 3 subsets with some overlap. I trained them separately and then created a submission with 3 subset model ensemble. This gave me 0.690440(1st place) in private LB, thus maybe splitting birds to subset is the key to win. But it was too risky to choose because I was not sure whether I was just overfitting the public LB.</p>\n<h2>Int8 quantize</h2>\n<p>Thanks to the unlabeled soundscape, we can perform quantization this year.</p>\n<p>I used nncf to quantize encoder part of sed model. Quantizing CNN model led to 0.01 decrease in LB score, so I didn't perform quantization for CNN.</p>\n<p>Quantization reduced inference time by 20 minutes for sed v2s model. (100 minutes to 80 minutes)</p>\n<pre><code> nncf\n openvino  ov\n\nnncf_dataset = nncf.Dataset(pytorch_dataset, transform_fn)\nmodel = ov.Core().read_model(path_to_model)\nquantized_model = nncf.quantize(\n    model, nncf_dataset,\n    target_device=nncf.TargetDevice.CPU,\n    subset_size=,\n    fast_bias_correction=,\n    preset=nncf.QuantizationPreset.MIXED,\n)\n\nov.save_model(quantized_model, quantized_model_path)\n</code></pre>\n<h2>ensemble strategy</h2>\n<ul>\n<li>logit average of 3 models</li>\n<li>rank average of 3 models</li>\n</ul>\n<h2>Things didn't work for me.</h2>\n<p>Too many…</p>\n<ul>\n<li>Adding validation data to training led to 0.01 LB decrease. It is very weird, but I have to accept it and be faithful…</li>\n<li>Calculating species weight to deal with label shift and covariance shift</li>\n<li>Giving more weight to audios recorded in india according to geometric information (latitude 8~21, longitude 72.5 ~ 79)</li>\n<li>changing weight of the sampler significantly decrease the LB. Sampling every species equally performed best</li>\n<li>Filtering out 2024 train data duplicates not only by id but also by 'author + primary_label'(proposed <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412808\" target=\"_blank\">here</a>)</li>\n<li>using embedding extracted from birdnet and bird-vocalization-classifier proposed by <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412708\" target=\"_blank\">last year's 6th place</a></li>\n</ul>\n<h1>Sources</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412707\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/discussion/412707</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412753\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/discussion/412753</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412922\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/discussion/412922</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412808\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/discussion/412808</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412708\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/discussion/412708</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/327193\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/327193</a></li>\n<li><a href=\"https://docs.openvino.ai/2023.3/basic_quantization_flow.html\" target=\"_blank\">https://docs.openvino.ai/2023.3/basic_quantization_flow.html</a></li>\n</ul>",
      "rawMarkdown": "Congratulations to all the winners! Thanks to Kaggle and Cornell Lab of Ornithology for hosting this interesting competition.\n\nThis year's birdclef is really a hard one. Unlike padded CMAP which suppress the impact of species with very few positive labels, AUC treat every species equally and is very sensitive.\n\nI am happy to survive the final shake. Let me introduce my solution.\n\nThanks to every competitor who gave me inspiration. Special thanks to atsunorifujita, martynoveduard, loonypenguin, vladimirsydor, anonamename\n\n# Context\n\n- Business context: https://www.kaggle.com/competitions/birdclef-2024/overview\n- Data context: https://www.kaggle.com/competitions/birdclef-2024/data\n\n# Overview of the approach\n\nMy final submission is a combination of 2 SED model and 1 CNN model. The training scheme is almost the same as my [last year's 2nd place solution](https://www.kaggle.com/competitions/birdclef-2023/discussion/412707).\nThe new stuffs I added include\n\n- Extract training sample from soundscape using Birdnet and Bird-vocalization-classifier\n- Knowledge distillation proposed by [last year's 4th place](https://www.kaggle.com/competitions/birdclef-2023/discussion/412753)\n- Sumixup proposed by [last year's 7th place](https://www.kaggle.com/competitions/birdclef-2023/discussion/412922)\n- Int8 quantize\n- Split species to serveral subset (most important for me)\n\nThe key to win this competition is to split species to several subset and train the subset separately, which enables model to focus more on each species compared to train them all.\n\nMy final submission is an ensemble of 2 models with all species and 1 model with 66 species which have relatively small sample size.\nActually I had an experiment which split species to 3 subset, which gave me 0.690440(1st place) in private LB, but it was too risky to choose because I was not sure whether I was just overfitting the public LB.\n\n# Details of the submission\n\n## Models\n\n- SED with tf_efficientnetv2_s_in21k (all species)\n- SED with seresnext26t_32x4d (66 rare species)\n- CNN with resnet34d (all species)\n\n## Extra data on xeno-canto\n\nLike last year's situation, I collected extra audio from xeno-canto and trained a baseline SED v2s model. The baseline score was 0.65.\n\n## Extracting train sample from soundscape using Birdnet and Bird-vocalization-classifier\n\nBirdnet covers 181 of 182 species and Bird-vocalization-classifier covers 180 of 182 species. I mainly used birdnet and extracted niwpig1 with bird-vocalization-classifier. I extracted the 15 second audio clip with threshold 0.3.\n\nThis boosted my baseline score to 0.68.\n\n## Knowledge Distillation\n\nI implemented the same knowledge distillation scheme as last year's 4th place. This further boosted my baseline score to 0.70.\n\n## Further Extracting train sample from soundscape using trained models\n\nI used my own trained models to further extract train samples from soundscape, this gave me a little boost in LB.\n\n## Species subset\n\nInspired by [3rd place in birdclef 2022](https://www.kaggle.com/competitions/birdclef-2022/discussion/327193), I tried to train a model with 66 species which have relatively small sample size, and I found that this model outperformed model trained with all species on 43 of 66 species.\n\nI was not sure whether this will also happen in private LB, but I decided to ensemble 1 model trained on 66 species into my final submission.\n\nI also tried to split species into 3 subsets with some overlap. I trained them separately and then created a submission with 3 subset model ensemble. This gave me 0.690440(1st place) in private LB, thus maybe splitting birds to subset is the key to win. But it was too risky to choose because I was not sure whether I was just overfitting the public LB.\n\n## Int8 quantize\n\nThanks to the unlabeled soundscape, we can perform quantization this year.\n\nI used nncf to quantize encoder part of sed model. Quantizing CNN model led to 0.01 decrease in LB score, so I didn't perform quantization for CNN.\n\nQuantization reduced inference time by 20 minutes for sed v2s model. (100 minutes to 80 minutes)\n\n```python\nimport nncf\nimport openvino as ov\n\nnncf_dataset = nncf.Dataset(pytorch_dataset, transform_fn)\nmodel = ov.Core().read_model(path_to_model)\nquantized_model = nncf.quantize(\n    model, nncf_dataset,\n    target_device=nncf.TargetDevice.CPU,\n    subset_size=300,\n    fast_bias_correction=True,\n    preset=nncf.QuantizationPreset.MIXED,\n)\n\nov.save_model(quantized_model, quantized_model_path)\n```\n\n\n\n## ensemble strategy\n\n- logit average of 3 models\n- rank average of 3 models\n\n## Things didn't work for me.\n\nToo many...\n\n- Adding validation data to training led to 0.01 LB decrease. It is very weird, but I have to accept it and be faithful...\n- Calculating species weight to deal with label shift and covariance shift\n- Giving more weight to audios recorded in india according to geometric information (latitude 8~21, longitude 72.5 ~ 79)\n- changing weight of the sampler significantly decrease the LB. Sampling every species equally performed best\n- Filtering out 2024 train data duplicates not only by id but also by 'author + primary_label'(proposed [here](https://www.kaggle.com/competitions/birdclef-2023/discussion/412808))\n- using embedding extracted from birdnet and bird-vocalization-classifier proposed by [last year's 6th place](https://www.kaggle.com/competitions/birdclef-2023/discussion/412708)\n\n# Sources\n\n- https://www.kaggle.com/competitions/birdclef-2023/discussion/412707\n- https://www.kaggle.com/competitions/birdclef-2023/discussion/412753\n- https://www.kaggle.com/competitions/birdclef-2023/discussion/412922\n- https://www.kaggle.com/competitions/birdclef-2023/discussion/412808\n- https://www.kaggle.com/competitions/birdclef-2023/discussion/412708\n- https://www.kaggle.com/competitions/birdclef-2022/discussion/327193\n- https://docs.openvino.ai/2023.3/basic_quantization_flow.html\n",
      "votes": 38
    },
    {
      "id": 2866276,
      "postDate": "2024-06-11T08:28:28.797Z",
      "content": "<p>Congratulations for gold medal !!!!!!!!</p>",
      "rawMarkdown": "Congratulations for gold medal !!!!!!!!",
      "votes": 1,
      "replies": [
        {
          "id": 2866290,
          "postDate": "2024-06-11T08:31:55.283Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2867416,
      "postDate": "2024-06-11T19:49:56.873Z",
      "content": "<p>Congratulations for gold!! I Had a beginner level doubt like are you feeding the model single channel spectrogram like (1,256,625) example shape or  3 channel by stacking the single channel?</p>",
      "rawMarkdown": "Congratulations for gold!! I Had a beginner level doubt like are you feeding the model single channel spectrogram like (1,256,625) example shape or  3 channel by stacking the single channel?",
      "replies": [
        {
          "id": 2867620,
          "postDate": "2024-06-12T01:28:24.803Z",
          "content": "<p>Thank you!<br>\nMel spectrogram is single channel. And I used <a href=\"https://pytorch.org/audio/main/generated/torchaudio.functional.compute_deltas.html\" target=\"_blank\">compute delta</a> and stack to 3 channel.</p>\n<pre><code> ():\n    input_tensor = input_tensor.transpose(,)\n    input_tensor = compute_deltas(input_tensor)\n    input_tensor = input_tensor.transpose(,)\n     input_tensor\n\n\n ():\n    delta_1 = make_delta(x)\n    delta_2 = make_delta(delta_1)\n    x = torch.cat([x,delta_1,delta_2], dim=)\n     x\n</code></pre>",
          "rawMarkdown": "Thank you!\nMel spectrogram is single channel. And I used [compute delta](https://pytorch.org/audio/main/generated/torchaudio.functional.compute_deltas.html) and stack to 3 channel.\n\n```python\ndef make_delta(\n    input_tensor: torch.Tensor\n):\n    input_tensor = input_tensor.transpose(3,2)\n    input_tensor = compute_deltas(input_tensor)\n    input_tensor = input_tensor.transpose(3,2)\n    return input_tensor\n\n\ndef image_delta(x):\n    delta_1 = make_delta(x)\n    delta_2 = make_delta(delta_1)\n    x = torch.cat([x,delta_1,delta_2], dim=1)\n    return x\n```",
          "votes": 2,
          "replies": [
            {
              "id": 2867848,
              "postDate": "2024-06-12T05:46:38.927Z",
              "content": "<p>Thank you for response was searching for something like this from long time</p>",
              "rawMarkdown": "Thank you for response was searching for something like this from long time"
            }
          ]
        }
      ]
    },
    {
      "id": 2866865,
      "postDate": "2024-06-11T14:44:03.743Z",
      "content": "<p>Thanks for sharing your solution again, I learnt a lot from what you shared last year!</p>\n<ol>\n<li><p>You mention “Sampling every species equally”. Does this mean that you randomly sampled a species with weight <code>1/182</code>, then sampled a datapoint within that species, or does it mean that each data point was seen once each epoch?</p></li>\n<li><p>Also, could you clarify how you create the 10-sec segment during training? Is the same as 2023 (eg. 60sec waveform -&gt; split to 6 * 10sec -&gt; np.sum(audios,axis=0))?</p></li>\n</ol>",
      "rawMarkdown": "Thanks for sharing your solution again, I learnt a lot from what you shared last year!\n\n1. You mention “Sampling every species equally”. Does this mean that you randomly sampled a species with weight `1/182`, then sampled a datapoint within that species, or does it mean that each data point was seen once each epoch?\n\n2. Also, could you clarify how you create the 10-sec segment during training? Is the same as 2023 (eg. 60sec waveform -> split to 6 * 10sec -> np.sum(audios,axis=0))?",
      "replies": [
        {
          "id": 2867625,
          "postDate": "2024-06-12T01:38:11.410Z",
          "content": "<p>Thank you!.</p>\n<ol>\n<li><p>Yes, the weight for each species is 1/182, but it is not a two step sampling. I used weighted sampler provided by pytorch. Let's assume there are 3 species (a,b,c). a has 3 samples, b has 2 samples, c has 1 sample. Then the weight for sample of a, b, c is 1/3, 1/3, 1/3, 1/2, 1/2, 1/1. You can find the weight for each species is the same here.</p></li>\n<li><p>For SED model, I simply randomly selected 10 sec from the audio clip. For audio extracted from unlabeled soundscape( I clipped 15sec),  it is 15sec waveform -&gt; 10sec, 5sec -&gt; 10sec, 10sec(pad 0) -&gt;  np.sum(audios,axis=0))</p></li>\n</ol>",
          "rawMarkdown": "Thank you!.\n\n1. Yes, the weight for each species is 1/182, but it is not a two step sampling. I used weighted sampler provided by pytorch. Let's assume there are 3 species (a,b,c). a has 3 samples, b has 2 samples, c has 1 sample. Then the weight for sample of a, b, c is 1/3, 1/3, 1/3, 1/2, 1/2, 1/1. You can find the weight for each species is the same here.\n\n2. For SED model, I simply randomly selected 10 sec from the audio clip. For audio extracted from unlabeled soundscape( I clipped 15sec),  it is 15sec waveform -> 10sec, 5sec -> 10sec, 10sec(pad 0) ->  np.sum(audios,axis=0))",
          "votes": 1,
          "replies": [
            {
              "id": 2868595,
              "postDate": "2024-06-12T14:00:27.670Z",
              "content": "<p>Got it. Thanks!</p>",
              "rawMarkdown": "Got it. Thanks!"
            }
          ]
        }
      ]
    },
    {
      "id": 2866366,
      "postDate": "2024-06-11T09:43:10.200Z",
      "content": "<p>Congratulations on winning the gold medal. I would like to ask why you thought of int8 quantization.</p>",
      "rawMarkdown": "Congratulations on winning the gold medal. I would like to ask why you thought of int8 quantization.",
      "replies": [
        {
          "id": 2866389,
          "postDate": "2024-06-11T09:59:35.740Z",
          "content": "<p>Thank you! The same to your team, Congratulations.<br>\nWell, my best single model's inference time is around 95min, which drove me to find some way to make it faster. And I came up with int8 quantization. <br>\nI searched the way to perform quantize and found that what you only need is a calibration dataset, and unlabeled soundscape provided by the host this year is just perfect for it!</p>",
          "rawMarkdown": "Thank you! The same to your team, Congratulations.\nWell, my best single model's inference time is around 95min, which drove me to find some way to make it faster. And I came up with int8 quantization. \nI searched the way to perform quantize and found that what you only need is a calibration dataset, and unlabeled soundscape provided by the host this year is just perfect for it!",
          "votes": 1,
          "replies": [
            {
              "id": 2866406,
              "postDate": "2024-06-11T10:06:41.317Z",
              "content": "<p>I understand, thanks for the reply!</p>",
              "rawMarkdown": "I understand, thanks for the reply!"
            }
          ]
        },
        {
          "id": 2866403,
          "postDate": "2024-06-11T10:05:58.200Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2866262,
      "postDate": "2024-06-11T08:24:42.193Z",
      "content": "<p>Maybe reduce training epochs when training with full data. But yeah, no significant improvement from my experience.</p>",
      "rawMarkdown": "Maybe reduce training epochs when training with full data. But yeah, no significant improvement from my experience.",
      "replies": [
        {
          "id": 2866293,
          "postDate": "2024-06-11T08:33:04.983Z",
          "content": "<p>I trained quite long epoch because I used knowledge distillation. <br>\nI think you are right, my CNN model shown significant LB decrease in long epoch…Sed model is more stable. Still, it is quite weird for me…</p>",
          "rawMarkdown": "I trained quite long epoch because I used knowledge distillation. \nI think you are right, my CNN model shown significant LB decrease in long epoch...Sed model is more stable. Still, it is quite weird for me...",
          "votes": 1
        },
        {
          "id": 2866312,
          "postDate": "2024-06-11T08:43:19.180Z",
          "content": "<p>I only try cnn models. I gauss different models have different fitting capabilities. Even for cnn models, different backbones perform differently, some(rexnet) can have really high CV, yet low leaderboard scores. </p>",
          "rawMarkdown": "I only try cnn models. I gauss different models have different fitting capabilities. Even for cnn models, different backbones perform differently, some(rexnet) can have really high CV, yet low leaderboard scores. ",
          "votes": 1,
          "replies": [
            {
              "id": 2866325,
              "postDate": "2024-06-11T09:05:03.363Z",
              "content": "<p>The same. For me, tf_efficientnet_b3ns performed really bad this year. Maybe it fits species not contained in test set better.</p>",
              "rawMarkdown": "The same. For me, tf_efficientnet_b3ns performed really bad this year. Maybe it fits species not contained in test set better."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2866276,
      "author_name": "Aaditya Porwal",
      "author_url": "",
      "post_date": "2024-06-11T08:28:28.797000",
      "content": "<p>Congratulations for gold medal !!!!!!!!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2866290,
          "author_name": "RihanPiggy",
          "author_url": "",
          "post_date": "2024-06-11T08:31:55.283000",
          "content": "<p>Thank you!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2867416,
      "author_name": "Ayush Solanki",
      "author_url": "",
      "post_date": "2024-06-11T19:49:56.873000",
      "content": "<p>Congratulations for gold!! I Had a beginner level doubt like are you feeding the model single channel spectrogram like (1,256,625) example shape or  3 channel by stacking the single channel?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2867620,
          "author_name": "RihanPiggy",
          "author_url": "",
          "post_date": "2024-06-12T01:28:24.803000",
          "content": "<p>Thank you!<br>\nMel spectrogram is single channel. And I used <a href=\"https://pytorch.org/audio/main/generated/torchaudio.functional.compute_deltas.html\" target=\"_blank\">compute delta</a> and stack to 3 channel.</p>\n<pre><code> ():\n    input_tensor = input_tensor.transpose(,)\n    input_tensor = compute_deltas(input_tensor)\n    input_tensor = input_tensor.transpose(,)\n     input_tensor\n\n\n ():\n    delta_1 = make_delta(x)\n    delta_2 = make_delta(delta_1)\n    x = torch.cat([x,delta_1,delta_2], dim=)\n     x\n</code></pre>",
          "votes": 2,
          "replies": [
            {
              "id": 2867848,
              "author_name": "Ayush Solanki",
              "author_url": "",
              "post_date": "2024-06-12T05:46:38.927000",
              "content": "<p>Thank you for response was searching for something like this from long time</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2866865,
      "author_name": "Bartley",
      "author_url": "",
      "post_date": "2024-06-11T14:44:03.743000",
      "content": "<p>Thanks for sharing your solution again, I learnt a lot from what you shared last year!</p>\n<ol>\n<li><p>You mention “Sampling every species equally”. Does this mean that you randomly sampled a species with weight <code>1/182</code>, then sampled a datapoint within that species, or does it mean that each data point was seen once each epoch?</p></li>\n<li><p>Also, could you clarify how you create the 10-sec segment during training? Is the same as 2023 (eg. 60sec waveform -&gt; split to 6 * 10sec -&gt; np.sum(audios,axis=0))?</p></li>\n</ol>",
      "votes": 0,
      "replies": [
        {
          "id": 2867625,
          "author_name": "RihanPiggy",
          "author_url": "",
          "post_date": "2024-06-12T01:38:11.410000",
          "content": "<p>Thank you!.</p>\n<ol>\n<li><p>Yes, the weight for each species is 1/182, but it is not a two step sampling. I used weighted sampler provided by pytorch. Let's assume there are 3 species (a,b,c). a has 3 samples, b has 2 samples, c has 1 sample. Then the weight for sample of a, b, c is 1/3, 1/3, 1/3, 1/2, 1/2, 1/1. You can find the weight for each species is the same here.</p></li>\n<li><p>For SED model, I simply randomly selected 10 sec from the audio clip. For audio extracted from unlabeled soundscape( I clipped 15sec),  it is 15sec waveform -&gt; 10sec, 5sec -&gt; 10sec, 10sec(pad 0) -&gt;  np.sum(audios,axis=0))</p></li>\n</ol>",
          "votes": 1,
          "replies": [
            {
              "id": 2868595,
              "author_name": "Bartley",
              "author_url": "",
              "post_date": "2024-06-12T14:00:27.670000",
              "content": "<p>Got it. Thanks!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2866366,
      "author_name": "Donghui Zhang",
      "author_url": "",
      "post_date": "2024-06-11T09:43:10.200000",
      "content": "<p>Congratulations on winning the gold medal. I would like to ask why you thought of int8 quantization.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2866389,
          "author_name": "RihanPiggy",
          "author_url": "",
          "post_date": "2024-06-11T09:59:35.740000",
          "content": "<p>Thank you! The same to your team, Congratulations.<br>\nWell, my best single model's inference time is around 95min, which drove me to find some way to make it faster. And I came up with int8 quantization. <br>\nI searched the way to perform quantize and found that what you only need is a calibration dataset, and unlabeled soundscape provided by the host this year is just perfect for it!</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2866406,
              "author_name": "Donghui Zhang",
              "author_url": "",
              "post_date": "2024-06-11T10:06:41.317000",
              "content": "<p>I understand, thanks for the reply!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2866403,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-06-11T10:05:58.200000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2866262,
      "author_name": "Aphysict",
      "author_url": "",
      "post_date": "2024-06-11T08:24:42.193000",
      "content": "<p>Maybe reduce training epochs when training with full data. But yeah, no significant improvement from my experience.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2866293,
          "author_name": "RihanPiggy",
          "author_url": "",
          "post_date": "2024-06-11T08:33:04.983000",
          "content": "<p>I trained quite long epoch because I used knowledge distillation. <br>\nI think you are right, my CNN model shown significant LB decrease in long epoch…Sed model is more stable. Still, it is quite weird for me…</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2866312,
          "author_name": "Aphysict",
          "author_url": "",
          "post_date": "2024-06-11T08:43:19.180000",
          "content": "<p>I only try cnn models. I gauss different models have different fitting capabilities. Even for cnn models, different backbones perform differently, some(rexnet) can have really high CV, yet low leaderboard scores. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2866325,
              "author_name": "RihanPiggy",
              "author_url": "",
              "post_date": "2024-06-11T09:05:03.363000",
              "content": "<p>The same. For me, tf_efficientnet_b3ns performed really bad this year. Maybe it fits species not contained in test set better.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2866078": "Congratulations to all the winners! Thanks to Kaggle and Cornell Lab of Ornithology for hosting this interesting competition.\n\nThis year's birdclef is really a hard one. Unlike padded CMAP which suppress the impact of species with very few positive labels, AUC treat every species equally and is very sensitive.\n\nI am happy to survive the final shake. Let me introduce my solution.\n\nThanks to every competitor who gave me inspiration. Special thanks to atsunorifujita, martynoveduard, loonypenguin, vladimirsydor, anonamename\n\n# Context\n\n- Business context: https://www.kaggle.com/competitions/birdclef-2024/overview\n- Data context: https://www.kaggle.com/competitions/birdclef-2024/data\n\n# Overview of the approach\n\nMy final submission is a combination of 2 SED model and 1 CNN model. The training scheme is almost the same as my [last year's 2nd place solution](https://www.kaggle.com/competitions/birdclef-2023/discussion/412707).\nThe new stuffs I added include\n\n- Extract training sample from soundscape using Birdnet and Bird-vocalization-classifier\n- Knowledge distillation proposed by [last year's 4th place](https://www.kaggle.com/competitions/birdclef-2023/discussion/412753)\n- Sumixup proposed by [last year's 7th place](https://www.kaggle.com/competitions/birdclef-2023/discussion/412922)\n- Int8 quantize\n- Split species to serveral subset (most important for me)\n\nThe key to win this competition is to split species to several subset and train the subset separately, which enables model to focus more on each species compared to train them all.\n\nMy final submission is an ensemble of 2 models with all species and 1 model with 66 species which have relatively small sample size.\nActually I had an experiment which split species to 3 subset, which gave me 0.690440(1st place) in private LB, but it was too risky to choose because I was not sure whether I was just overfitting the public LB.\n\n# Details of the submission\n\n## Models\n\n- SED with tf_efficientnetv2_s_in21k (all species)\n- SED with seresnext26t_32x4d (66 rare species)\n- CNN with resnet34d (all species)\n\n## Extra data on xeno-canto\n\nLike last year's situation, I collected extra audio from xeno-canto and trained a baseline SED v2s model. The baseline score was 0.65.\n\n## Extracting train sample from soundscape using Birdnet and Bird-vocalization-classifier\n\nBirdnet covers 181 of 182 species and Bird-vocalization-classifier covers 180 of 182 species. I mainly used birdnet and extracted niwpig1 with bird-vocalization-classifier. I extracted the 15 second audio clip with threshold 0.3.\n\nThis boosted my baseline score to 0.68.\n\n## Knowledge Distillation\n\nI implemented the same knowledge distillation scheme as last year's 4th place. This further boosted my baseline score to 0.70.\n\n## Further Extracting train sample from soundscape using trained models\n\nI used my own trained models to further extract train samples from soundscape, this gave me a little boost in LB.\n\n## Species subset\n\nInspired by [3rd place in birdclef 2022](https://www.kaggle.com/competitions/birdclef-2022/discussion/327193), I tried to train a model with 66 species which have relatively small sample size, and I found that this model outperformed model trained with all species on 43 of 66 species.\n\nI was not sure whether this will also happen in private LB, but I decided to ensemble 1 model trained on 66 species into my final submission.\n\nI also tried to split species into 3 subsets with some overlap. I trained them separately and then created a submission with 3 subset model ensemble. This gave me 0.690440(1st place) in private LB, thus maybe splitting birds to subset is the key to win. But it was too risky to choose because I was not sure whether I was just overfitting the public LB.\n\n## Int8 quantize\n\nThanks to the unlabeled soundscape, we can perform quantization this year.\n\nI used nncf to quantize encoder part of sed model. Quantizing CNN model led to 0.01 decrease in LB score, so I didn't perform quantization for CNN.\n\nQuantization reduced inference time by 20 minutes for sed v2s model. (100 minutes to 80 minutes)\n\n```python\nimport nncf\nimport openvino as ov\n\nnncf_dataset = nncf.Dataset(pytorch_dataset, transform_fn)\nmodel = ov.Core().read_model(path_to_model)\nquantized_model = nncf.quantize(\n    model, nncf_dataset,\n    target_device=nncf.TargetDevice.CPU,\n    subset_size=300,\n    fast_bias_correction=True,\n    preset=nncf.QuantizationPreset.MIXED,\n)\n\nov.save_model(quantized_model, quantized_model_path)\n```\n\n\n\n## ensemble strategy\n\n- logit average of 3 models\n- rank average of 3 models\n\n## Things didn't work for me.\n\nToo many...\n\n- Adding validation data to training led to 0.01 LB decrease. It is very weird, but I have to accept it and be faithful...\n- Calculating species weight to deal with label shift and covariance shift\n- Giving more weight to audios recorded in india according to geometric information (latitude 8~21, longitude 72.5 ~ 79)\n- changing weight of the sampler significantly decrease the LB. Sampling every species equally performed best\n- Filtering out 2024 train data duplicates not only by id but also by 'author + primary_label'(proposed [here](https://www.kaggle.com/competitions/birdclef-2023/discussion/412808))\n- using embedding extracted from birdnet and bird-vocalization-classifier proposed by [last year's 6th place](https://www.kaggle.com/competitions/birdclef-2023/discussion/412708)\n\n# Sources\n\n- https://www.kaggle.com/competitions/birdclef-2023/discussion/412707\n- https://www.kaggle.com/competitions/birdclef-2023/discussion/412753\n- https://www.kaggle.com/competitions/birdclef-2023/discussion/412922\n- https://www.kaggle.com/competitions/birdclef-2023/discussion/412808\n- https://www.kaggle.com/competitions/birdclef-2023/discussion/412708\n- https://www.kaggle.com/competitions/birdclef-2022/discussion/327193\n- https://docs.openvino.ai/2023.3/basic_quantization_flow.html\n",
    "2866276": "Congratulations for gold medal !!!!!!!!",
    "2867416": "Congratulations for gold!! I Had a beginner level doubt like are you feeding the model single channel spectrogram like (1,256,625) example shape or  3 channel by stacking the single channel?",
    "2866865": "Thanks for sharing your solution again, I learnt a lot from what you shared last year!\n\n1. You mention “Sampling every species equally”. Does this mean that you randomly sampled a species with weight `1/182`, then sampled a datapoint within that species, or does it mean that each data point was seen once each epoch?\n\n2. Also, could you clarify how you create the 10-sec segment during training? Is the same as 2023 (eg. 60sec waveform -> split to 6 * 10sec -> np.sum(audios,axis=0))?",
    "2866366": "Congratulations on winning the gold medal. I would like to ask why you thought of int8 quantization.",
    "2866262": "Maybe reduce training epochs when training with full data. But yeah, no significant improvement from my experience."
  }
}