{
  "id": 412768,
  "title": "12th place solution: 8 CNN models ensemble with OpenVINO",
  "url": "/competitions/birdclef-2023/discussion/412768",
  "author_name": "",
  "post_date": "2023-05-25T06:00:51.437369Z",
  "votes": 21,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Many thanks to Kaggle team, organizers from Cornell Lab of Ornithology and all participants. </p>\n<p>As a newbie in both deep learning and audio signal processing, I am glad that I secure the gold medal after the final shake up (barely😂). </p>\n<p>My final solution is an ensemble of eight different models trained with different backbones or spectrogram configurations. Each of those models is based on the <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">BirdCLEF 2021 2nd place solution</a>, and inspired by other reimplementations in BirdCLEF 2022. </p>\n<h2>Model Architecture</h2>\n<p>I use the CNN architecture in the BirdCLEF 2021 2nd place solution. As mentioned, the final solution includes four different backbones: </p>\n<ul>\n<li><code>seresnext26t_32x4d</code></li>\n<li><code>tf_efficientnetv2_s_in21k</code></li>\n<li><code>tf_efficientnet_b3_ns</code></li>\n<li><code>eca_nfnet_l0</code></li>\n</ul>\n<p>For each backbone, two different spectrogram configurations are used, the same as the <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/326979\" target=\"_blank\">11th place solution</a> last year.</p>\n<h2>Data augmentations</h2>\n<p>To improve the performance of the single model, I implement both waveform and spec augmentations. <br>\nWaveform augmentations:</p>\n<ul>\n<li>AddBackgroundNoise (Besides the popular noise datasets, I also create a <a href=\"https://www.kaggle.com/code/aphysict/2022test-nocall/notebook\" target=\"_blank\">dataset</a> from last year's test soundscapes)</li>\n<li>Gain</li>\n<li>Shift</li>\n</ul>\n<p>Spec augmentations:</p>\n<ul>\n<li>mixup and cutmix</li>\n<li><a href=\"https://github.com/frednam93/FilterAugSED/tree/main\" target=\"_blank\">FilterAugment</a> which applies random weights on randomly selected frequency bands. (This performs better than LowPassFilter on public LB but worse on private LB)</li>\n<li>time and freq mask</li>\n</ul>\n<h2>Ensemble</h2>\n<p>In this competition, I spend a lot of effort to speed-up the infer speed of the single model. And finally, I implement a pytorch-ONNX-OpenVINO workflow on kaggle, that can bring the infer time of the single model to ~15min.</p>\n<p>The first thing to do is to replace <code>ta.transforms.MelSpectrogram</code> with a customized <code>Spectrogram</code> from <a href=\"https://github.com/adobe-research/convmelspec\" target=\"_blank\">Convmelspec</a>, because onnx doesn't support STFT operator on python 3.7. Also remember to use fixed size in <code>avg_pool2d</code>. </p>\n<p>Then transform the pytorch model to ONNX, and convert the ONNX model to OpenVINO. The interesting thing is, this does not speed the infer up if you run the infer on the 200 soundscapes in the notebook. However, OpenVINO does speed things up on the submission end by ~50%. I guess the intel skylake cpu has some extra power with OpenVINO. </p>\n<p><a href=\"https://www.kaggle.com/aphysict/fork-of-birdclef-submit-rewrite-onnx\" target=\"_blank\">Here</a> is my best single model submission notebook (private LB 0.72463), which includes the conversion and the infer. The final submission only includes the infer part. </p>\n<p>Overall, my solution has no fancy tricks other than the OpenVINO part. The performance gain from OpenVINO also surprises me, and leaves me a very short time to build an 8 models ensemble. I am looking forward to seeing other top solutions and learning more about deep learning! </p>",
  "messages": [
    {
      "id": "2273350",
      "postDate": "05/25/2023 06:00:51",
      "content": "<p>Many thanks to Kaggle team, organizers from Cornell Lab of Ornithology and all participants. </p>\n<p>As a newbie in both deep learning and audio signal processing, I am glad that I secure the gold medal after the final shake up (barely😂). </p>\n<p>My final solution is an ensemble of eight different models trained with different backbones or spectrogram configurations. Each of those models is based on the <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">BirdCLEF 2021 2nd place solution</a>, and inspired by other reimplementations in BirdCLEF 2022. </p>\n<h2>Model Architecture</h2>\n<p>I use the CNN architecture in the BirdCLEF 2021 2nd place solution. As mentioned, the final solution includes four different backbones: </p>\n<ul>\n<li><code>seresnext26t_32x4d</code></li>\n<li><code>tf_efficientnetv2_s_in21k</code></li>\n<li><code>tf_efficientnet_b3_ns</code></li>\n<li><code>eca_nfnet_l0</code></li>\n</ul>\n<p>For each backbone, two different spectrogram configurations are used, the same as the <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/326979\" target=\"_blank\">11th place solution</a> last year.</p>\n<h2>Data augmentations</h2>\n<p>To improve the performance of the single model, I implement both waveform and spec augmentations. <br>\nWaveform augmentations:</p>\n<ul>\n<li>AddBackgroundNoise (Besides the popular noise datasets, I also create a <a href=\"https://www.kaggle.com/code/aphysict/2022test-nocall/notebook\" target=\"_blank\">dataset</a> from last year's test soundscapes)</li>\n<li>Gain</li>\n<li>Shift</li>\n</ul>\n<p>Spec augmentations:</p>\n<ul>\n<li>mixup and cutmix</li>\n<li><a href=\"https://github.com/frednam93/FilterAugSED/tree/main\" target=\"_blank\">FilterAugment</a> which applies random weights on randomly selected frequency bands. (This performs better than LowPassFilter on public LB but worse on private LB)</li>\n<li>time and freq mask</li>\n</ul>\n<h2>Ensemble</h2>\n<p>In this competition, I spend a lot of effort to speed-up the infer speed of the single model. And finally, I implement a pytorch-ONNX-OpenVINO workflow on kaggle, that can bring the infer time of the single model to ~15min.</p>\n<p>The first thing to do is to replace <code>ta.transforms.MelSpectrogram</code> with a customized <code>Spectrogram</code> from <a href=\"https://github.com/adobe-research/convmelspec\" target=\"_blank\">Convmelspec</a>, because onnx doesn't support STFT operator on python 3.7. Also remember to use fixed size in <code>avg_pool2d</code>. </p>\n<p>Then transform the pytorch model to ONNX, and convert the ONNX model to OpenVINO. The interesting thing is, this does not speed the infer up if you run the infer on the 200 soundscapes in the notebook. However, OpenVINO does speed things up on the submission end by ~50%. I guess the intel skylake cpu has some extra power with OpenVINO. </p>\n<p><a href=\"https://www.kaggle.com/aphysict/fork-of-birdclef-submit-rewrite-onnx\" target=\"_blank\">Here</a> is my best single model submission notebook (private LB 0.72463), which includes the conversion and the infer. The final submission only includes the infer part. </p>\n<p>Overall, my solution has no fancy tricks other than the OpenVINO part. The performance gain from OpenVINO also surprises me, and leaves me a very short time to build an 8 models ensemble. I am looking forward to seeing other top solutions and learning more about deep learning! </p>",
      "rawMarkdown": "Many thanks to Kaggle team, organizers from Cornell Lab of Ornithology and all participants. \n\nAs a newbie in both deep learning and audio signal processing, I am glad that I secure the gold medal after the final shake up (barely😂). \n\nMy final solution is an ensemble of eight different models trained with different backbones or spectrogram configurations. Each of those models is based on the [BirdCLEF 2021 2nd place solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463), and inspired by other reimplementations in BirdCLEF 2022. \n\n## Model Architecture\nI use the CNN architecture in the BirdCLEF 2021 2nd place solution. As mentioned, the final solution includes four different backbones: \n- `seresnext26t_32x4d`\n- `tf_efficientnetv2_s_in21k`\n- `tf_efficientnet_b3_ns`\n- `eca_nfnet_l0`\n\nFor each backbone, two different spectrogram configurations are used, the same as the [11th place solution](https://www.kaggle.com/competitions/birdclef-2022/discussion/326979) last year.\n\n## Data augmentations\nTo improve the performance of the single model, I implement both waveform and spec augmentations. \nWaveform augmentations:\n- AddBackgroundNoise (Besides the popular noise datasets, I also create a [dataset](https://www.kaggle.com/code/aphysict/2022test-nocall/notebook) from last year's test soundscapes)\n- Gain\n- Shift\n\nSpec augmentations:\n- mixup and cutmix\n- [FilterAugment](https://github.com/frednam93/FilterAugSED/tree/main) which applies random weights on randomly selected frequency bands. (This performs better than LowPassFilter on public LB but worse on private LB)\n- time and freq mask\n\n## Ensemble\nIn this competition, I spend a lot of effort to speed-up the infer speed of the single model. And finally, I implement a pytorch-ONNX-OpenVINO workflow on kaggle, that can bring the infer time of the single model to ~15min.\n\nThe first thing to do is to replace `ta.transforms.MelSpectrogram` with a customized `Spectrogram` from [Convmelspec](https://github.com/adobe-research/convmelspec), because onnx doesn't support STFT operator on python 3.7. Also remember to use fixed size in `avg_pool2d`. \n\nThen transform the pytorch model to ONNX, and convert the ONNX model to OpenVINO. The interesting thing is, this does not speed the infer up if you run the infer on the 200 soundscapes in the notebook. However, OpenVINO does speed things up on the submission end by ~50%. I guess the intel skylake cpu has some extra power with OpenVINO. \n\n[Here](https://www.kaggle.com/aphysict/fork-of-birdclef-submit-rewrite-onnx) is my best single model submission notebook (private LB 0.72463), which includes the conversion and the infer. The final submission only includes the infer part. \n\nOverall, my solution has no fancy tricks other than the OpenVINO part. The performance gain from OpenVINO also surprises me, and leaves me a very short time to build an 8 models ensemble. I am looking forward to seeing other top solutions and learning more about deep learning!",
      "votes": null
    },
    {
      "id": "2273497",
      "postDate": "05/25/2023 07:34:03",
      "content": "<p>Congratulations for solo gold! My best single model scored 0.7218 (private LB) which is slightly worse than your best single model. My initial ideas are ensemble several CNNs models and several PANNS models, but I have no ideas how to ensemble more models so that I just ensemble 3 PANNS models. My final submission only scored 0.7333(private LB). Seems inference acceleration is an import part to win the competition. Anyway, Congratulations for solo gold!</p>",
      "rawMarkdown": "Congratulations for solo gold! My best single model scored 0.7218 (private LB) which is slightly worse than your best single model. My initial ideas are ensemble several CNNs models and several PANNS models, but I have no ideas how to ensemble more models so that I just ensemble 3 PANNS models. My final submission only scored 0.7333(private LB). Seems inference acceleration is an import part to win the competition. Anyway, Congratulations for solo gold!",
      "votes": null
    },
    {
      "id": "2273549",
      "postDate": "05/25/2023 08:15:13",
      "content": "<p>That's amazing! Thanks for sharing and Congratulations 🎉💥</p>",
      "rawMarkdown": "That's amazing! Thanks for sharing and Congratulations 🎉💥",
      "votes": null
    },
    {
      "id": "2273576",
      "postDate": "05/25/2023 08:37:11",
      "content": "<p>Thanks, that makes two of us. Without openvino I can only ensemble 3 CNN models, so I thought I should focus on CNN models. Then I found openvino, however I didn't have enough time to learn PANNS models. I would definitely try powerful PANNS next time! <br>\nPs: I had a conference in your school three years ago. it's a beautiful university!</p>",
      "rawMarkdown": "Thanks, that makes two of us. Without openvino I can only ensemble 3 CNN models, so I thought I should focus on CNN models. Then I found openvino, however I didn't have enough time to learn PANNS models. I would definitely try powerful PANNS next time! \nPs: I had a conference in your school three years ago. it's a beautiful university!",
      "votes": null
    },
    {
      "id": "2275056",
      "postDate": "05/26/2023 13:06:25",
      "content": "<p>Thanks for the helpful description! Could you also post a link to your training code, e.g. on github?</p>",
      "rawMarkdown": "Thanks for the helpful description! Could you also post a link to your training code, e.g. on github?",
      "votes": null
    },
    {
      "id": "2276679",
      "postDate": "05/27/2023 05:57:03",
      "content": "<p>Thanks for your comment. I was using Jupyterlab on a sever in this comp, so the code is a little dirty. I suggest you to take a look this <a href=\"https://github.com/dazzle-me/birdclef-2022-3rd-place-solution\" target=\"_blank\">github repository</a>, because ~80% of codes are the same and the difference largely comes from the validation metric and the infer part. </p>",
      "rawMarkdown": "Thanks for your comment. I was using Jupyterlab on a sever in this comp, so the code is a little dirty. I suggest you to take a look this [github repository](https://github.com/dazzle-me/birdclef-2022-3rd-place-solution), because ~80% of codes are the same and the difference largely comes from the validation metric and the infer part.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2273497,
      "author_name": "shtljw",
      "author_url": "",
      "post_date": "05/25/2023 07:34:03",
      "content": "<p>Congratulations for solo gold! My best single model scored 0.7218 (private LB) which is slightly worse than your best single model. My initial ideas are ensemble several CNNs models and several PANNS models, but I have no ideas how to ensemble more models so that I just ensemble 3 PANNS models. My final submission only scored 0.7333(private LB). Seems inference acceleration is an import part to win the competition. Anyway, Congratulations for solo gold!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2273576,
          "author_name": "aphysict",
          "author_url": "",
          "post_date": "05/25/2023 08:37:11",
          "content": "<p>Thanks, that makes two of us. Without openvino I can only ensemble 3 CNN models, so I thought I should focus on CNN models. Then I found openvino, however I didn't have enough time to learn PANNS models. I would definitely try powerful PANNS next time! <br>\nPs: I had a conference in your school three years ago. it's a beautiful university!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2273549,
      "author_name": "scipygaurav",
      "author_url": "",
      "post_date": "05/25/2023 08:15:13",
      "content": "<p>That's amazing! Thanks for sharing and Congratulations 🎉💥</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2275056,
      "author_name": "janhuus",
      "author_url": "",
      "post_date": "05/26/2023 13:06:25",
      "content": "<p>Thanks for the helpful description! Could you also post a link to your training code, e.g. on github?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2276679,
          "author_name": "aphysict",
          "author_url": "",
          "post_date": "05/27/2023 05:57:03",
          "content": "<p>Thanks for your comment. I was using Jupyterlab on a sever in this comp, so the code is a little dirty. I suggest you to take a look this <a href=\"https://github.com/dazzle-me/birdclef-2022-3rd-place-solution\" target=\"_blank\">github repository</a>, because ~80% of codes are the same and the difference largely comes from the validation metric and the infer part. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2273350": "Many thanks to Kaggle team, organizers from Cornell Lab of Ornithology and all participants. \n\nAs a newbie in both deep learning and audio signal processing, I am glad that I secure the gold medal after the final shake up (barely😂). \n\nMy final solution is an ensemble of eight different models trained with different backbones or spectrogram configurations. Each of those models is based on the [BirdCLEF 2021 2nd place solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463), and inspired by other reimplementations in BirdCLEF 2022. \n\n## Model Architecture\nI use the CNN architecture in the BirdCLEF 2021 2nd place solution. As mentioned, the final solution includes four different backbones: \n- `seresnext26t_32x4d`\n- `tf_efficientnetv2_s_in21k`\n- `tf_efficientnet_b3_ns`\n- `eca_nfnet_l0`\n\nFor each backbone, two different spectrogram configurations are used, the same as the [11th place solution](https://www.kaggle.com/competitions/birdclef-2022/discussion/326979) last year.\n\n## Data augmentations\nTo improve the performance of the single model, I implement both waveform and spec augmentations. \nWaveform augmentations:\n- AddBackgroundNoise (Besides the popular noise datasets, I also create a [dataset](https://www.kaggle.com/code/aphysict/2022test-nocall/notebook) from last year's test soundscapes)\n- Gain\n- Shift\n\nSpec augmentations:\n- mixup and cutmix\n- [FilterAugment](https://github.com/frednam93/FilterAugSED/tree/main) which applies random weights on randomly selected frequency bands. (This performs better than LowPassFilter on public LB but worse on private LB)\n- time and freq mask\n\n## Ensemble\nIn this competition, I spend a lot of effort to speed-up the infer speed of the single model. And finally, I implement a pytorch-ONNX-OpenVINO workflow on kaggle, that can bring the infer time of the single model to ~15min.\n\nThe first thing to do is to replace `ta.transforms.MelSpectrogram` with a customized `Spectrogram` from [Convmelspec](https://github.com/adobe-research/convmelspec), because onnx doesn't support STFT operator on python 3.7. Also remember to use fixed size in `avg_pool2d`. \n\nThen transform the pytorch model to ONNX, and convert the ONNX model to OpenVINO. The interesting thing is, this does not speed the infer up if you run the infer on the 200 soundscapes in the notebook. However, OpenVINO does speed things up on the submission end by ~50%. I guess the intel skylake cpu has some extra power with OpenVINO. \n\n[Here](https://www.kaggle.com/aphysict/fork-of-birdclef-submit-rewrite-onnx) is my best single model submission notebook (private LB 0.72463), which includes the conversion and the infer. The final submission only includes the infer part. \n\nOverall, my solution has no fancy tricks other than the OpenVINO part. The performance gain from OpenVINO also surprises me, and leaves me a very short time to build an 8 models ensemble. I am looking forward to seeing other top solutions and learning more about deep learning!",
    "2273497": "Congratulations for solo gold! My best single model scored 0.7218 (private LB) which is slightly worse than your best single model. My initial ideas are ensemble several CNNs models and several PANNS models, but I have no ideas how to ensemble more models so that I just ensemble 3 PANNS models. My final submission only scored 0.7333(private LB). Seems inference acceleration is an import part to win the competition. Anyway, Congratulations for solo gold!",
    "2273549": "That's amazing! Thanks for sharing and Congratulations 🎉💥",
    "2273576": "Thanks, that makes two of us. Without openvino I can only ensemble 3 CNN models, so I thought I should focus on CNN models. Then I found openvino, however I didn't have enough time to learn PANNS models. I would definitely try powerful PANNS next time! \nPs: I had a conference in your school three years ago. it's a beautiful university!",
    "2275056": "Thanks for the helpful description! Could you also post a link to your training code, e.g. on github?",
    "2276679": "Thanks for your comment. I was using Jupyterlab on a sever in this comp, so the code is a little dirty. I suggest you to take a look this [github repository](https://github.com/dazzle-me/birdclef-2022-3rd-place-solution), because ~80% of codes are the same and the difference largely comes from the validation metric and the infer part."
  },
  "source": "meta"
}