{
  "id": 412794,
  "title": "9th Place Solution: 7 CNN Models Ensemble",
  "url": "/competitions/birdclef-2023/writeups/synergy-9th-place-solution-7-cnn-models-ensemble",
  "author_name": "",
  "post_date": "2023-05-25T09:09:37.669516800Z",
  "votes": 26,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Thanks to the organizers for hosting BirdClef 2023 again this year and giving us a learning experience.</p>\n<p>Heartfelt Congratulations to the winners and everyone who got to learn and experience.</p>\n<p>I would also like to express my gratitude to my talented teammates <a href=\"https://www.kaggle.com/ivanaerlic\" target=\"_blank\">@ivanaerlic</a> and <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a></p>\n<h1>Training</h1>\n<p>We used efficientnet_b0, eca_nfnet_l0 and convnext_tiny architectures in our final submission.</p>\n<p>The training was done in multiple rounds using datasets from all previous years. All models were trained on 5s clips</p>\n<p>For eca_nfnet_l0 and convnext_tiny<br>\nHop size 320, n_mels 64, image_size = 64x501<br>\nFor efficientnet_b0<br>\nHop size 320, n_mels 64, image_size (Bilinear image size increase) = 128x1002</p>\n<p>The training was done in multiple rounds:</p>\n<p><strong>Round 1:</strong></p>\n<p>Hop size 512, n_mels 64, image_size = 64x313<br>\n(Bilinear image size increase) = 128x636</p>\n<p>Models : Eca nfnet l0 (128x636), convnext tiny (128x636).</p>\n<p>Training data: 2022 &amp; 2023</p>\n<p>Augmentations: HFlip, Random Cutout, Pink Noise, Gaussian Noise, Noise Injection, Random Volume</p>\n<p>Type of Training: Competition Labels</p>\n<p>CV: ~.79 (Low)</p>\n<p><strong>Round 2:</strong></p>\n<p>Hop size 320, n_mels 64, image_size = 64x501</p>\n<p>Models : Efficientnet B0 (128x1002 Bilinear increase), Eca nfnet l0 (64x501), convnext tiny (64x501).</p>\n<p>Augmentations: HFlip, Random Cutout, Pink Noise, Gaussian Noise, Noise Injection, Random Volume</p>\n<p>Training data: 2022 &amp; 2023</p>\n<p>Type of Training: Distillation</p>\n<p>CV: ~.79 (Slight increase)</p>\n<p><strong>Round 3:</strong></p>\n<p>Hop size 320, n_mels 64, image_size = 64x501</p>\n<p>Models : Efficientnet B0 (128x1002 Bilinear increase), Eca nfnet l0 (64x501), convnext tiny (64x501).</p>\n<p>Training data: 2021 &amp; 2022 &amp; 2023</p>\n<p>Extra methods: AWP, SWA, Attention Head</p>\n<p>Augmentations: HFlip, Random Cutout, Pixel Dropout</p>\n<p>Type of Training: Competition Labels</p>\n<p>CV: ~.8</p>\n<p><strong>Round 4:</strong></p>\n<p>Hop size 320, n_mels 64, image_size = 64x501</p>\n<p>Models : Efficientnet B0 (128x1002 Bilinear increase), Eca nfnet l0 (64x501), convnext tiny (64x501).</p>\n<p>Training data: 2023</p>\n<p>Augmentations: HFlip, Random Cutout, Pixel Dropout, Pink Noise, Gaussian Noise, Noise Injection, Random Volume</p>\n<p>Extra methods: SWA, Attention Head</p>\n<p>Type of Training: Pseudo Labels </p>\n<p>CV: ~.81 (mild leak) </p>\n<p>Our CV was the first 5 seconds of the clips, this CV although was not always consistent with LB, gave us a good enough idea of what we could expect.</p>\n<p>Pseudo labeling improved our score a lot. We used pseudo-labels to make new secondary labels and gave them a hard probability 0.5</p>\n<p>For convnext model, it was first fine tuned to increase the score on the targets where secondary label exists, although this sacrificed the primary label score, this adds diversity to our ensembles</p>\n<h1>Inference</h1>\n<p>We used OpenVino to increase the inference speed of our models</p>\n<p>We also saw that loading an audio and giving it in batches, the first batch of inference is especially slower than the rest, so we made it so that it loads 8 audios at once, makes a data loader for all 8 audios and then forward pass the model, doing this we increased the inference speed further by 20%.</p>\n<p>So we were able to submit 7 models, 2 of efficientnet_b0 + 2 of eca_nfnet_l0 + 3 of convnext_tiny</p>\n<p>We also noticed that sigmoid after weight averaging was more effective than sigmoid before weight averaging, also we were using softmax activation which gave a better score than sigmoid up until we reached top of .82 (public LB), after that sigmoid started to work better.</p>\n<h1>BirdNet</h1>\n<p>Our best private submission turned out to be when we ensembled our models with BirdNet, although it was like our top 10 submissions of public and did give us a lot of variance, we did not select this because we were a little afraid of it being a dice roll on private LB.</p>\n<h1>Post Processing</h1>\n<p>Looking at the last year’s competition, we tried out increasing the confidence of a class by a factor of 1.5 if any prediction of that class in the audio file is more than .96</p>\n<p>This did work on a lot of submissions and was in fact amongst our top subs, but in the end, it did not play well on private LB and we did not select it because of the risk of overfitting the public LB.</p>\n<h1>What did not work</h1>\n<p>SED models<br>\nQ-Transform<br>\nMore than 5s training<br>\nRolling Mean with last and next 5s segment</p>\n<p>More than 60 submissions of ours had timeout errors and on the last second day of the competition, 4 of 5 submissions had encountered timeout error while before that the chance of getting a timeout was only 20-30% in the same notebook.</p>",
  "messages": [
    {
      "id": "2273599",
      "postDate": "05/25/2023 09:09:37",
      "content": "<p>Thanks to the organizers for hosting BirdClef 2023 again this year and giving us a learning experience.</p>\n<p>Heartfelt Congratulations to the winners and everyone who got to learn and experience.</p>\n<p>I would also like to express my gratitude to my talented teammates <a href=\"https://www.kaggle.com/ivanaerlic\" target=\"_blank\">@ivanaerlic</a> and <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a></p>\n<h1>Training</h1>\n<p>We used efficientnet_b0, eca_nfnet_l0 and convnext_tiny architectures in our final submission.</p>\n<p>The training was done in multiple rounds using datasets from all previous years. All models were trained on 5s clips</p>\n<p>For eca_nfnet_l0 and convnext_tiny<br>\nHop size 320, n_mels 64, image_size = 64x501<br>\nFor efficientnet_b0<br>\nHop size 320, n_mels 64, image_size (Bilinear image size increase) = 128x1002</p>\n<p>The training was done in multiple rounds:</p>\n<p><strong>Round 1:</strong></p>\n<p>Hop size 512, n_mels 64, image_size = 64x313<br>\n(Bilinear image size increase) = 128x636</p>\n<p>Models : Eca nfnet l0 (128x636), convnext tiny (128x636).</p>\n<p>Training data: 2022 &amp; 2023</p>\n<p>Augmentations: HFlip, Random Cutout, Pink Noise, Gaussian Noise, Noise Injection, Random Volume</p>\n<p>Type of Training: Competition Labels</p>\n<p>CV: ~.79 (Low)</p>\n<p><strong>Round 2:</strong></p>\n<p>Hop size 320, n_mels 64, image_size = 64x501</p>\n<p>Models : Efficientnet B0 (128x1002 Bilinear increase), Eca nfnet l0 (64x501), convnext tiny (64x501).</p>\n<p>Augmentations: HFlip, Random Cutout, Pink Noise, Gaussian Noise, Noise Injection, Random Volume</p>\n<p>Training data: 2022 &amp; 2023</p>\n<p>Type of Training: Distillation</p>\n<p>CV: ~.79 (Slight increase)</p>\n<p><strong>Round 3:</strong></p>\n<p>Hop size 320, n_mels 64, image_size = 64x501</p>\n<p>Models : Efficientnet B0 (128x1002 Bilinear increase), Eca nfnet l0 (64x501), convnext tiny (64x501).</p>\n<p>Training data: 2021 &amp; 2022 &amp; 2023</p>\n<p>Extra methods: AWP, SWA, Attention Head</p>\n<p>Augmentations: HFlip, Random Cutout, Pixel Dropout</p>\n<p>Type of Training: Competition Labels</p>\n<p>CV: ~.8</p>\n<p><strong>Round 4:</strong></p>\n<p>Hop size 320, n_mels 64, image_size = 64x501</p>\n<p>Models : Efficientnet B0 (128x1002 Bilinear increase), Eca nfnet l0 (64x501), convnext tiny (64x501).</p>\n<p>Training data: 2023</p>\n<p>Augmentations: HFlip, Random Cutout, Pixel Dropout, Pink Noise, Gaussian Noise, Noise Injection, Random Volume</p>\n<p>Extra methods: SWA, Attention Head</p>\n<p>Type of Training: Pseudo Labels </p>\n<p>CV: ~.81 (mild leak) </p>\n<p>Our CV was the first 5 seconds of the clips, this CV although was not always consistent with LB, gave us a good enough idea of what we could expect.</p>\n<p>Pseudo labeling improved our score a lot. We used pseudo-labels to make new secondary labels and gave them a hard probability 0.5</p>\n<p>For convnext model, it was first fine tuned to increase the score on the targets where secondary label exists, although this sacrificed the primary label score, this adds diversity to our ensembles</p>\n<h1>Inference</h1>\n<p>We used OpenVino to increase the inference speed of our models</p>\n<p>We also saw that loading an audio and giving it in batches, the first batch of inference is especially slower than the rest, so we made it so that it loads 8 audios at once, makes a data loader for all 8 audios and then forward pass the model, doing this we increased the inference speed further by 20%.</p>\n<p>So we were able to submit 7 models, 2 of efficientnet_b0 + 2 of eca_nfnet_l0 + 3 of convnext_tiny</p>\n<p>We also noticed that sigmoid after weight averaging was more effective than sigmoid before weight averaging, also we were using softmax activation which gave a better score than sigmoid up until we reached top of .82 (public LB), after that sigmoid started to work better.</p>\n<h1>BirdNet</h1>\n<p>Our best private submission turned out to be when we ensembled our models with BirdNet, although it was like our top 10 submissions of public and did give us a lot of variance, we did not select this because we were a little afraid of it being a dice roll on private LB.</p>\n<h1>Post Processing</h1>\n<p>Looking at the last year’s competition, we tried out increasing the confidence of a class by a factor of 1.5 if any prediction of that class in the audio file is more than .96</p>\n<p>This did work on a lot of submissions and was in fact amongst our top subs, but in the end, it did not play well on private LB and we did not select it because of the risk of overfitting the public LB.</p>\n<h1>What did not work</h1>\n<p>SED models<br>\nQ-Transform<br>\nMore than 5s training<br>\nRolling Mean with last and next 5s segment</p>\n<p>More than 60 submissions of ours had timeout errors and on the last second day of the competition, 4 of 5 submissions had encountered timeout error while before that the chance of getting a timeout was only 20-30% in the same notebook.</p>",
      "rawMarkdown": "Thanks to the organizers for hosting BirdClef 2023 again this year and giving us a learning experience.\n\nHeartfelt Congratulations to the winners and everyone who got to learn and experience.\n\nI would also like to express my gratitude to my talented teammates @ivanaerlic and @nischaydnk\n\n# Training\n\nWe used efficientnet_b0, eca_nfnet_l0 and convnext_tiny architectures in our final submission.\n\nThe training was done in multiple rounds using datasets from all previous years. All models were trained on 5s clips\n\nFor eca_nfnet_l0 and convnext_tiny\nHop size 320, n_mels 64, image_size = 64x501\nFor efficientnet_b0\nHop size 320, n_mels 64, image_size (Bilinear image size increase) = 128x1002\n\nThe training was done in multiple rounds:\n\n**Round 1:**\n\nHop size 512, n_mels 64, image_size = 64x313\n(Bilinear image size increase) = 128x636\n\nModels : Eca nfnet l0 (128x636), convnext tiny (128x636).\n\nTraining data: 2022 & 2023\n\nAugmentations: HFlip, Random Cutout, Pink Noise, Gaussian Noise, Noise Injection, Random Volume\n\nType of Training: Competition Labels\n\nCV: ~.79 (Low)\n\n**Round 2:**\n\nHop size 320, n_mels 64, image_size = 64x501\n\nModels : Efficientnet B0 (128x1002 Bilinear increase), Eca nfnet l0 (64x501), convnext tiny (64x501).\n\nAugmentations: HFlip, Random Cutout, Pink Noise, Gaussian Noise, Noise Injection, Random Volume\n\nTraining data: 2022 & 2023\n\nType of Training: Distillation\n\nCV: ~.79 (Slight increase)\n\n**Round 3:**\n\nHop size 320, n_mels 64, image_size = 64x501\n\nModels : Efficientnet B0 (128x1002 Bilinear increase), Eca nfnet l0 (64x501), convnext tiny (64x501).\n\nTraining data: 2021 & 2022 & 2023\n\nExtra methods: AWP, SWA, Attention Head\n\nAugmentations: HFlip, Random Cutout, Pixel Dropout\n\nType of Training: Competition Labels\n\nCV: ~.8\n\n**Round 4:**\n\nHop size 320, n_mels 64, image_size = 64x501\n\nModels : Efficientnet B0 (128x1002 Bilinear increase), Eca nfnet l0 (64x501), convnext tiny (64x501).\n\nTraining data: 2023\n\nAugmentations: HFlip, Random Cutout, Pixel Dropout, Pink Noise, Gaussian Noise, Noise Injection, Random Volume\n\nExtra methods: SWA, Attention Head\n\nType of Training: Pseudo Labels \n\nCV: ~.81 (mild leak) \n\nOur CV was the first 5 seconds of the clips, this CV although was not always consistent with LB, gave us a good enough idea of what we could expect.\n\nPseudo labeling improved our score a lot. We used pseudo-labels to make new secondary labels and gave them a hard probability 0.5\n\nFor convnext model, it was first fine tuned to increase the score on the targets where secondary label exists, although this sacrificed the primary label score, this adds diversity to our ensembles\n\n# Inference\n\nWe used OpenVino to increase the inference speed of our models\n\nWe also saw that loading an audio and giving it in batches, the first batch of inference is especially slower than the rest, so we made it so that it loads 8 audios at once, makes a data loader for all 8 audios and then forward pass the model, doing this we increased the inference speed further by 20%.\n\nSo we were able to submit 7 models, 2 of efficientnet_b0 + 2 of eca_nfnet_l0 + 3 of convnext_tiny\n\nWe also noticed that sigmoid after weight averaging was more effective than sigmoid before weight averaging, also we were using softmax activation which gave a better score than sigmoid up until we reached top of .82 (public LB), after that sigmoid started to work better.\n\n# BirdNet\n\nOur best private submission turned out to be when we ensembled our models with BirdNet, although it was like our top 10 submissions of public and did give us a lot of variance, we did not select this because we were a little afraid of it being a dice roll on private LB.\n\n# Post Processing\n\nLooking at the last year’s competition, we tried out increasing the confidence of a class by a factor of 1.5 if any prediction of that class in the audio file is more than .96\n\nThis did work on a lot of submissions and was in fact amongst our top subs, but in the end, it did not play well on private LB and we did not select it because of the risk of overfitting the public LB.\n\n# What did not work\n\nSED models\nQ-Transform\nMore than 5s training\nRolling Mean with last and next 5s segment\n\nMore than 60 submissions of ours had timeout errors and on the last second day of the competition, 4 of 5 submissions had encountered timeout error while before that the chance of getting a timeout was only 20-30% in the same notebook.",
      "votes": null
    },
    {
      "id": "2273615",
      "postDate": "05/25/2023 09:30:14",
      "content": "<p>Congratulations! Here again, OpenVino. Speeding up inference is a great learning experience for me.</p>",
      "rawMarkdown": "Congratulations! Here again, OpenVino. Speeding up inference is a great learning experience for me.",
      "votes": null
    },
    {
      "id": "2273631",
      "postDate": "05/25/2023 10:00:39",
      "content": "<p>Congratulations and thanks for sharing! What is the private LB score for single BirdNET ? I am wondering if I use BirdNET correctly.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing! What is the private LB score for single BirdNET ? I am wondering if I use BirdNET correctly.",
      "votes": null
    },
    {
      "id": "2273641",
      "postDate": "05/25/2023 10:06:08",
      "content": "<p>A lot of people on Public .82 will get to .83 if they use BirdNet, I think it would be a 0.005 increase on average between both public and private… atleast that is what we observed</p>",
      "rawMarkdown": "A lot of people on Public .82 will get to .83 if they use BirdNet, I think it would be a 0.005 increase on average between both public and private... atleast that is what we observed",
      "votes": null
    },
    {
      "id": "2273646",
      "postDate": "05/25/2023 10:08:00",
      "content": "<p>Congratulations to you and thanks for sharing </p>",
      "rawMarkdown": "Congratulations to you and thanks for sharing",
      "votes": null
    },
    {
      "id": "2273659",
      "postDate": "05/25/2023 10:18:26",
      "content": "<p>Thanks for the elaboration, I am actually asking about BirdNet performance alone, not the ensemble. Do you retrain the BirdNet or directly use BirdNet and average results?</p>",
      "rawMarkdown": "Thanks for the elaboration, I am actually asking about BirdNet performance alone, not the ensemble. Do you retrain the BirdNet or directly use BirdNet and average results?",
      "votes": null
    },
    {
      "id": "2273667",
      "postDate": "05/25/2023 10:24:11",
      "content": "<p>I submitted some birdnet models but never included them in my ensemble (nor did I concat its embeddings to my imagenet backbone embeddings). Here are my scores with birdnet:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F443651%2Fd0fd9d07194abf17e6f599b1a1727348%2FScreenshot%20from%202023-05-25%2012-23-57.png?generation=1685010249453468&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I submitted some birdnet models but never included them in my ensemble (nor did I concat its embeddings to my imagenet backbone embeddings). Here are my scores with birdnet:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F443651%2Fd0fd9d07194abf17e6f599b1a1727348%2FScreenshot%20from%202023-05-25%2012-23-57.png?generation=1685010249453468&alt=media)",
      "votes": null
    },
    {
      "id": "2273687",
      "postDate": "05/25/2023 10:37:28",
      "content": "<p>Thanks for the birdnet results, I guess I need to check my code.</p>",
      "rawMarkdown": "Thanks for the birdnet results, I guess I need to check my code.",
      "votes": null
    },
    {
      "id": "2273699",
      "postDate": "05/25/2023 10:43:44",
      "content": "<p>BirdNet models on their own do not do good as they were not trained with all our competitions classes, they only work when ensembled on their classes</p>",
      "rawMarkdown": "BirdNet models on their own do not do good as they were not trained with all our competitions classes, they only work when ensembled on their classes",
      "votes": null
    },
    {
      "id": "2273709",
      "postDate": "05/25/2023 10:53:29",
      "content": "<p>Have you checked the <a href=\"https://github.com/kahst/BirdNET-Analyzer/tree/main#training\" target=\"_blank\">BirdNET-Analyzer v2.3</a>? They claimed that \"You can train your own custom classifier on top of BirdNET\". However I guess you can only train the classifier head, not the BirdNET. </p>",
      "rawMarkdown": "Have you checked the [BirdNET-Analyzer v2.3](https://github.com/kahst/BirdNET-Analyzer/tree/main#training)? They claimed that \"You can train your own custom classifier on top of BirdNET\". However I guess you can only train the classifier head, not the BirdNET.",
      "votes": null
    },
    {
      "id": "2276951",
      "postDate": "05/27/2023 10:35:46",
      "content": "<p>Congratulations and thanks for sharing! <br>\nIt is nice to see the birdnet worked!<br>\nWould you mind to share the code of converting eca_nfnet_l0 to openvino?<br>\nI had trouble in converting so I gave up using eca_nfnet_l0, big thanks!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing! \nIt is nice to see the birdnet worked!\nWould you mind to share the code of converting eca_nfnet_l0 to openvino?\nI had trouble in converting so I gave up using eca_nfnet_l0, big thanks!",
      "votes": null
    },
    {
      "id": "2286254",
      "postDate": "06/03/2023 10:33:36",
      "content": "<p>Great work team and thanks for sharing the writeup!</p>",
      "rawMarkdown": "Great work team and thanks for sharing the writeup!",
      "votes": null
    },
    {
      "id": "2286300",
      "postDate": "06/03/2023 11:26:24",
      "content": "<p>Sorry for a very late reply, I somehow missed this comment</p>\n<p><a href=\"https://www.kaggle.com/code/harshitsheoran/openvino-opt/notebook\" target=\"_blank\">https://www.kaggle.com/code/harshitsheoran/openvino-opt/notebook</a></p>\n<p>I think it was that ScaledStdConv2d was using batch norm and the default was going for training=True, which was causing error, so I remade the function and replaced the layers</p>",
      "rawMarkdown": "Sorry for a very late reply, I somehow missed this comment\n\nhttps://www.kaggle.com/code/harshitsheoran/openvino-opt/notebook\n\nI think it was that ScaledStdConv2d was using batch norm and the default was going for training=True, which was causing error, so I remade the function and replaced the layers",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2273615,
      "author_name": "atsunorifujita",
      "author_url": "",
      "post_date": "05/25/2023 09:30:14",
      "content": "<p>Congratulations! Here again, OpenVino. Speeding up inference is a great learning experience for me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2273631,
      "author_name": "aphysict",
      "author_url": "",
      "post_date": "05/25/2023 10:00:39",
      "content": "<p>Congratulations and thanks for sharing! What is the private LB score for single BirdNET ? I am wondering if I use BirdNET correctly.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2273641,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "05/25/2023 10:06:08",
          "content": "<p>A lot of people on Public .82 will get to .83 if they use BirdNet, I think it would be a 0.005 increase on average between both public and private… atleast that is what we observed</p>",
          "votes": null,
          "replies": [
            {
              "id": 2273659,
              "author_name": "aphysict",
              "author_url": "",
              "post_date": "05/25/2023 10:18:26",
              "content": "<p>Thanks for the elaboration, I am actually asking about BirdNet performance alone, not the ensemble. Do you retrain the BirdNet or directly use BirdNet and average results?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2273667,
                  "author_name": "group16",
                  "author_url": "",
                  "post_date": "05/25/2023 10:24:11",
                  "content": "<p>I submitted some birdnet models but never included them in my ensemble (nor did I concat its embeddings to my imagenet backbone embeddings). Here are my scores with birdnet:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F443651%2Fd0fd9d07194abf17e6f599b1a1727348%2FScreenshot%20from%202023-05-25%2012-23-57.png?generation=1685010249453468&amp;alt=media\" alt=\"\"></p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 2273687,
              "author_name": "aphysict",
              "author_url": "",
              "post_date": "05/25/2023 10:37:28",
              "content": "<p>Thanks for the birdnet results, I guess I need to check my code.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2273699,
                  "author_name": "harshitsheoran",
                  "author_url": "",
                  "post_date": "05/25/2023 10:43:44",
                  "content": "<p>BirdNet models on their own do not do good as they were not trained with all our competitions classes, they only work when ensembled on their classes</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 2273709,
              "author_name": "aphysict",
              "author_url": "",
              "post_date": "05/25/2023 10:53:29",
              "content": "<p>Have you checked the <a href=\"https://github.com/kahst/BirdNET-Analyzer/tree/main#training\" target=\"_blank\">BirdNET-Analyzer v2.3</a>? They claimed that \"You can train your own custom classifier on top of BirdNET\". However I guess you can only train the classifier head, not the BirdNET. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2273646,
      "author_name": "scipygaurav",
      "author_url": "",
      "post_date": "05/25/2023 10:08:00",
      "content": "<p>Congratulations to you and thanks for sharing </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2276951,
      "author_name": "honglihang",
      "author_url": "",
      "post_date": "05/27/2023 10:35:46",
      "content": "<p>Congratulations and thanks for sharing! <br>\nIt is nice to see the birdnet worked!<br>\nWould you mind to share the code of converting eca_nfnet_l0 to openvino?<br>\nI had trouble in converting so I gave up using eca_nfnet_l0, big thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2286300,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "06/03/2023 11:26:24",
          "content": "<p>Sorry for a very late reply, I somehow missed this comment</p>\n<p><a href=\"https://www.kaggle.com/code/harshitsheoran/openvino-opt/notebook\" target=\"_blank\">https://www.kaggle.com/code/harshitsheoran/openvino-opt/notebook</a></p>\n<p>I think it was that ScaledStdConv2d was using batch norm and the default was going for training=True, which was causing error, so I remade the function and replaced the layers</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2286254,
      "author_name": "pardeep19singh",
      "author_url": "",
      "post_date": "06/03/2023 10:33:36",
      "content": "<p>Great work team and thanks for sharing the writeup!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2273599": "Thanks to the organizers for hosting BirdClef 2023 again this year and giving us a learning experience.\n\nHeartfelt Congratulations to the winners and everyone who got to learn and experience.\n\nI would also like to express my gratitude to my talented teammates @ivanaerlic and @nischaydnk\n\n# Training\n\nWe used efficientnet_b0, eca_nfnet_l0 and convnext_tiny architectures in our final submission.\n\nThe training was done in multiple rounds using datasets from all previous years. All models were trained on 5s clips\n\nFor eca_nfnet_l0 and convnext_tiny\nHop size 320, n_mels 64, image_size = 64x501\nFor efficientnet_b0\nHop size 320, n_mels 64, image_size (Bilinear image size increase) = 128x1002\n\nThe training was done in multiple rounds:\n\n**Round 1:**\n\nHop size 512, n_mels 64, image_size = 64x313\n(Bilinear image size increase) = 128x636\n\nModels : Eca nfnet l0 (128x636), convnext tiny (128x636).\n\nTraining data: 2022 & 2023\n\nAugmentations: HFlip, Random Cutout, Pink Noise, Gaussian Noise, Noise Injection, Random Volume\n\nType of Training: Competition Labels\n\nCV: ~.79 (Low)\n\n**Round 2:**\n\nHop size 320, n_mels 64, image_size = 64x501\n\nModels : Efficientnet B0 (128x1002 Bilinear increase), Eca nfnet l0 (64x501), convnext tiny (64x501).\n\nAugmentations: HFlip, Random Cutout, Pink Noise, Gaussian Noise, Noise Injection, Random Volume\n\nTraining data: 2022 & 2023\n\nType of Training: Distillation\n\nCV: ~.79 (Slight increase)\n\n**Round 3:**\n\nHop size 320, n_mels 64, image_size = 64x501\n\nModels : Efficientnet B0 (128x1002 Bilinear increase), Eca nfnet l0 (64x501), convnext tiny (64x501).\n\nTraining data: 2021 & 2022 & 2023\n\nExtra methods: AWP, SWA, Attention Head\n\nAugmentations: HFlip, Random Cutout, Pixel Dropout\n\nType of Training: Competition Labels\n\nCV: ~.8\n\n**Round 4:**\n\nHop size 320, n_mels 64, image_size = 64x501\n\nModels : Efficientnet B0 (128x1002 Bilinear increase), Eca nfnet l0 (64x501), convnext tiny (64x501).\n\nTraining data: 2023\n\nAugmentations: HFlip, Random Cutout, Pixel Dropout, Pink Noise, Gaussian Noise, Noise Injection, Random Volume\n\nExtra methods: SWA, Attention Head\n\nType of Training: Pseudo Labels \n\nCV: ~.81 (mild leak) \n\nOur CV was the first 5 seconds of the clips, this CV although was not always consistent with LB, gave us a good enough idea of what we could expect.\n\nPseudo labeling improved our score a lot. We used pseudo-labels to make new secondary labels and gave them a hard probability 0.5\n\nFor convnext model, it was first fine tuned to increase the score on the targets where secondary label exists, although this sacrificed the primary label score, this adds diversity to our ensembles\n\n# Inference\n\nWe used OpenVino to increase the inference speed of our models\n\nWe also saw that loading an audio and giving it in batches, the first batch of inference is especially slower than the rest, so we made it so that it loads 8 audios at once, makes a data loader for all 8 audios and then forward pass the model, doing this we increased the inference speed further by 20%.\n\nSo we were able to submit 7 models, 2 of efficientnet_b0 + 2 of eca_nfnet_l0 + 3 of convnext_tiny\n\nWe also noticed that sigmoid after weight averaging was more effective than sigmoid before weight averaging, also we were using softmax activation which gave a better score than sigmoid up until we reached top of .82 (public LB), after that sigmoid started to work better.\n\n# BirdNet\n\nOur best private submission turned out to be when we ensembled our models with BirdNet, although it was like our top 10 submissions of public and did give us a lot of variance, we did not select this because we were a little afraid of it being a dice roll on private LB.\n\n# Post Processing\n\nLooking at the last year’s competition, we tried out increasing the confidence of a class by a factor of 1.5 if any prediction of that class in the audio file is more than .96\n\nThis did work on a lot of submissions and was in fact amongst our top subs, but in the end, it did not play well on private LB and we did not select it because of the risk of overfitting the public LB.\n\n# What did not work\n\nSED models\nQ-Transform\nMore than 5s training\nRolling Mean with last and next 5s segment\n\nMore than 60 submissions of ours had timeout errors and on the last second day of the competition, 4 of 5 submissions had encountered timeout error while before that the chance of getting a timeout was only 20-30% in the same notebook.",
    "2273615": "Congratulations! Here again, OpenVino. Speeding up inference is a great learning experience for me.",
    "2273631": "Congratulations and thanks for sharing! What is the private LB score for single BirdNET ? I am wondering if I use BirdNET correctly.",
    "2273641": "A lot of people on Public .82 will get to .83 if they use BirdNet, I think it would be a 0.005 increase on average between both public and private... atleast that is what we observed",
    "2273646": "Congratulations to you and thanks for sharing",
    "2273659": "Thanks for the elaboration, I am actually asking about BirdNet performance alone, not the ensemble. Do you retrain the BirdNet or directly use BirdNet and average results?",
    "2273667": "I submitted some birdnet models but never included them in my ensemble (nor did I concat its embeddings to my imagenet backbone embeddings). Here are my scores with birdnet:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F443651%2Fd0fd9d07194abf17e6f599b1a1727348%2FScreenshot%20from%202023-05-25%2012-23-57.png?generation=1685010249453468&alt=media)",
    "2273687": "Thanks for the birdnet results, I guess I need to check my code.",
    "2273699": "BirdNet models on their own do not do good as they were not trained with all our competitions classes, they only work when ensembled on their classes",
    "2273709": "Have you checked the [BirdNET-Analyzer v2.3](https://github.com/kahst/BirdNET-Analyzer/tree/main#training)? They claimed that \"You can train your own custom classifier on top of BirdNET\". However I guess you can only train the classifier head, not the BirdNET.",
    "2276951": "Congratulations and thanks for sharing! \nIt is nice to see the birdnet worked!\nWould you mind to share the code of converting eca_nfnet_l0 to openvino?\nI had trouble in converting so I gave up using eca_nfnet_l0, big thanks!",
    "2286254": "Great work team and thanks for sharing the writeup!",
    "2286300": "Sorry for a very late reply, I somehow missed this comment\n\nhttps://www.kaggle.com/code/harshitsheoran/openvino-opt/notebook\n\nI think it was that ScaledStdConv2d was using batch norm and the default was going for training=True, which was causing error, so I remade the function and replaced the layers"
  },
  "source": "meta"
}