{
  "id": 183212,
  "title": "101st Place Solution (Top 8%)",
  "url": "/competitions/birdsong-recognition/writeups/101st-place-solution-top-8",
  "author_name": "",
  "post_date": "2020-09-17T15:53:52.213Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>This competition was a toughie. Special thanks to my team members for their support <a href=\"https://www.kaggle.com/doanquanvietnamca\" target=\"_blank\">@doanquanvietnamca</a> , <a href=\"https://www.kaggle.com/truonghoang\" target=\"_blank\">@truonghoang</a> throughout this competition.</p>\n<p>As this was my first audio competition, in the beginning, I tried effnetb3 using fastai as a baseline, with precomputed spectrograms. This was able to beat the public LB baseline, giving a score of 0.554. However, I was unable to improve this model. My team and I subsequently tried several other new models in the process</p>\n<p><strong>Models tried:</strong></p>\n<ol>\n<li>EfficientNetB1</li>\n<li>ResNest50 (best model)</li>\n<li>MobileNetV2</li>\n<li>SED model, with effnetb0 backbone</li>\n<li>EfficientNetB3 (with pretrained noisystudent weights)</li>\n</ol>\n<p>ResNest50 performed well for this task, with the same classifier head provided by <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a>. </p>\n<p>Augmentations also worked well for us, namely:</p>\n<ol>\n<li>Adding gaussian noise</li>\n<li>Adding pink noise</li>\n<li>Adding background noise </li>\n</ol>\n<p>We also tried a model with 265 classes, with nocall as a class but that didn't work better than a normal 264 classifier model.</p>\n<p>We tried incorporating secondary labels to perform the multilabel classification task, it gave decent CV (up to 0.7x validation micro score), faltered in public LB (0.569) but was the best performing model for private LB (0.616). </p>\n<p>We were unable to bridge the gap between CV and LB, the correlation between CV and LB was not strong so in the end we had to trust our methodology and public LB.</p>\n<p>We settled for an ensemble of models between ResNest50 and MobileNetV2 which landed us with the bronze solution.</p>\n<p><strong>Summary of things that did not work for us:</strong></p>\n<ol>\n<li>SED ( could not get it to work somehow, will have to learn from training solution by <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> to understand more why)</li>\n<li>265 class classifier (including nocall class), decreased public and private LB</li>\n<li>Denoising on test set</li>\n<li>F1Loss (possibly due to batch size being too small)</li>\n<li>FocalLoss</li>\n</ol>\n<p>Special thanks to the organizers of this competition and fellow competitors whom I learn loads from (@hidehisaarai1213 , <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a>, <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>). Hope to join another audio competition to further my learning. Congrats to all winners in this competition.</p>",
  "messages": [
    {
      "id": "1012210",
      "postDate": "09/16/2020 01:06:02",
      "content": "<p>This competition was a toughie. Special thanks to my team members for their support <a href=\"https://www.kaggle.com/doanquanvietnamca\" target=\"_blank\">@doanquanvietnamca</a> , <a href=\"https://www.kaggle.com/truonghoang\" target=\"_blank\">@truonghoang</a> throughout this competition.</p>\n<p>As this was my first audio competition, in the beginning, I tried effnetb3 using fastai as a baseline, with precomputed spectrograms. This was able to beat the public LB baseline, giving a score of 0.554. However, I was unable to improve this model. My team and I subsequently tried several other new models in the process</p>\n<p><strong>Models tried:</strong></p>\n<ol>\n<li>EfficientNetB1</li>\n<li>ResNest50 (best model)</li>\n<li>MobileNetV2</li>\n<li>SED model, with effnetb0 backbone</li>\n<li>EfficientNetB3 (with pretrained noisystudent weights)</li>\n</ol>\n<p>ResNest50 performed well for this task, with the same classifier head provided by <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a>. </p>\n<p>Augmentations also worked well for us, namely:</p>\n<ol>\n<li>Adding gaussian noise</li>\n<li>Adding pink noise</li>\n<li>Adding background noise </li>\n</ol>\n<p>We also tried a model with 265 classes, with nocall as a class but that didn't work better than a normal 264 classifier model.</p>\n<p>We tried incorporating secondary labels to perform the multilabel classification task, it gave decent CV (up to 0.7x validation micro score), faltered in public LB (0.569) but was the best performing model for private LB (0.616). </p>\n<p>We were unable to bridge the gap between CV and LB, the correlation between CV and LB was not strong so in the end we had to trust our methodology and public LB.</p>\n<p>We settled for an ensemble of models between ResNest50 and MobileNetV2 which landed us with the bronze solution.</p>\n<p><strong>Summary of things that did not work for us:</strong></p>\n<ol>\n<li>SED ( could not get it to work somehow, will have to learn from training solution by <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> to understand more why)</li>\n<li>265 class classifier (including nocall class), decreased public and private LB</li>\n<li>Denoising on test set</li>\n<li>F1Loss (possibly due to batch size being too small)</li>\n<li>FocalLoss</li>\n</ol>\n<p>Special thanks to the organizers of this competition and fellow competitors whom I learn loads from (@hidehisaarai1213 , <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a>, <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>). Hope to join another audio competition to further my learning. Congrats to all winners in this competition.</p>",
      "rawMarkdown": "This competition was a toughie. Special thanks to my team members for their support @doanquanvietnamca , @truonghoang throughout this competition.\n\nAs this was my first audio competition, in the beginning, I tried effnetb3 using fastai as a baseline, with precomputed spectrograms. This was able to beat the public LB baseline, giving a score of 0.554. However, I was unable to improve this model. My team and I subsequently tried several other new models in the process\n\n**Models tried:**\n1. EfficientNetB1\n2. ResNest50 (best model)\n3. MobileNetV2\n4. SED model, with effnetb0 backbone\n5. EfficientNetB3 (with pretrained noisystudent weights)\n\nResNest50 performed well for this task, with the same classifier head provided by @hidehisaarai1213. \n\nAugmentations also worked well for us, namely:\n1. Adding gaussian noise\n2. Adding pink noise\n3. Adding background noise \n\nWe also tried a model with 265 classes, with nocall as a class but that didn't work better than a normal 264 classifier model.\n\nWe tried incorporating secondary labels to perform the multilabel classification task, it gave decent CV (up to 0.7x validation micro score), faltered in public LB (0.569) but was the best performing model for private LB (0.616). \n\nWe were unable to bridge the gap between CV and LB, the correlation between CV and LB was not strong so in the end we had to trust our methodology and public LB.\n\nWe settled for an ensemble of models between ResNest50 and MobileNetV2 which landed us with the bronze solution.\n\n**Summary of things that did not work for us:**\n1. SED ( could not get it to work somehow, will have to learn from training solution by @hidehisaarai1213 to understand more why)\n2. 265 class classifier (including nocall class), decreased public and private LB\n3. Denoising on test set\n4. F1Loss (possibly due to batch size being too small)\n5. FocalLoss\n\nSpecial thanks to the organizers of this competition and fellow competitors whom I learn loads from (@hidehisaarai1213 , @radek1, @kneroma, @ttahara). Hope to join another audio competition to further my learning. Congrats to all winners in this competition.",
      "votes": null
    },
    {
      "id": "1012940",
      "postDate": "09/16/2020 12:06:02",
      "content": "<p><a href=\"https://www.kaggle.com/alanchn31\" target=\"_blank\">@alanchn31</a>  Congrats to become competition Expert :)</p>",
      "rawMarkdown": "alanchn31  Congrats to become competition Expert :)",
      "votes": null
    },
    {
      "id": "1012947",
      "postDate": "09/16/2020 12:10:57",
      "content": "<p>Thank you! Congrats to you too on your great placing.</p>",
      "rawMarkdown": "Thank you! Congrats to you too on your great placing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1012940,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "09/16/2020 12:06:02",
      "content": "<p><a href=\"https://www.kaggle.com/alanchn31\" target=\"_blank\">@alanchn31</a>  Congrats to become competition Expert :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1012947,
          "author_name": "alanchn31",
          "author_url": "",
          "post_date": "09/16/2020 12:10:57",
          "content": "<p>Thank you! Congrats to you too on your great placing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1012210": "This competition was a toughie. Special thanks to my team members for their support @doanquanvietnamca , @truonghoang throughout this competition.\n\nAs this was my first audio competition, in the beginning, I tried effnetb3 using fastai as a baseline, with precomputed spectrograms. This was able to beat the public LB baseline, giving a score of 0.554. However, I was unable to improve this model. My team and I subsequently tried several other new models in the process\n\n**Models tried:**\n1. EfficientNetB1\n2. ResNest50 (best model)\n3. MobileNetV2\n4. SED model, with effnetb0 backbone\n5. EfficientNetB3 (with pretrained noisystudent weights)\n\nResNest50 performed well for this task, with the same classifier head provided by @hidehisaarai1213. \n\nAugmentations also worked well for us, namely:\n1. Adding gaussian noise\n2. Adding pink noise\n3. Adding background noise \n\nWe also tried a model with 265 classes, with nocall as a class but that didn't work better than a normal 264 classifier model.\n\nWe tried incorporating secondary labels to perform the multilabel classification task, it gave decent CV (up to 0.7x validation micro score), faltered in public LB (0.569) but was the best performing model for private LB (0.616). \n\nWe were unable to bridge the gap between CV and LB, the correlation between CV and LB was not strong so in the end we had to trust our methodology and public LB.\n\nWe settled for an ensemble of models between ResNest50 and MobileNetV2 which landed us with the bronze solution.\n\n**Summary of things that did not work for us:**\n1. SED ( could not get it to work somehow, will have to learn from training solution by @hidehisaarai1213 to understand more why)\n2. 265 class classifier (including nocall class), decreased public and private LB\n3. Denoising on test set\n4. F1Loss (possibly due to batch size being too small)\n5. FocalLoss\n\nSpecial thanks to the organizers of this competition and fellow competitors whom I learn loads from (@hidehisaarai1213 , @radek1, @kneroma, @ttahara). Hope to join another audio competition to further my learning. Congrats to all winners in this competition.",
    "1012940": "alanchn31  Congrats to become competition Expert :)",
    "1012947": "Thank you! Congrats to you too on your great placing."
  },
  "source": "meta"
}