{
  "id": 327393,
  "title": "15th place summary",
  "url": "/competitions/birdclef-2022/discussion/327393",
  "author_name": "furu-nag",
  "post_date": "2022-05-27T03:30:53.699000",
  "votes": 16,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thanks to team member and all participants.<br>\nI enjoyed the speech recognition competition very much,<br>\nbut I couldn’t grasp the behavior of the metric and LB until the end.<br>\nAlso, crinoid ’s EDA was very helpful in formulating a strategy.<br>\nThe detailed solution is shown below:</p>\n<p><strong>Melspectrogram preprocessing parameter</strong></p>\n<pre><code>sr = 32000\nn_mels = 128\nfmin = 20\nfmax = 16000\nn_fft = 32000//10\nhop_len = 3200//4\npower = 2\ntop_db = 80\n</code></pre>\n<p><strong>Model Architecture</strong><br>\nwe was using very simple CNN architecture.</p>\n<pre><code>class Model(nn.Module):\n    def __init__(self,name,pretrained=False):\n        super(Model, self).__init__()\n        self.model = timm.create_model(name,pretrained=pretrained, in_chans=1)\n        self.model.reset_classifier(num_classes=0) \n        in_features = self.model.num_features\n        self.fc = nn.Linear(in_features, cfg.CLASS_NUM)\n\n    def forward(self, x):\n        x = self.model(x)\n        x = self.fc(x)\n        return x\n</code></pre>\n<p>As a backbones we have used:<br>\n・eca_nfnet_l0<br>\n・convnext_tiny<br>\n・resnest50d<br>\nModel Training<br>\nwe have trained models on 5 sec and 7sec chunks and with secondary labels.<br>\nAt First, we made pretrain models by 2020, 2021, 2022 without scored birds competition data.<br>\nNext, we have finetunned pretrain models the scored birds datas.</p>\n<p><strong>Ensemble and TTA</strong><br>\nWe ensemble the single trained model (head 30s train tail5s validation 7s segment) and 5foldmodel (5s segment) and apply TTA:</p>\n<pre><code>if i &lt;= 2:\n                pred = model(images).sigmoid().detach().cpu().numpy()\n            else:\n                pred1 = model(images[:,:,:,0:201]).sigmoid().detach().cpu().numpy()\n                pred2 = model(images[:,:,:,40:241]).sigmoid().detach().cpu().numpy()\n                pred3 = model(images[:,:,:,80:281]).sigmoid().detach().cpu().numpy()\n                pred = 0.25*pred1 + 0.5*pred2 + 0.25*pred3\n</code></pre>\n<p><strong>Post Processing</strong></p>\n<p>Model weighted<br>\nModels with high LB are weighted larger and models with lower LB are weighted smaller.</p>\n<p>and then we apply voting method:<br>\nAfter weight averaging the model’s predictions, we conducted pred re-weighting by voting score.  Voting score was calculated by judging ,for each species, whether each model’s pred exceeds the given threshold or not. By doing this, we aimed to grab more rare bird’s TP.</p>\n<p>how to decide threshold<br>\nwe apply percentaile approach(2021 2nd).</p>",
  "messages": [
    {
      "id": 1802688,
      "postDate": "2022-05-27T03:30:53.700Z",
      "content": "<p>Thanks to team member and all participants.<br>\nI enjoyed the speech recognition competition very much,<br>\nbut I couldn’t grasp the behavior of the metric and LB until the end.<br>\nAlso, crinoid ’s EDA was very helpful in formulating a strategy.<br>\nThe detailed solution is shown below:</p>\n<p><strong>Melspectrogram preprocessing parameter</strong></p>\n<pre><code>sr = 32000\nn_mels = 128\nfmin = 20\nfmax = 16000\nn_fft = 32000//10\nhop_len = 3200//4\npower = 2\ntop_db = 80\n</code></pre>\n<p><strong>Model Architecture</strong><br>\nwe was using very simple CNN architecture.</p>\n<pre><code>class Model(nn.Module):\n    def __init__(self,name,pretrained=False):\n        super(Model, self).__init__()\n        self.model = timm.create_model(name,pretrained=pretrained, in_chans=1)\n        self.model.reset_classifier(num_classes=0) \n        in_features = self.model.num_features\n        self.fc = nn.Linear(in_features, cfg.CLASS_NUM)\n\n    def forward(self, x):\n        x = self.model(x)\n        x = self.fc(x)\n        return x\n</code></pre>\n<p>As a backbones we have used:<br>\n・eca_nfnet_l0<br>\n・convnext_tiny<br>\n・resnest50d<br>\nModel Training<br>\nwe have trained models on 5 sec and 7sec chunks and with secondary labels.<br>\nAt First, we made pretrain models by 2020, 2021, 2022 without scored birds competition data.<br>\nNext, we have finetunned pretrain models the scored birds datas.</p>\n<p><strong>Ensemble and TTA</strong><br>\nWe ensemble the single trained model (head 30s train tail5s validation 7s segment) and 5foldmodel (5s segment) and apply TTA:</p>\n<pre><code>if i &lt;= 2:\n                pred = model(images).sigmoid().detach().cpu().numpy()\n            else:\n                pred1 = model(images[:,:,:,0:201]).sigmoid().detach().cpu().numpy()\n                pred2 = model(images[:,:,:,40:241]).sigmoid().detach().cpu().numpy()\n                pred3 = model(images[:,:,:,80:281]).sigmoid().detach().cpu().numpy()\n                pred = 0.25*pred1 + 0.5*pred2 + 0.25*pred3\n</code></pre>\n<p><strong>Post Processing</strong></p>\n<p>Model weighted<br>\nModels with high LB are weighted larger and models with lower LB are weighted smaller.</p>\n<p>and then we apply voting method:<br>\nAfter weight averaging the model’s predictions, we conducted pred re-weighting by voting score.  Voting score was calculated by judging ,for each species, whether each model’s pred exceeds the given threshold or not. By doing this, we aimed to grab more rare bird’s TP.</p>\n<p>how to decide threshold<br>\nwe apply percentaile approach(2021 2nd).</p>",
      "rawMarkdown": "Thanks to team member and all participants.\nI enjoyed the speech recognition competition very much,\nbut I couldn’t grasp the behavior of the metric and LB until the end.\nAlso, crinoid ’s EDA was very helpful in formulating a strategy.\nThe detailed solution is shown below:\n\n**Melspectrogram preprocessing parameter**\n```\nsr = 32000\nn_mels = 128\nfmin = 20\nfmax = 16000\nn_fft = 32000//10\nhop_len = 3200//4\npower = 2\ntop_db = 80\n```\n**Model Architecture**\nwe was using very simple CNN architecture.\n```\nclass Model(nn.Module):\n    def __init__(self,name,pretrained=False):\n        super(Model, self).__init__()\n        self.model = timm.create_model(name,pretrained=pretrained, in_chans=1)\n        self.model.reset_classifier(num_classes=0) \n        in_features = self.model.num_features\n        self.fc = nn.Linear(in_features, cfg.CLASS_NUM)\n\n    def forward(self, x):\n        x = self.model(x)\n        x = self.fc(x)\n        return x\n```\n\nAs a backbones we have used:\n・eca_nfnet_l0\n・convnext_tiny\n・resnest50d\nModel Training\nwe have trained models on 5 sec and 7sec chunks and with secondary labels.\nAt First, we made pretrain models by 2020, 2021, 2022 without scored birds competition data.\nNext, we have finetunned pretrain models the scored birds datas.\n\n**Ensemble and TTA**\nWe ensemble the single trained model (head 30s train tail5s validation 7s segment) and 5foldmodel (5s segment) and apply TTA:\n```\nif i <= 2:\n                pred = model(images).sigmoid().detach().cpu().numpy()\n            else:\n                pred1 = model(images[:,:,:,0:201]).sigmoid().detach().cpu().numpy()\n                pred2 = model(images[:,:,:,40:241]).sigmoid().detach().cpu().numpy()\n                pred3 = model(images[:,:,:,80:281]).sigmoid().detach().cpu().numpy()\n                pred = 0.25*pred1 + 0.5*pred2 + 0.25*pred3\n```\n\n**Post Processing**\n\nModel weighted\nModels with high LB are weighted larger and models with lower LB are weighted smaller.\n\nand then we apply voting method:\nAfter weight averaging the model’s predictions, we conducted pred re-weighting by voting score.  Voting score was calculated by judging ,for each species, whether each model’s pred exceeds the given threshold or not. By doing this, we aimed to grab more rare bird’s TP.\n\nhow to decide threshold\nwe apply percentaile approach(2021 2nd).",
      "votes": 16
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1802688": "Thanks to team member and all participants.\nI enjoyed the speech recognition competition very much,\nbut I couldn’t grasp the behavior of the metric and LB until the end.\nAlso, crinoid ’s EDA was very helpful in formulating a strategy.\nThe detailed solution is shown below:\n\n**Melspectrogram preprocessing parameter**\n```\nsr = 32000\nn_mels = 128\nfmin = 20\nfmax = 16000\nn_fft = 32000//10\nhop_len = 3200//4\npower = 2\ntop_db = 80\n```\n**Model Architecture**\nwe was using very simple CNN architecture.\n```\nclass Model(nn.Module):\n    def __init__(self,name,pretrained=False):\n        super(Model, self).__init__()\n        self.model = timm.create_model(name,pretrained=pretrained, in_chans=1)\n        self.model.reset_classifier(num_classes=0) \n        in_features = self.model.num_features\n        self.fc = nn.Linear(in_features, cfg.CLASS_NUM)\n\n    def forward(self, x):\n        x = self.model(x)\n        x = self.fc(x)\n        return x\n```\n\nAs a backbones we have used:\n・eca_nfnet_l0\n・convnext_tiny\n・resnest50d\nModel Training\nwe have trained models on 5 sec and 7sec chunks and with secondary labels.\nAt First, we made pretrain models by 2020, 2021, 2022 without scored birds competition data.\nNext, we have finetunned pretrain models the scored birds datas.\n\n**Ensemble and TTA**\nWe ensemble the single trained model (head 30s train tail5s validation 7s segment) and 5foldmodel (5s segment) and apply TTA:\n```\nif i <= 2:\n                pred = model(images).sigmoid().detach().cpu().numpy()\n            else:\n                pred1 = model(images[:,:,:,0:201]).sigmoid().detach().cpu().numpy()\n                pred2 = model(images[:,:,:,40:241]).sigmoid().detach().cpu().numpy()\n                pred3 = model(images[:,:,:,80:281]).sigmoid().detach().cpu().numpy()\n                pred = 0.25*pred1 + 0.5*pred2 + 0.25*pred3\n```\n\n**Post Processing**\n\nModel weighted\nModels with high LB are weighted larger and models with lower LB are weighted smaller.\n\nand then we apply voting method:\nAfter weight averaging the model’s predictions, we conducted pred re-weighting by voting score.  Voting score was calculated by judging ,for each species, whether each model’s pred exceeds the given threshold or not. By doing this, we aimed to grab more rare bird’s TP.\n\nhow to decide threshold\nwe apply percentaile approach(2021 2nd)."
  }
}