{
  "id": 183196,
  "title": "167 Place Solution",
  "url": "/competitions/birdsong-recognition/writeups/takamichi-toda-167-place-solution",
  "author_name": "",
  "post_date": "2020-09-16T00:04:06.508385200Z",
  "votes": 13,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Congrats to all winners! And thanks to organizers!<br>\nMy score is not good, but sharing my gotten knowledge.</p>\n<h3>modeling</h3>\n<p>My base model is DenseNet201 trained ImageNet.<br>\nMost people used ResNet, but it not work for me.</p>\n<p>I read <a href=\"https://arxiv.org/pdf/1908.02876.pdf\" target=\"_blank\">this paper</a>, and apply applied to my model.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1483555%2F4d0acfb9be5bb112c566bc8d55627aa2%2F2020-09-15%2018.45.50.png?generation=1600214606636136&amp;alt=media\" alt=\"\"></p>\n<p>code is:</p>\n<pre><code>class BirdcallNet(nn.Module):\n    def __init__(self):\n        super(BirdcallNet_densenet, self).__init__()\n        densenet = densenet161(pretrained=config.PRETRAINED)\n        self.features = densenet.features\n        self.l8_a = nn.Conv1d(2208, config.N_LABEL, 1, bias=False)\n        self.l8_b = nn.Conv1d(2208, config.N_LABEL, 1, bias=False)\n\n    def forward(self, x, perm=None, gamma=None):\n        # input: (batch, channel, Hz, time)\n        frames_num = x.shape[3]\n        x = x.transpose(3, 2)  # (batch, channel, time, Hz)\n        h = self.features(x)  # (batch, unit, time, Hz)\n\n        h = F.relu(h, inplace=True)\n        h  = torch.mean(h, dim=3)  # (batch, unit, time)\n\n        xa = self.l8_a(h)  # (batch, n_class, time)\n        xb = self.l8_b(h)  # (batch, n_class, time)\n        xb = torch.softmax(xb, dim=2)\n\n        pseudo_label = (xa.sigmoid() &gt;= 0.5).float()\n        clipwise_preds = torch.sum(xa * xb, dim=2)\n        attention_preds = xb\n\n        return clipwise_preds, attention_preds, pseudo_label\n</code></pre>\n<h4>Parameters</h4>\n<ul>\n<li>BCE loss</li>\n<li>5-fold CV</li>\n<li>Adam<ul>\n<li>learning_rate=1e-3</li>\n<li>55 epoch</li>\n<li>batch size=64</li></ul></li>\n<li>CosineAnnealingWarmRestarts<ul>\n<li>T=10</li></ul></li>\n</ul>\n<h3>Data Augumentation</h3>\n<ul>\n<li>Adjust Gamma</li>\n<li>Spec Augmentation Freq</li>\n<li>MixUp</li>\n</ul>\n<h3>Other Technique</h3>\n<ul>\n<li>SWA</li>\n</ul>\n<h3>Not Work for me</h3>\n<ul>\n<li>Denoise(I tried my best, but it didn't work …)</li>\n<li>\"nocall\" prediction by PANNs trained model</li>\n<li>Transformer</li>\n<li>WaveNet</li>\n<li>use secondary_labels</li>\n<li>optimize threshold</li>\n<li>Focal Loss</li>\n<li>CutMix</li>\n<li>Label Smoothing</li>\n<li>Spec Augmentation Time</li>\n<li>DIfferent learning rate for each layer</li>\n<li>Multi Sample Dropout</li>\n</ul>\n<h3>My Code</h3>\n<p><a href=\"https://github.com/trtd56/Birdcall\" target=\"_blank\">https://github.com/trtd56/Birdcall</a></p>\n<h3>My Blog(Japanese)</h3>\n<p><a href=\"https://www.ai-shift.co.jp/techblog/1271\" target=\"_blank\">https://www.ai-shift.co.jp/techblog/1271</a></p>\n<p>Thank you!!</p>",
  "messages": [
    {
      "id": "1012127",
      "postDate": "09/16/2020 00:04:06",
      "content": "<p>Congrats to all winners! And thanks to organizers!<br>\nMy score is not good, but sharing my gotten knowledge.</p>\n<h3>modeling</h3>\n<p>My base model is DenseNet201 trained ImageNet.<br>\nMost people used ResNet, but it not work for me.</p>\n<p>I read <a href=\"https://arxiv.org/pdf/1908.02876.pdf\" target=\"_blank\">this paper</a>, and apply applied to my model.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1483555%2F4d0acfb9be5bb112c566bc8d55627aa2%2F2020-09-15%2018.45.50.png?generation=1600214606636136&amp;alt=media\" alt=\"\"></p>\n<p>code is:</p>\n<pre><code>class BirdcallNet(nn.Module):\n    def __init__(self):\n        super(BirdcallNet_densenet, self).__init__()\n        densenet = densenet161(pretrained=config.PRETRAINED)\n        self.features = densenet.features\n        self.l8_a = nn.Conv1d(2208, config.N_LABEL, 1, bias=False)\n        self.l8_b = nn.Conv1d(2208, config.N_LABEL, 1, bias=False)\n\n    def forward(self, x, perm=None, gamma=None):\n        # input: (batch, channel, Hz, time)\n        frames_num = x.shape[3]\n        x = x.transpose(3, 2)  # (batch, channel, time, Hz)\n        h = self.features(x)  # (batch, unit, time, Hz)\n\n        h = F.relu(h, inplace=True)\n        h  = torch.mean(h, dim=3)  # (batch, unit, time)\n\n        xa = self.l8_a(h)  # (batch, n_class, time)\n        xb = self.l8_b(h)  # (batch, n_class, time)\n        xb = torch.softmax(xb, dim=2)\n\n        pseudo_label = (xa.sigmoid() &gt;= 0.5).float()\n        clipwise_preds = torch.sum(xa * xb, dim=2)\n        attention_preds = xb\n\n        return clipwise_preds, attention_preds, pseudo_label\n</code></pre>\n<h4>Parameters</h4>\n<ul>\n<li>BCE loss</li>\n<li>5-fold CV</li>\n<li>Adam<ul>\n<li>learning_rate=1e-3</li>\n<li>55 epoch</li>\n<li>batch size=64</li></ul></li>\n<li>CosineAnnealingWarmRestarts<ul>\n<li>T=10</li></ul></li>\n</ul>\n<h3>Data Augumentation</h3>\n<ul>\n<li>Adjust Gamma</li>\n<li>Spec Augmentation Freq</li>\n<li>MixUp</li>\n</ul>\n<h3>Other Technique</h3>\n<ul>\n<li>SWA</li>\n</ul>\n<h3>Not Work for me</h3>\n<ul>\n<li>Denoise(I tried my best, but it didn't work …)</li>\n<li>\"nocall\" prediction by PANNs trained model</li>\n<li>Transformer</li>\n<li>WaveNet</li>\n<li>use secondary_labels</li>\n<li>optimize threshold</li>\n<li>Focal Loss</li>\n<li>CutMix</li>\n<li>Label Smoothing</li>\n<li>Spec Augmentation Time</li>\n<li>DIfferent learning rate for each layer</li>\n<li>Multi Sample Dropout</li>\n</ul>\n<h3>My Code</h3>\n<p><a href=\"https://github.com/trtd56/Birdcall\" target=\"_blank\">https://github.com/trtd56/Birdcall</a></p>\n<h3>My Blog(Japanese)</h3>\n<p><a href=\"https://www.ai-shift.co.jp/techblog/1271\" target=\"_blank\">https://www.ai-shift.co.jp/techblog/1271</a></p>\n<p>Thank you!!</p>",
      "rawMarkdown": "Congrats to all winners! And thanks to organizers!\nMy score is not good, but sharing my gotten knowledge.\n\n### modeling\nMy base model is DenseNet201 trained ImageNet.\nMost people used ResNet, but it not work for me.\n\nI read [this paper](https://arxiv.org/pdf/1908.02876.pdf), and apply applied to my model.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1483555%2F4d0acfb9be5bb112c566bc8d55627aa2%2F2020-09-15%2018.45.50.png?generation=1600214606636136&alt=media)\n\ncode is:\n```python\nclass BirdcallNet(nn.Module):\n    def __init__(self):\n        super(BirdcallNet_densenet, self).__init__()\n        densenet = densenet161(pretrained=config.PRETRAINED)\n        self.features = densenet.features\n        self.l8_a = nn.Conv1d(2208, config.N_LABEL, 1, bias=False)\n        self.l8_b = nn.Conv1d(2208, config.N_LABEL, 1, bias=False)\n\n    def forward(self, x, perm=None, gamma=None):\n        # input: (batch, channel, Hz, time)\n        frames_num = x.shape[3]\n        x = x.transpose(3, 2)  # (batch, channel, time, Hz)\n        h = self.features(x)  # (batch, unit, time, Hz)\n\n        h = F.relu(h, inplace=True)\n        h  = torch.mean(h, dim=3)  # (batch, unit, time)\n\n        xa = self.l8_a(h)  # (batch, n_class, time)\n        xb = self.l8_b(h)  # (batch, n_class, time)\n        xb = torch.softmax(xb, dim=2)\n\n        pseudo_label = (xa.sigmoid() >= 0.5).float()\n        clipwise_preds = torch.sum(xa * xb, dim=2)\n        attention_preds = xb\n\n        return clipwise_preds, attention_preds, pseudo_label\n```\n\n#### Parameters\n- BCE loss\n- 5-fold CV\n- Adam\n - learning_rate=1e-3\n - 55 epoch\n - batch size=64\n- CosineAnnealingWarmRestarts\n - T=10\n\n### Data Augumentation\n- Adjust Gamma\n- Spec Augmentation Freq\n- MixUp\n\n### Other Technique\n- SWA\n\n### Not Work for me\n- Denoise(I tried my best, but it didn't work ...)\n- \"nocall\" prediction by PANNs trained model\n- Transformer\n- WaveNet\n- use secondary_labels\n- optimize threshold\n- Focal Loss\n- CutMix\n- Label Smoothing\n- Spec Augmentation Time\n- DIfferent learning rate for each layer\n- Multi Sample Dropout\n\n\n### My Code\nhttps://github.com/trtd56/Birdcall\n\n### My Blog(Japanese)\nhttps://www.ai-shift.co.jp/techblog/1271\n\nThank you!!",
      "votes": null
    },
    {
      "id": "1012131",
      "postDate": "09/16/2020 00:08:54",
      "content": "<p>Thank you for participating in the challenge and for posting your code!</p>",
      "rawMarkdown": "Thank you for participating in the challenge and for posting your code!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1012131,
      "author_name": "holgerklinck",
      "author_url": "",
      "post_date": "09/16/2020 00:08:54",
      "content": "<p>Thank you for participating in the challenge and for posting your code!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1012127": "Congrats to all winners! And thanks to organizers!\nMy score is not good, but sharing my gotten knowledge.\n\n### modeling\nMy base model is DenseNet201 trained ImageNet.\nMost people used ResNet, but it not work for me.\n\nI read [this paper](https://arxiv.org/pdf/1908.02876.pdf), and apply applied to my model.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1483555%2F4d0acfb9be5bb112c566bc8d55627aa2%2F2020-09-15%2018.45.50.png?generation=1600214606636136&alt=media)\n\ncode is:\n```python\nclass BirdcallNet(nn.Module):\n    def __init__(self):\n        super(BirdcallNet_densenet, self).__init__()\n        densenet = densenet161(pretrained=config.PRETRAINED)\n        self.features = densenet.features\n        self.l8_a = nn.Conv1d(2208, config.N_LABEL, 1, bias=False)\n        self.l8_b = nn.Conv1d(2208, config.N_LABEL, 1, bias=False)\n\n    def forward(self, x, perm=None, gamma=None):\n        # input: (batch, channel, Hz, time)\n        frames_num = x.shape[3]\n        x = x.transpose(3, 2)  # (batch, channel, time, Hz)\n        h = self.features(x)  # (batch, unit, time, Hz)\n\n        h = F.relu(h, inplace=True)\n        h  = torch.mean(h, dim=3)  # (batch, unit, time)\n\n        xa = self.l8_a(h)  # (batch, n_class, time)\n        xb = self.l8_b(h)  # (batch, n_class, time)\n        xb = torch.softmax(xb, dim=2)\n\n        pseudo_label = (xa.sigmoid() >= 0.5).float()\n        clipwise_preds = torch.sum(xa * xb, dim=2)\n        attention_preds = xb\n\n        return clipwise_preds, attention_preds, pseudo_label\n```\n\n#### Parameters\n- BCE loss\n- 5-fold CV\n- Adam\n - learning_rate=1e-3\n - 55 epoch\n - batch size=64\n- CosineAnnealingWarmRestarts\n - T=10\n\n### Data Augumentation\n- Adjust Gamma\n- Spec Augmentation Freq\n- MixUp\n\n### Other Technique\n- SWA\n\n### Not Work for me\n- Denoise(I tried my best, but it didn't work ...)\n- \"nocall\" prediction by PANNs trained model\n- Transformer\n- WaveNet\n- use secondary_labels\n- optimize threshold\n- Focal Loss\n- CutMix\n- Label Smoothing\n- Spec Augmentation Time\n- DIfferent learning rate for each layer\n- Multi Sample Dropout\n\n\n### My Code\nhttps://github.com/trtd56/Birdcall\n\n### My Blog(Japanese)\nhttps://www.ai-shift.co.jp/techblog/1271\n\nThank you!!",
    "1012131": "Thank you for participating in the challenge and for posting your code!"
  },
  "source": "meta"
}