{
  "id": 511851,
  "title": "18th place",
  "url": "/competitions/birdclef-2024/writeups/moyashii-18th-place",
  "author_name": "",
  "post_date": "2024-06-12T11:21:07.690Z",
  "votes": 15,
  "comment_count": 3,
  "views": 0,
  "content": "<h1>Summary</h1>\n<p>I used various training setups (spectrogram settings, model types, data augmentation) to perform logit distillation and assembled an ensemble of four distillation models with different configurations.</p>\n<h1>Distillation</h1>\n<p>Using the $L_{MLD}$ from <a href=\"https://arxiv.org/abs/2308.06453\" target=\"_blank\">Multi-Label Knowledge Distillation</a> led to an improvement in accuracy.</p>\n<pre><code>     () -&gt; torch.Tensor:\n        \n        loss_focal = self.focal_weight * self.focal_loss(logits_student, target.())\n\n        \n        prob_student = torch.sigmoid(logits_student)\n        prob_teacher = torch.sigmoid(logits_teacher)\n        prob_student = torch.clamp(prob_student, =self.eps, =-self.eps)\n        prob_teacher = torch.clamp(prob_teacher, =self.eps, =-self.eps)\n        loss_mld = self.kl_div_loss(torch.log(prob_student), prob_teacher) + self.kl_div_loss(torch.log( - prob_student),  - prob_teacher)\n        loss_mld = loss_mld.mean()\n        loss_mld = (epoch / self.warmup, ) * loss_mld\n\n         loss_focal + loss_mld\n</code></pre>\n<p>◆Pattern 1</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ECA NFNet L0 (teacher)</td>\n<td>0.685730</td>\n<td>0.637945</td>\n</tr>\n<tr>\n<td>EfficientNet B0 (student)</td>\n<td>0.681507</td>\n<td>0.638541</td>\n</tr>\n<tr>\n<td>Focal + KD</td>\n<td>0.674996</td>\n<td>0.639671</td>\n</tr>\n<tr>\n<td>Focal + DKD</td>\n<td>0.657009</td>\n<td>0.629826</td>\n</tr>\n<tr>\n<td>Focal + MLD</td>\n<td><strong>0.695467</strong></td>\n<td><strong>0.655356</strong></td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<p>◆Pattern 2</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ECA NFNet L1 (teacher)</td>\n<td>0.690174</td>\n<td>0.630097</td>\n</tr>\n<tr>\n<td>EfficientNet B0 (student)</td>\n<td>0.661563</td>\n<td>0.622148</td>\n</tr>\n<tr>\n<td>Focal + KD</td>\n<td>N/A</td>\n<td>N/A</td>\n</tr>\n<tr>\n<td>Focal + DKD</td>\n<td>N/A</td>\n<td>N/A</td>\n</tr>\n<tr>\n<td>Focal + MLD</td>\n<td><strong>0.704317</strong></td>\n<td><strong>0.633722</strong></td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<p>◆Pattern 3</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ECA NFNet L1 (teacher)</td>\n<td>0.680272</td>\n<td>0.658853</td>\n</tr>\n<tr>\n<td>EfficientNet B1 (student)</td>\n<td>0.661441</td>\n<td>0.620913</td>\n</tr>\n<tr>\n<td>Focal + KD</td>\n<td>N/A</td>\n<td>N/A</td>\n</tr>\n<tr>\n<td>Focal + DKD</td>\n<td>N/A</td>\n<td>N/A</td>\n</tr>\n<tr>\n<td>Focal + MLD</td>\n<td><strong>0.693386</strong></td>\n<td><strong>0.667936</strong></td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<p>KD: Logit KD, DKD: <a href=\"https://arxiv.org/abs/2203.08679\" target=\"_blank\">Decoupled Knowledge Distillation</a>, MLD: $L_{MLD}$ from <a href=\"https://arxiv.org/abs/2308.06453\" target=\"_blank\">Multi-Label Knowledge Distillation</a></p>\n<h1>Reflections</h1>\n<ul>\n<li>✅ Distillation with MLD Loss proved effective.</li>\n<li>✅ Having models with varied training settings helped us withstand shake and create a slightly more accurate solution than the public notebooks.</li>\n<li>❌ We struggled to determine the best approach.</li>\n</ul>\n<h1>Appendix</h1>\n<p>Code for converting <code>eca_nfnet_l*</code> from timm to OpenVINO format.</p>\n<pre><code> copy\n functools  reduce\n\n torch.nn  nn\n torch.nn.functional  F\n\n timm\n\n\n () -&gt; nn.Module:\n    names = access_string.split(sep=)\n     reduce(, names, module)\n\n\n () -&gt; nn.Module:\n    converted_model = copy.deepcopy(model)\n    module_table = (converted_model.named_modules())\n     name, m  converted_model.named_modules():\n         (m, timm.layers.std_conv.ScaledStdConv2d):\n            scaled_weight = F.batch_norm(\n                m.weight.reshape(, m.out_channels, -), , ,\n                weight=(m.gain * m.scale).view(-),\n                training=, momentum=, eps=m.eps).reshape_as(m.weight).detach()\n\n            bias = m.bias   \n            conv = nn.Conv2d(m.in_channels, m.out_channels, m.kernel_size,\n                             stride=m.stride, padding=m.padding, dilation=m.dilation,\n                             groups=m.groups, bias=bias, padding_mode=m.padding_mode)\n            conv.weight.data = scaled_weight\n             bias:\n                conv.bias.data = m.bias\n\n            \n            parent_name, child_name = name.rsplit(, )\n            parent_module = module_table[parent_name]\n            (parent_module, child_name, conv)\n\n     converted_model\n</code></pre>",
  "messages": [
    {
      "id": "2868417",
      "postDate": "06/12/2024 11:12:49",
      "content": "<h1>Summary</h1>\n<p>I used various training setups (spectrogram settings, model types, data augmentation) to perform logit distillation and assembled an ensemble of four distillation models with different configurations.</p>\n<h1>Distillation</h1>\n<p>Using the $L_{MLD}$ from <a href=\"https://arxiv.org/abs/2308.06453\" target=\"_blank\">Multi-Label Knowledge Distillation</a> led to an improvement in accuracy.</p>\n<pre><code>     () -&gt; torch.Tensor:\n        \n        loss_focal = self.focal_weight * self.focal_loss(logits_student, target.())\n\n        \n        prob_student = torch.sigmoid(logits_student)\n        prob_teacher = torch.sigmoid(logits_teacher)\n        prob_student = torch.clamp(prob_student, =self.eps, =-self.eps)\n        prob_teacher = torch.clamp(prob_teacher, =self.eps, =-self.eps)\n        loss_mld = self.kl_div_loss(torch.log(prob_student), prob_teacher) + self.kl_div_loss(torch.log( - prob_student),  - prob_teacher)\n        loss_mld = loss_mld.mean()\n        loss_mld = (epoch / self.warmup, ) * loss_mld\n\n         loss_focal + loss_mld\n</code></pre>\n<p>◆Pattern 1</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ECA NFNet L0 (teacher)</td>\n<td>0.685730</td>\n<td>0.637945</td>\n</tr>\n<tr>\n<td>EfficientNet B0 (student)</td>\n<td>0.681507</td>\n<td>0.638541</td>\n</tr>\n<tr>\n<td>Focal + KD</td>\n<td>0.674996</td>\n<td>0.639671</td>\n</tr>\n<tr>\n<td>Focal + DKD</td>\n<td>0.657009</td>\n<td>0.629826</td>\n</tr>\n<tr>\n<td>Focal + MLD</td>\n<td><strong>0.695467</strong></td>\n<td><strong>0.655356</strong></td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<p>◆Pattern 2</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ECA NFNet L1 (teacher)</td>\n<td>0.690174</td>\n<td>0.630097</td>\n</tr>\n<tr>\n<td>EfficientNet B0 (student)</td>\n<td>0.661563</td>\n<td>0.622148</td>\n</tr>\n<tr>\n<td>Focal + KD</td>\n<td>N/A</td>\n<td>N/A</td>\n</tr>\n<tr>\n<td>Focal + DKD</td>\n<td>N/A</td>\n<td>N/A</td>\n</tr>\n<tr>\n<td>Focal + MLD</td>\n<td><strong>0.704317</strong></td>\n<td><strong>0.633722</strong></td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<p>◆Pattern 3</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ECA NFNet L1 (teacher)</td>\n<td>0.680272</td>\n<td>0.658853</td>\n</tr>\n<tr>\n<td>EfficientNet B1 (student)</td>\n<td>0.661441</td>\n<td>0.620913</td>\n</tr>\n<tr>\n<td>Focal + KD</td>\n<td>N/A</td>\n<td>N/A</td>\n</tr>\n<tr>\n<td>Focal + DKD</td>\n<td>N/A</td>\n<td>N/A</td>\n</tr>\n<tr>\n<td>Focal + MLD</td>\n<td><strong>0.693386</strong></td>\n<td><strong>0.667936</strong></td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<p>KD: Logit KD, DKD: <a href=\"https://arxiv.org/abs/2203.08679\" target=\"_blank\">Decoupled Knowledge Distillation</a>, MLD: $L_{MLD}$ from <a href=\"https://arxiv.org/abs/2308.06453\" target=\"_blank\">Multi-Label Knowledge Distillation</a></p>\n<h1>Reflections</h1>\n<ul>\n<li>✅ Distillation with MLD Loss proved effective.</li>\n<li>✅ Having models with varied training settings helped us withstand shake and create a slightly more accurate solution than the public notebooks.</li>\n<li>❌ We struggled to determine the best approach.</li>\n</ul>\n<h1>Appendix</h1>\n<p>Code for converting <code>eca_nfnet_l*</code> from timm to OpenVINO format.</p>\n<pre><code> copy\n functools  reduce\n\n torch.nn  nn\n torch.nn.functional  F\n\n timm\n\n\n () -&gt; nn.Module:\n    names = access_string.split(sep=)\n     reduce(, names, module)\n\n\n () -&gt; nn.Module:\n    converted_model = copy.deepcopy(model)\n    module_table = (converted_model.named_modules())\n     name, m  converted_model.named_modules():\n         (m, timm.layers.std_conv.ScaledStdConv2d):\n            scaled_weight = F.batch_norm(\n                m.weight.reshape(, m.out_channels, -), , ,\n                weight=(m.gain * m.scale).view(-),\n                training=, momentum=, eps=m.eps).reshape_as(m.weight).detach()\n\n            bias = m.bias   \n            conv = nn.Conv2d(m.in_channels, m.out_channels, m.kernel_size,\n                             stride=m.stride, padding=m.padding, dilation=m.dilation,\n                             groups=m.groups, bias=bias, padding_mode=m.padding_mode)\n            conv.weight.data = scaled_weight\n             bias:\n                conv.bias.data = m.bias\n\n            \n            parent_name, child_name = name.rsplit(, )\n            parent_module = module_table[parent_name]\n            (parent_module, child_name, conv)\n\n     converted_model\n</code></pre>",
      "rawMarkdown": "# Summary\n\nI used various training setups (spectrogram settings, model types, data augmentation) to perform logit distillation and assembled an ensemble of four distillation models with different configurations.\n\n# Distillation\n\nUsing the $L_{MLD}$ from [Multi-Label Knowledge Distillation](https://arxiv.org/abs/2308.06453) led to an improvement in accuracy.\n\n```python\n    def forward(\n        self,\n        logits_student: torch.Tensor,\n        logits_teacher: torch.Tensor,\n        target: torch.Tensor,\n        epoch: int,\n    ) -> torch.Tensor:\n        # focal loss\n        loss_focal = self.focal_weight * self.focal_loss(logits_student, target.float())\n\n        # mld loss\n        prob_student = torch.sigmoid(logits_student)\n        prob_teacher = torch.sigmoid(logits_teacher)\n        prob_student = torch.clamp(prob_student, min=self.eps, max=1-self.eps)\n        prob_teacher = torch.clamp(prob_teacher, min=self.eps, max=1-self.eps)\n        loss_mld = self.kl_div_loss(torch.log(prob_student), prob_teacher) + self.kl_div_loss(torch.log(1 - prob_student), 1 - prob_teacher)\n        loss_mld = loss_mld.mean()\n        loss_mld = min(epoch / self.warmup, 1.0) * loss_mld\n\n        return loss_focal + loss_mld\n```\n\n◆Pattern 1\n\n| Model                     | Public   | Private  |\n| ------------------------- | -------- | -------- |\n| ECA NFNet L0 (teacher)    | 0.685730 | 0.637945 |\n| EfficientNet B0 (student) | 0.681507 | 0.638541 |\n| Focal + KD                | 0.674996 | 0.639671 |\n| Focal + DKD               | 0.657009 | 0.629826 |\n| Focal + MLD               | **0.695467** | **0.655356** |\n\n<br>\n\n◆Pattern 2\n\n| Model                     | Public   | Private  |\n| ------------------------- | -------- | -------- |\n| ECA NFNet L1 (teacher)    | 0.690174 | 0.630097 |\n| EfficientNet B0 (student) | 0.661563 | 0.622148 |\n| Focal + KD                | N/A      | N/A      |\n| Focal + DKD               | N/A      | N/A      |\n| Focal + MLD               | **0.704317** | **0.633722** |\n\n<br>\n\n◆Pattern 3\n\n| Model                     | Public   | Private  |\n| ------------------------- | -------- | -------- |\n| ECA NFNet L1 (teacher)    | 0.680272 | 0.658853 |\n| EfficientNet B1 (student) | 0.661441 | 0.620913 |\n| Focal + KD                | N/A      | N/A      |\n| Focal + DKD               | N/A      | N/A      |\n| Focal + MLD               | **0.693386** | **0.667936** |\n\n<br>\n\nKD: Logit KD, DKD: [Decoupled Knowledge Distillation](https://arxiv.org/abs/2203.08679), MLD: $L_{MLD}$ from [Multi-Label Knowledge Distillation](https://arxiv.org/abs/2308.06453)\n\n# Reflections\n\n* ✅ Distillation with MLD Loss proved effective.\n* ✅ Having models with varied training settings helped us withstand shake and create a slightly more accurate solution than the public notebooks.\n* ❌ We struggled to determine the best approach.\n\n# Appendix\n\nCode for converting `eca_nfnet_l*` from timm to OpenVINO format.\n\n```python\nimport copy\nfrom functools import reduce\n\nimport torch.nn as nn\nimport torch.nn.functional as F\n\nimport timm\n\n\ndef get_module_by_name(module: nn.Module, access_string: str) -> nn.Module:\n    names = access_string.split(sep='.')\n    return reduce(getattr, names, module)\n\n\ndef convert_scaled_std_conv2d_to_conv2d(model: nn.Module) -> nn.Module:\n    converted_model = copy.deepcopy(model)\n    module_table = dict(converted_model.named_modules())\n    for name, m in converted_model.named_modules():\n        if isinstance(m, timm.layers.std_conv.ScaledStdConv2d):\n            scaled_weight = F.batch_norm(\n                m.weight.reshape(1, m.out_channels, -1), None, None,\n                weight=(m.gain * m.scale).view(-1),\n                training=True, momentum=0., eps=m.eps).reshape_as(m.weight).detach()\n\n            bias = m.bias is not None\n            conv = nn.Conv2d(m.in_channels, m.out_channels, m.kernel_size,\n                             stride=m.stride, padding=m.padding, dilation=m.dilation,\n                             groups=m.groups, bias=bias, padding_mode=m.padding_mode)\n            conv.weight.data = scaled_weight\n            if bias:\n                conv.bias.data = m.bias\n\n            # replace ScaledStdConv2d to nn.Conv2d\n            parent_name, child_name = name.rsplit('.', 1)\n            parent_module = module_table[parent_name]\n            setattr(parent_module, child_name, conv)\n\n    return converted_model\n```",
      "votes": null
    },
    {
      "id": "2868703",
      "postDate": "06/12/2024 15:11:45",
      "content": "<p>Kudos on your remarkable achievement! <a href=\"https://www.kaggle.com/akinosora\" target=\"_blank\">@akinosora</a> </p>",
      "rawMarkdown": "Kudos on your remarkable achievement! @akinosora",
      "votes": null
    },
    {
      "id": "2878306",
      "postDate": "06/18/2024 21:12:22",
      "content": "<p>Hi, thanks for the notes!<br>\nWhat were you distilling from? Pretrained model (birdnet, bird-vocalization-classifier) or larger custom-trained models?</p>",
      "rawMarkdown": "Hi, thanks for the notes!\nWhat were you distilling from? Pretrained model (birdnet, bird-vocalization-classifier) or larger custom-trained models?",
      "votes": null
    },
    {
      "id": "2890934",
      "postDate": "06/26/2024 12:07:05",
      "content": "<p>larger custom-trained models (the models listed as teachers in the table).</p>",
      "rawMarkdown": "larger custom-trained models (the models listed as teachers in the table).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2868703,
      "author_name": "mubashirsidiki",
      "author_url": "",
      "post_date": "06/12/2024 15:11:45",
      "content": "<p>Kudos on your remarkable achievement! <a href=\"https://www.kaggle.com/akinosora\" target=\"_blank\">@akinosora</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2878306,
      "author_name": "tomdenton",
      "author_url": "",
      "post_date": "06/18/2024 21:12:22",
      "content": "<p>Hi, thanks for the notes!<br>\nWhat were you distilling from? Pretrained model (birdnet, bird-vocalization-classifier) or larger custom-trained models?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2890934,
          "author_name": "akinosora",
          "author_url": "",
          "post_date": "06/26/2024 12:07:05",
          "content": "<p>larger custom-trained models (the models listed as teachers in the table).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2868417": "# Summary\n\nI used various training setups (spectrogram settings, model types, data augmentation) to perform logit distillation and assembled an ensemble of four distillation models with different configurations.\n\n# Distillation\n\nUsing the $L_{MLD}$ from [Multi-Label Knowledge Distillation](https://arxiv.org/abs/2308.06453) led to an improvement in accuracy.\n\n```python\n    def forward(\n        self,\n        logits_student: torch.Tensor,\n        logits_teacher: torch.Tensor,\n        target: torch.Tensor,\n        epoch: int,\n    ) -> torch.Tensor:\n        # focal loss\n        loss_focal = self.focal_weight * self.focal_loss(logits_student, target.float())\n\n        # mld loss\n        prob_student = torch.sigmoid(logits_student)\n        prob_teacher = torch.sigmoid(logits_teacher)\n        prob_student = torch.clamp(prob_student, min=self.eps, max=1-self.eps)\n        prob_teacher = torch.clamp(prob_teacher, min=self.eps, max=1-self.eps)\n        loss_mld = self.kl_div_loss(torch.log(prob_student), prob_teacher) + self.kl_div_loss(torch.log(1 - prob_student), 1 - prob_teacher)\n        loss_mld = loss_mld.mean()\n        loss_mld = min(epoch / self.warmup, 1.0) * loss_mld\n\n        return loss_focal + loss_mld\n```\n\n◆Pattern 1\n\n| Model                     | Public   | Private  |\n| ------------------------- | -------- | -------- |\n| ECA NFNet L0 (teacher)    | 0.685730 | 0.637945 |\n| EfficientNet B0 (student) | 0.681507 | 0.638541 |\n| Focal + KD                | 0.674996 | 0.639671 |\n| Focal + DKD               | 0.657009 | 0.629826 |\n| Focal + MLD               | **0.695467** | **0.655356** |\n\n<br>\n\n◆Pattern 2\n\n| Model                     | Public   | Private  |\n| ------------------------- | -------- | -------- |\n| ECA NFNet L1 (teacher)    | 0.690174 | 0.630097 |\n| EfficientNet B0 (student) | 0.661563 | 0.622148 |\n| Focal + KD                | N/A      | N/A      |\n| Focal + DKD               | N/A      | N/A      |\n| Focal + MLD               | **0.704317** | **0.633722** |\n\n<br>\n\n◆Pattern 3\n\n| Model                     | Public   | Private  |\n| ------------------------- | -------- | -------- |\n| ECA NFNet L1 (teacher)    | 0.680272 | 0.658853 |\n| EfficientNet B1 (student) | 0.661441 | 0.620913 |\n| Focal + KD                | N/A      | N/A      |\n| Focal + DKD               | N/A      | N/A      |\n| Focal + MLD               | **0.693386** | **0.667936** |\n\n<br>\n\nKD: Logit KD, DKD: [Decoupled Knowledge Distillation](https://arxiv.org/abs/2203.08679), MLD: $L_{MLD}$ from [Multi-Label Knowledge Distillation](https://arxiv.org/abs/2308.06453)\n\n# Reflections\n\n* ✅ Distillation with MLD Loss proved effective.\n* ✅ Having models with varied training settings helped us withstand shake and create a slightly more accurate solution than the public notebooks.\n* ❌ We struggled to determine the best approach.\n\n# Appendix\n\nCode for converting `eca_nfnet_l*` from timm to OpenVINO format.\n\n```python\nimport copy\nfrom functools import reduce\n\nimport torch.nn as nn\nimport torch.nn.functional as F\n\nimport timm\n\n\ndef get_module_by_name(module: nn.Module, access_string: str) -> nn.Module:\n    names = access_string.split(sep='.')\n    return reduce(getattr, names, module)\n\n\ndef convert_scaled_std_conv2d_to_conv2d(model: nn.Module) -> nn.Module:\n    converted_model = copy.deepcopy(model)\n    module_table = dict(converted_model.named_modules())\n    for name, m in converted_model.named_modules():\n        if isinstance(m, timm.layers.std_conv.ScaledStdConv2d):\n            scaled_weight = F.batch_norm(\n                m.weight.reshape(1, m.out_channels, -1), None, None,\n                weight=(m.gain * m.scale).view(-1),\n                training=True, momentum=0., eps=m.eps).reshape_as(m.weight).detach()\n\n            bias = m.bias is not None\n            conv = nn.Conv2d(m.in_channels, m.out_channels, m.kernel_size,\n                             stride=m.stride, padding=m.padding, dilation=m.dilation,\n                             groups=m.groups, bias=bias, padding_mode=m.padding_mode)\n            conv.weight.data = scaled_weight\n            if bias:\n                conv.bias.data = m.bias\n\n            # replace ScaledStdConv2d to nn.Conv2d\n            parent_name, child_name = name.rsplit('.', 1)\n            parent_module = module_table[parent_name]\n            setattr(parent_module, child_name, conv)\n\n    return converted_model\n```",
    "2868703": "Kudos on your remarkable achievement! @akinosora",
    "2878306": "Hi, thanks for the notes!\nWhat were you distilling from? Pretrained model (birdnet, bird-vocalization-classifier) or larger custom-trained models?",
    "2890934": "larger custom-trained models (the models listed as teachers in the table)."
  },
  "source": "meta"
}