{
  "id": 220436,
  "title": "21st place solution - FP co-teaching with loss improvements",
  "url": "/competitions/rfcx-species-audio-detection/writeups/vasiliy-kotov-21st-place-solution-fp-co-teaching-w",
  "author_name": "",
  "post_date": "2021-02-18T21:32:20.757Z",
  "votes": 22,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Congratulations to the top finishers!</p>\n<p>The whole my solution is here: <a href=\"https://github.com/MPGek/mpgek-rfcx-species-audio-detection-public\" target=\"_blank\">https://github.com/MPGek/mpgek-rfcx-species-audio-detection-public</a>.<br>\nI left the best configs that were used in the final submission.</p>\n<h2>Summary</h2>\n<p>Summary of my solution:</p>\n<ul>\n<li>General augmentations with frequency band augmentations:<ul>\n<li>Frequency band filtering</li>\n<li>Frequency band mixup</li></ul></li>\n<li>TP training with a combined loss of the BCE for confident samples and LSoft 0.7 for noisy samples</li>\n<li>FP training with a BCE for confident samples and ignored losses for samples with noisy samples</li>\n<li>FP co-teaching training with loss as described in the co-teaching paper, but with the extra loss for the high loss samples</li>\n<li>Ensemble of TP, FP, FP co-teaching results.</li>\n</ul>\n<h2>Spectrogram and random crop</h2>\n<p>For the training, I used a random crop with the centered sample in 10 seconds and a random crop for 6 seconds.<br>\nFor validation and prediction, I used 6 seconds crop with a stride in 2 seconds with maximizing outputs.<br>\nFor the mel spectrogram, I used the following parameters:</p>\n<ul>\n<li>mels count: 380</li>\n<li>FTT: 4096</li>\n<li>window length: 1536</li>\n<li>hop length: 400</li>\n<li>fmin: 50</li>\n<li>fmax: 15000</li>\n</ul>\n<p>So one sample in 6 seconds produced an image with size 720 x 380 pixels</p>\n<h2>Augmentations</h2>\n<p><strong>Augmentation that improves LB:</strong></p>\n<ul>\n<li>Gaussian noise</li>\n<li>Random crop + resize with size reduction in 40 and 20 pixels</li>\n<li>Frequency band filtering - based on f_min and f_max of the sample I set 0 values to all mels that lower or higher than f_min and f_max with a sigmoid transition to remove sharp edges<ul>\n<li>Frequency band mixup - for some samples I used frequency band filtering and then I mixed it with different samples with a band that higher than f_max and with a band that lower than f_min. So I managed to get a single image with the 3 mixed samples.</li></ul></li>\n</ul>\n<p><strong>Example of the Frequency band filtering (top - original sample, bottom - sample after filtering):</strong><br>\n<img src=\"https://github.com/MPGek/mpgek-rfcx-species-audio-detection-public/raw/main/img/Band%20filtering.png\" alt=\"\"><br>\n<strong>Example of the Frequency band mixup (top - original sample, bottom - sample after mixup):</strong><br>\n<img src=\"https://raw.githubusercontent.com/MPGek/mpgek-rfcx-species-audio-detection-public/main/img/Band%20mixup.png\" alt=\"\"><br>\n<strong>Augmentation that doesn't improve LB:</strong></p>\n<ul>\n<li>SpecAugment</li>\n<li>Sum mixup</li>\n</ul>\n<h2>Network topology</h2>\n<p>I got the best results on EfficientNetB2, B4, B7 (noisy students weights) with a simple FC to 24 classes after the adaptive average pool.</p>\n<p>I tried to use different head but all of them gave the same or worse result:</p>\n<ul>\n<li>the hyper column</li>\n<li>the hyper column with auxiliary losses after each pooling</li>\n<li>extra spatial and channel attentions blocks - CBAM</li>\n<li>dense convolutions</li>\n</ul>\n<h2>TP training</h2>\n<p>Based on the post where described that every sound file can present unlabeled samples I have to work noisy samples.</p>\n<p>I split all samples into confident samples and noisy samples:</p>\n<ul>\n<li>confident samples - all sigmoid outputs for classes which present in the train_tp.csv file (1 in targets tensor)</li>\n<li>noisy samples - all sigmoid outputs for classes which not described in the train_tp.csv file (0 in target tensor)</li>\n</ul>\n<p>For the confident samples, I used simple BCE. For the noisy samples, I used LSoft with beta 0.7 (<a href=\"https://arxiv.org/pdf/1901.01189.pdf)\" target=\"_blank\">https://arxiv.org/pdf/1901.01189.pdf)</a>. <br>\nLSoft:</p>\n<pre><code>def forward(self, input: torch.Tensor, target: torch.Tensor):\n        with torch.no_grad():\n            pred = torch.sigmoid(input)\n            target_update = self.beta * target + (1 - self.beta) * pred\n        loss = F.binary_cross_entropy_with_logits(input, target_update, reduction=self.reduction)\n        return loss\n</code></pre>\n<p>In the loss function, I flatten all outputs (even batch dim) to a linear array. And split items into 2 arrays where targets were 1 and where targets were 0.</p>\n<p>With LSoft I got on EfficientNetB7 0.912-0.915 LB.<br>\nWithout LSoft - BCE for all samples I got only about 0.895-0.900.</p>\n<h2>FP training</h2>\n<p>For the FP training, I used a dataset with undersampling of the FP samples. Each epoch had all TP samples and the same count of the FP samples.<br>\nI used batch sampling to provide a balanced batch of TP/FP samples - after each TP I added FP with the same species id.</p>\n<p>In the loss function, I calculate loss only for those sigmoid outputs that present in train_tp.csv or train_fp.csv. So all noisy samples are ignored.</p>\n<h2>FP co-teaching training</h2>\n<p>I have tried to find a way how to use FOCI or SELFIE to work with noisy data, but all of them use historical predictions of each sample. With my random crop and frequency band mixup it's almost impossible. Even shift for 0.5-1 seconds can add a new species to the sample. So historical data will be incorrect.</p>\n<p>I tried co-teaching training because it doesn't require historical data.<br>\nPaper: <a href=\"https://arxiv.org/pdf/1804.06872.pdf\" target=\"_blank\">https://arxiv.org/pdf/1804.06872.pdf</a><br>\nCode sample: <a href=\"https://github.com/bhanML/Co-teaching\" target=\"_blank\">https://github.com/bhanML/Co-teaching</a></p>\n<p>When I implemented co-teaching training I have only 5 days before the deadline.<br>\nThe first experiments with co-teaching gave 0.830 LB for the TP and 0.880 LB for the FP. So it looked like a bad experiment.</p>\n<p>I tried to improve loss function by adding high loss samples with changed targets (by default co-teaching should ignore high loss samples as wrong samples).</p>\n<p>The final loss function consists of:</p>\n<ul>\n<li>50% lowest loss samples (all confident samples mandatory added to this part of loss with scale factor 2)</li>\n<li>45% ignored losses</li>\n<li>5% highest loss samples with the changed target (1 for predictions with sigmoid &gt;= 0.5 and 0 for sigmoid &lt; 0.5)<br>\nThe loss implementation is here: <a href=\"https://github.com/MPGek/mpgek-rfcx-species-audio-detection-public/blob/main/model/forward_passes_coteaching.py\" target=\"_blank\">https://github.com/MPGek/mpgek-rfcx-species-audio-detection-public/blob/main/model/forward_passes_coteaching.py</a></li>\n</ul>\n<p>When I invented this loss I had only 3 days before the deadline.<br>\nThis loss has a good potential for future experiments. I trained only 2 experiments with this loss, both with the same hyperparameters. The first one had a bug so it produces extremely high metrics and used wrong epochs to submission - however, even with the bug it produces a good LB score comparing to the original FP training.</p>\n<h2>Folds and best epoch metric</h2>\n<p>To chose the best epoch in all training I used the BCE which calculated only on confident samples.<br>\nIn some experiments I used 5 stratified KFold, in some, I used 7 folds with stratified shuffle split with test size 0.3.</p>\n<h2>Ensembles</h2>\n<p>The TP training with EfficientNetB7 gave me only 0.912-0.915 on the public LB.<br>\nThe FP training with EfficientNetB2-B4 gave only 0.887-0.893 LB.<br>\nThe ensemble of the TP and FP gave 0.929 LB.</p>\n<p>The FP co-teaching training on simple EfficientNetB2 gave me 0.925 LB (a quite good improvement from the original FP with 0.893)</p>\n<p><strong>The final ensemble consists of all best experiments (0.941 public LB and 0.944 private LB):</strong></p>\n<ul>\n<li>TP EfficientNetB7 with 0.915</li>\n<li>FP EfficientNetB2-B4 with 0.893</li>\n<li>FP co-teaching EfficientNetB2 with 0.925</li>\n</ul>",
  "messages": [
    {
      "id": "1208460",
      "postDate": "02/18/2021 09:46:55",
      "content": "<p>Congratulations to the top finishers!</p>\n<p>The whole my solution is here: <a href=\"https://github.com/MPGek/mpgek-rfcx-species-audio-detection-public\" target=\"_blank\">https://github.com/MPGek/mpgek-rfcx-species-audio-detection-public</a>.<br>\nI left the best configs that were used in the final submission.</p>\n<h2>Summary</h2>\n<p>Summary of my solution:</p>\n<ul>\n<li>General augmentations with frequency band augmentations:<ul>\n<li>Frequency band filtering</li>\n<li>Frequency band mixup</li></ul></li>\n<li>TP training with a combined loss of the BCE for confident samples and LSoft 0.7 for noisy samples</li>\n<li>FP training with a BCE for confident samples and ignored losses for samples with noisy samples</li>\n<li>FP co-teaching training with loss as described in the co-teaching paper, but with the extra loss for the high loss samples</li>\n<li>Ensemble of TP, FP, FP co-teaching results.</li>\n</ul>\n<h2>Spectrogram and random crop</h2>\n<p>For the training, I used a random crop with the centered sample in 10 seconds and a random crop for 6 seconds.<br>\nFor validation and prediction, I used 6 seconds crop with a stride in 2 seconds with maximizing outputs.<br>\nFor the mel spectrogram, I used the following parameters:</p>\n<ul>\n<li>mels count: 380</li>\n<li>FTT: 4096</li>\n<li>window length: 1536</li>\n<li>hop length: 400</li>\n<li>fmin: 50</li>\n<li>fmax: 15000</li>\n</ul>\n<p>So one sample in 6 seconds produced an image with size 720 x 380 pixels</p>\n<h2>Augmentations</h2>\n<p><strong>Augmentation that improves LB:</strong></p>\n<ul>\n<li>Gaussian noise</li>\n<li>Random crop + resize with size reduction in 40 and 20 pixels</li>\n<li>Frequency band filtering - based on f_min and f_max of the sample I set 0 values to all mels that lower or higher than f_min and f_max with a sigmoid transition to remove sharp edges<ul>\n<li>Frequency band mixup - for some samples I used frequency band filtering and then I mixed it with different samples with a band that higher than f_max and with a band that lower than f_min. So I managed to get a single image with the 3 mixed samples.</li></ul></li>\n</ul>\n<p><strong>Example of the Frequency band filtering (top - original sample, bottom - sample after filtering):</strong><br>\n<img src=\"https://github.com/MPGek/mpgek-rfcx-species-audio-detection-public/raw/main/img/Band%20filtering.png\" alt=\"\"><br>\n<strong>Example of the Frequency band mixup (top - original sample, bottom - sample after mixup):</strong><br>\n<img src=\"https://raw.githubusercontent.com/MPGek/mpgek-rfcx-species-audio-detection-public/main/img/Band%20mixup.png\" alt=\"\"><br>\n<strong>Augmentation that doesn't improve LB:</strong></p>\n<ul>\n<li>SpecAugment</li>\n<li>Sum mixup</li>\n</ul>\n<h2>Network topology</h2>\n<p>I got the best results on EfficientNetB2, B4, B7 (noisy students weights) with a simple FC to 24 classes after the adaptive average pool.</p>\n<p>I tried to use different head but all of them gave the same or worse result:</p>\n<ul>\n<li>the hyper column</li>\n<li>the hyper column with auxiliary losses after each pooling</li>\n<li>extra spatial and channel attentions blocks - CBAM</li>\n<li>dense convolutions</li>\n</ul>\n<h2>TP training</h2>\n<p>Based on the post where described that every sound file can present unlabeled samples I have to work noisy samples.</p>\n<p>I split all samples into confident samples and noisy samples:</p>\n<ul>\n<li>confident samples - all sigmoid outputs for classes which present in the train_tp.csv file (1 in targets tensor)</li>\n<li>noisy samples - all sigmoid outputs for classes which not described in the train_tp.csv file (0 in target tensor)</li>\n</ul>\n<p>For the confident samples, I used simple BCE. For the noisy samples, I used LSoft with beta 0.7 (<a href=\"https://arxiv.org/pdf/1901.01189.pdf)\" target=\"_blank\">https://arxiv.org/pdf/1901.01189.pdf)</a>. <br>\nLSoft:</p>\n<pre><code>def forward(self, input: torch.Tensor, target: torch.Tensor):\n        with torch.no_grad():\n            pred = torch.sigmoid(input)\n            target_update = self.beta * target + (1 - self.beta) * pred\n        loss = F.binary_cross_entropy_with_logits(input, target_update, reduction=self.reduction)\n        return loss\n</code></pre>\n<p>In the loss function, I flatten all outputs (even batch dim) to a linear array. And split items into 2 arrays where targets were 1 and where targets were 0.</p>\n<p>With LSoft I got on EfficientNetB7 0.912-0.915 LB.<br>\nWithout LSoft - BCE for all samples I got only about 0.895-0.900.</p>\n<h2>FP training</h2>\n<p>For the FP training, I used a dataset with undersampling of the FP samples. Each epoch had all TP samples and the same count of the FP samples.<br>\nI used batch sampling to provide a balanced batch of TP/FP samples - after each TP I added FP with the same species id.</p>\n<p>In the loss function, I calculate loss only for those sigmoid outputs that present in train_tp.csv or train_fp.csv. So all noisy samples are ignored.</p>\n<h2>FP co-teaching training</h2>\n<p>I have tried to find a way how to use FOCI or SELFIE to work with noisy data, but all of them use historical predictions of each sample. With my random crop and frequency band mixup it's almost impossible. Even shift for 0.5-1 seconds can add a new species to the sample. So historical data will be incorrect.</p>\n<p>I tried co-teaching training because it doesn't require historical data.<br>\nPaper: <a href=\"https://arxiv.org/pdf/1804.06872.pdf\" target=\"_blank\">https://arxiv.org/pdf/1804.06872.pdf</a><br>\nCode sample: <a href=\"https://github.com/bhanML/Co-teaching\" target=\"_blank\">https://github.com/bhanML/Co-teaching</a></p>\n<p>When I implemented co-teaching training I have only 5 days before the deadline.<br>\nThe first experiments with co-teaching gave 0.830 LB for the TP and 0.880 LB for the FP. So it looked like a bad experiment.</p>\n<p>I tried to improve loss function by adding high loss samples with changed targets (by default co-teaching should ignore high loss samples as wrong samples).</p>\n<p>The final loss function consists of:</p>\n<ul>\n<li>50% lowest loss samples (all confident samples mandatory added to this part of loss with scale factor 2)</li>\n<li>45% ignored losses</li>\n<li>5% highest loss samples with the changed target (1 for predictions with sigmoid &gt;= 0.5 and 0 for sigmoid &lt; 0.5)<br>\nThe loss implementation is here: <a href=\"https://github.com/MPGek/mpgek-rfcx-species-audio-detection-public/blob/main/model/forward_passes_coteaching.py\" target=\"_blank\">https://github.com/MPGek/mpgek-rfcx-species-audio-detection-public/blob/main/model/forward_passes_coteaching.py</a></li>\n</ul>\n<p>When I invented this loss I had only 3 days before the deadline.<br>\nThis loss has a good potential for future experiments. I trained only 2 experiments with this loss, both with the same hyperparameters. The first one had a bug so it produces extremely high metrics and used wrong epochs to submission - however, even with the bug it produces a good LB score comparing to the original FP training.</p>\n<h2>Folds and best epoch metric</h2>\n<p>To chose the best epoch in all training I used the BCE which calculated only on confident samples.<br>\nIn some experiments I used 5 stratified KFold, in some, I used 7 folds with stratified shuffle split with test size 0.3.</p>\n<h2>Ensembles</h2>\n<p>The TP training with EfficientNetB7 gave me only 0.912-0.915 on the public LB.<br>\nThe FP training with EfficientNetB2-B4 gave only 0.887-0.893 LB.<br>\nThe ensemble of the TP and FP gave 0.929 LB.</p>\n<p>The FP co-teaching training on simple EfficientNetB2 gave me 0.925 LB (a quite good improvement from the original FP with 0.893)</p>\n<p><strong>The final ensemble consists of all best experiments (0.941 public LB and 0.944 private LB):</strong></p>\n<ul>\n<li>TP EfficientNetB7 with 0.915</li>\n<li>FP EfficientNetB2-B4 with 0.893</li>\n<li>FP co-teaching EfficientNetB2 with 0.925</li>\n</ul>",
      "rawMarkdown": "Congratulations to the top finishers!\n\nThe whole my solution is here: https://github.com/MPGek/mpgek-rfcx-species-audio-detection-public.\nI left the best configs that were used in the final submission.\n\n## Summary\nSummary of my solution:\n- General augmentations with frequency band augmentations:\n - Frequency band filtering\n - Frequency band mixup\n- TP training with a combined loss of the BCE for confident samples and LSoft 0.7 for noisy samples\n- FP training with a BCE for confident samples and ignored losses for samples with noisy samples\n- FP co-teaching training with loss as described in the co-teaching paper, but with the extra loss for the high loss samples\n- Ensemble of TP, FP, FP co-teaching results.\n\n## Spectrogram and random crop\nFor the training, I used a random crop with the centered sample in 10 seconds and a random crop for 6 seconds.\nFor validation and prediction, I used 6 seconds crop with a stride in 2 seconds with maximizing outputs.\nFor the mel spectrogram, I used the following parameters:\n- mels count: 380\n- FTT: 4096\n- window length: 1536\n- hop length: 400\n- fmin: 50\n- fmax: 15000\n\nSo one sample in 6 seconds produced an image with size 720 x 380 pixels\n\n## Augmentations\n**Augmentation that improves LB:**\n- Gaussian noise\n- Random crop + resize with size reduction in 40 and 20 pixels\n- Frequency band filtering - based on f_min and f_max of the sample I set 0 values to all mels that lower or higher than f_min and f_max with a sigmoid transition to remove sharp edges\n - Frequency band mixup - for some samples I used frequency band filtering and then I mixed it with different samples with a band that higher than f_max and with a band that lower than f_min. So I managed to get a single image with the 3 mixed samples.\n\n**Example of the Frequency band filtering (top - original sample, bottom - sample after filtering):**\n![](https://github.com/MPGek/mpgek-rfcx-species-audio-detection-public/raw/main/img/Band%20filtering.png)\n**Example of the Frequency band mixup (top - original sample, bottom - sample after mixup):**\n![](https://raw.githubusercontent.com/MPGek/mpgek-rfcx-species-audio-detection-public/main/img/Band%20mixup.png)\n**Augmentation that doesn't improve LB:**\n- SpecAugment\n- Sum mixup\n\n## Network topology\n\nI got the best results on EfficientNetB2, B4, B7 (noisy students weights) with a simple FC to 24 classes after the adaptive average pool.\n\nI tried to use different head but all of them gave the same or worse result:\n- the hyper column\n- the hyper column with auxiliary losses after each pooling\n- extra spatial and channel attentions blocks - CBAM\n- dense convolutions\n\n## TP training\nBased on the post where described that every sound file can present unlabeled samples I have to work noisy samples.\n\nI split all samples into confident samples and noisy samples:\n- confident samples - all sigmoid outputs for classes which present in the train_tp.csv file (1 in targets tensor)\n- noisy samples - all sigmoid outputs for classes which not described in the train_tp.csv file (0 in target tensor)\n\nFor the confident samples, I used simple BCE. For the noisy samples, I used LSoft with beta 0.7 (https://arxiv.org/pdf/1901.01189.pdf). \nLSoft:\n```\ndef forward(self, input: torch.Tensor, target: torch.Tensor):\n        with torch.no_grad():\n            pred = torch.sigmoid(input)\n            target_update = self.beta * target + (1 - self.beta) * pred\n        loss = F.binary_cross_entropy_with_logits(input, target_update, reduction=self.reduction)\n        return loss\n```\n\nIn the loss function, I flatten all outputs (even batch dim) to a linear array. And split items into 2 arrays where targets were 1 and where targets were 0.\n\nWith LSoft I got on EfficientNetB7 0.912-0.915 LB.\nWithout LSoft - BCE for all samples I got only about 0.895-0.900.\n\n## FP training\nFor the FP training, I used a dataset with undersampling of the FP samples. Each epoch had all TP samples and the same count of the FP samples.\nI used batch sampling to provide a balanced batch of TP/FP samples - after each TP I added FP with the same species id.\n\nIn the loss function, I calculate loss only for those sigmoid outputs that present in train_tp.csv or train_fp.csv. So all noisy samples are ignored.\n\n## FP co-teaching training\nI have tried to find a way how to use FOCI or SELFIE to work with noisy data, but all of them use historical predictions of each sample. With my random crop and frequency band mixup it's almost impossible. Even shift for 0.5-1 seconds can add a new species to the sample. So historical data will be incorrect.\n\nI tried co-teaching training because it doesn't require historical data.\nPaper: https://arxiv.org/pdf/1804.06872.pdf\nCode sample: https://github.com/bhanML/Co-teaching\n\nWhen I implemented co-teaching training I have only 5 days before the deadline.\nThe first experiments with co-teaching gave 0.830 LB for the TP and 0.880 LB for the FP. So it looked like a bad experiment.\n\nI tried to improve loss function by adding high loss samples with changed targets (by default co-teaching should ignore high loss samples as wrong samples).\n\nThe final loss function consists of:\n- 50% lowest loss samples (all confident samples mandatory added to this part of loss with scale factor 2)\n- 45% ignored losses\n- 5% highest loss samples with the changed target (1 for predictions with sigmoid >= 0.5 and 0 for sigmoid < 0.5)\nThe loss implementation is here: https://github.com/MPGek/mpgek-rfcx-species-audio-detection-public/blob/main/model/forward_passes_coteaching.py\n\nWhen I invented this loss I had only 3 days before the deadline.\nThis loss has a good potential for future experiments. I trained only 2 experiments with this loss, both with the same hyperparameters. The first one had a bug so it produces extremely high metrics and used wrong epochs to submission - however, even with the bug it produces a good LB score comparing to the original FP training.\n\n## Folds and best epoch metric\nTo chose the best epoch in all training I used the BCE which calculated only on confident samples.\nIn some experiments I used 5 stratified KFold, in some, I used 7 folds with stratified shuffle split with test size 0.3.\n\n## Ensembles\nThe TP training with EfficientNetB7 gave me only 0.912-0.915 on the public LB.\nThe FP training with EfficientNetB2-B4 gave only 0.887-0.893 LB.\nThe ensemble of the TP and FP gave 0.929 LB.\n\nThe FP co-teaching training on simple EfficientNetB2 gave me 0.925 LB (a quite good improvement from the original FP with 0.893)\n\n**The final ensemble consists of all best experiments (0.941 public LB and 0.944 private LB):**\n- TP EfficientNetB7 with 0.915\n- FP EfficientNetB2-B4 with 0.893\n- FP co-teaching EfficientNetB2 with 0.925",
      "votes": null
    },
    {
      "id": "1208693",
      "postDate": "02/18/2021 12:47:12",
      "content": "<p>Good job, congrats on results and thanks for sharing solution and code <a href=\"https://www.kaggle.com/xmpgek\" target=\"_blank\">@xmpgek</a> </p>",
      "rawMarkdown": "Good job, congrats on results and thanks for sharing solution and code @xmpgek",
      "votes": null
    },
    {
      "id": "1209457",
      "postDate": "02/18/2021 22:56:41",
      "content": "<p>Thanks for sharing.  And congrats on the good result.  Also, thanks for mentioning yet another teacher/student paper I didn't know.  I'll read it for sure.</p>",
      "rawMarkdown": "Thanks for sharing.  And congrats on the good result.  Also, thanks for mentioning yet another teacher/student paper I didn't know.  I'll read it for sure.",
      "votes": null
    },
    {
      "id": "1219713",
      "postDate": "02/27/2021 06:57:26",
      "content": "<p>Thanks, I have starred!</p>",
      "rawMarkdown": "Thanks, I have starred!",
      "votes": null
    },
    {
      "id": "1255080",
      "postDate": "03/28/2021 11:58:36",
      "content": "<p>I want to ask something about your code. In <code>RainforestModel</code>, the <code>self.is_base_resnet</code> is always set to <code>False</code>, why is that?</p>",
      "rawMarkdown": "I want to ask something about your code. In `RainforestModel`, the `self.is_base_resnet` is always set to `False`, why is that?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1208693,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "02/18/2021 12:47:12",
      "content": "<p>Good job, congrats on results and thanks for sharing solution and code <a href=\"https://www.kaggle.com/xmpgek\" target=\"_blank\">@xmpgek</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209457,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/18/2021 22:56:41",
      "content": "<p>Thanks for sharing.  And congrats on the good result.  Also, thanks for mentioning yet another teacher/student paper I didn't know.  I'll read it for sure.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1219713,
      "author_name": "wubinbai",
      "author_url": "",
      "post_date": "02/27/2021 06:57:26",
      "content": "<p>Thanks, I have starred!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1255080,
      "author_name": "faraksuli",
      "author_url": "",
      "post_date": "03/28/2021 11:58:36",
      "content": "<p>I want to ask something about your code. In <code>RainforestModel</code>, the <code>self.is_base_resnet</code> is always set to <code>False</code>, why is that?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1208460": "Congratulations to the top finishers!\n\nThe whole my solution is here: https://github.com/MPGek/mpgek-rfcx-species-audio-detection-public.\nI left the best configs that were used in the final submission.\n\n## Summary\nSummary of my solution:\n- General augmentations with frequency band augmentations:\n - Frequency band filtering\n - Frequency band mixup\n- TP training with a combined loss of the BCE for confident samples and LSoft 0.7 for noisy samples\n- FP training with a BCE for confident samples and ignored losses for samples with noisy samples\n- FP co-teaching training with loss as described in the co-teaching paper, but with the extra loss for the high loss samples\n- Ensemble of TP, FP, FP co-teaching results.\n\n## Spectrogram and random crop\nFor the training, I used a random crop with the centered sample in 10 seconds and a random crop for 6 seconds.\nFor validation and prediction, I used 6 seconds crop with a stride in 2 seconds with maximizing outputs.\nFor the mel spectrogram, I used the following parameters:\n- mels count: 380\n- FTT: 4096\n- window length: 1536\n- hop length: 400\n- fmin: 50\n- fmax: 15000\n\nSo one sample in 6 seconds produced an image with size 720 x 380 pixels\n\n## Augmentations\n**Augmentation that improves LB:**\n- Gaussian noise\n- Random crop + resize with size reduction in 40 and 20 pixels\n- Frequency band filtering - based on f_min and f_max of the sample I set 0 values to all mels that lower or higher than f_min and f_max with a sigmoid transition to remove sharp edges\n - Frequency band mixup - for some samples I used frequency band filtering and then I mixed it with different samples with a band that higher than f_max and with a band that lower than f_min. So I managed to get a single image with the 3 mixed samples.\n\n**Example of the Frequency band filtering (top - original sample, bottom - sample after filtering):**\n![](https://github.com/MPGek/mpgek-rfcx-species-audio-detection-public/raw/main/img/Band%20filtering.png)\n**Example of the Frequency band mixup (top - original sample, bottom - sample after mixup):**\n![](https://raw.githubusercontent.com/MPGek/mpgek-rfcx-species-audio-detection-public/main/img/Band%20mixup.png)\n**Augmentation that doesn't improve LB:**\n- SpecAugment\n- Sum mixup\n\n## Network topology\n\nI got the best results on EfficientNetB2, B4, B7 (noisy students weights) with a simple FC to 24 classes after the adaptive average pool.\n\nI tried to use different head but all of them gave the same or worse result:\n- the hyper column\n- the hyper column with auxiliary losses after each pooling\n- extra spatial and channel attentions blocks - CBAM\n- dense convolutions\n\n## TP training\nBased on the post where described that every sound file can present unlabeled samples I have to work noisy samples.\n\nI split all samples into confident samples and noisy samples:\n- confident samples - all sigmoid outputs for classes which present in the train_tp.csv file (1 in targets tensor)\n- noisy samples - all sigmoid outputs for classes which not described in the train_tp.csv file (0 in target tensor)\n\nFor the confident samples, I used simple BCE. For the noisy samples, I used LSoft with beta 0.7 (https://arxiv.org/pdf/1901.01189.pdf). \nLSoft:\n```\ndef forward(self, input: torch.Tensor, target: torch.Tensor):\n        with torch.no_grad():\n            pred = torch.sigmoid(input)\n            target_update = self.beta * target + (1 - self.beta) * pred\n        loss = F.binary_cross_entropy_with_logits(input, target_update, reduction=self.reduction)\n        return loss\n```\n\nIn the loss function, I flatten all outputs (even batch dim) to a linear array. And split items into 2 arrays where targets were 1 and where targets were 0.\n\nWith LSoft I got on EfficientNetB7 0.912-0.915 LB.\nWithout LSoft - BCE for all samples I got only about 0.895-0.900.\n\n## FP training\nFor the FP training, I used a dataset with undersampling of the FP samples. Each epoch had all TP samples and the same count of the FP samples.\nI used batch sampling to provide a balanced batch of TP/FP samples - after each TP I added FP with the same species id.\n\nIn the loss function, I calculate loss only for those sigmoid outputs that present in train_tp.csv or train_fp.csv. So all noisy samples are ignored.\n\n## FP co-teaching training\nI have tried to find a way how to use FOCI or SELFIE to work with noisy data, but all of them use historical predictions of each sample. With my random crop and frequency band mixup it's almost impossible. Even shift for 0.5-1 seconds can add a new species to the sample. So historical data will be incorrect.\n\nI tried co-teaching training because it doesn't require historical data.\nPaper: https://arxiv.org/pdf/1804.06872.pdf\nCode sample: https://github.com/bhanML/Co-teaching\n\nWhen I implemented co-teaching training I have only 5 days before the deadline.\nThe first experiments with co-teaching gave 0.830 LB for the TP and 0.880 LB for the FP. So it looked like a bad experiment.\n\nI tried to improve loss function by adding high loss samples with changed targets (by default co-teaching should ignore high loss samples as wrong samples).\n\nThe final loss function consists of:\n- 50% lowest loss samples (all confident samples mandatory added to this part of loss with scale factor 2)\n- 45% ignored losses\n- 5% highest loss samples with the changed target (1 for predictions with sigmoid >= 0.5 and 0 for sigmoid < 0.5)\nThe loss implementation is here: https://github.com/MPGek/mpgek-rfcx-species-audio-detection-public/blob/main/model/forward_passes_coteaching.py\n\nWhen I invented this loss I had only 3 days before the deadline.\nThis loss has a good potential for future experiments. I trained only 2 experiments with this loss, both with the same hyperparameters. The first one had a bug so it produces extremely high metrics and used wrong epochs to submission - however, even with the bug it produces a good LB score comparing to the original FP training.\n\n## Folds and best epoch metric\nTo chose the best epoch in all training I used the BCE which calculated only on confident samples.\nIn some experiments I used 5 stratified KFold, in some, I used 7 folds with stratified shuffle split with test size 0.3.\n\n## Ensembles\nThe TP training with EfficientNetB7 gave me only 0.912-0.915 on the public LB.\nThe FP training with EfficientNetB2-B4 gave only 0.887-0.893 LB.\nThe ensemble of the TP and FP gave 0.929 LB.\n\nThe FP co-teaching training on simple EfficientNetB2 gave me 0.925 LB (a quite good improvement from the original FP with 0.893)\n\n**The final ensemble consists of all best experiments (0.941 public LB and 0.944 private LB):**\n- TP EfficientNetB7 with 0.915\n- FP EfficientNetB2-B4 with 0.893\n- FP co-teaching EfficientNetB2 with 0.925",
    "1208693": "Good job, congrats on results and thanks for sharing solution and code @xmpgek",
    "1209457": "Thanks for sharing.  And congrats on the good result.  Also, thanks for mentioning yet another teacher/student paper I didn't know.  I'll read it for sure.",
    "1219713": "Thanks, I have starred!",
    "1255080": "I want to ask something about your code. In `RainforestModel`, the `self.is_base_resnet` is always set to `False`, why is that?"
  },
  "source": "meta"
}