{
  "id": 220760,
  "title": "2nd place solution",
  "url": "/competitions/rfcx-species-audio-detection/writeups/selim-seferbekov-2nd-place-solution",
  "author_name": "",
  "post_date": "2021-02-19T14:21:45.960Z",
  "votes": 91,
  "comment_count": 40,
  "views": 0,
  "content": "<h2>Overview</h2>\n<p>I trained simple classification models (24 binary classes) with logmel spectrograms :</p>\n<ol>\n<li>bootstrap stage: models are trained on TP/FP with masked BCE loss</li>\n<li>generate soft pseudo labels with 0.5 second sliding window </li>\n<li>train models with pseudo labels and also sample (with p=0.5) places with TP/FP - this partially solves confirmation bias problem. </li>\n</ol>\n<p>Rounds of pseudo labeling and retraining (points 2,3) were repeated until the score on public LB didn't improve. Depending on the settings it took  around 4-10 rounds to converge.</p>\n<p>My initial models that gave 0.86 on TP/FP alone easily reached 0.96x with pseudo labeling . After this success I gave this challenge a 5 weeks break as I lost any motivation to improve my score :)<br>\nLater to my surprise it was extremely hard to beat 0.97 even with improved first stage models.</p>\n<h2>Melspectrogram parameters</h2>\n<ul>\n<li>256 mel bins</li>\n<li>512 hop length</li>\n<li>original SR</li>\n<li>4096 nfft</li>\n</ul>\n<h2>FreqConv (CoordConv for frequency)</h2>\n<p>After my first successful experiment with pseudo labeling that reached 0.969 on public LB  I tried to just swap encoders and blend models but this did not bring any improvements. <br>\nSo I visualised the data for different classes and understood that when working with mel spectrograms for this task we don’t need translation invariance and classes really depend on both frequency and patterns.<br>\nI added a channel to CNN input which contains the number of mel bin scaled to 0-1. This significantly improved validation metrics and after this change log loss on crops around TP after the first round of training with pseudo labels was around 0.04 (same for crops around FP). Though it only slightly improved results on the LB.</p>\n<h2>First stage</h2>\n<p>For the first stage I used all tp/fp information without any sampling and made crops around the center of the signal.</p>\n<p><strong>Augmentations</strong></p>\n<ul>\n<li>time warping</li>\n<li>random frequency masking below TP/FP signal</li>\n<li>random frequency masking above TP/FP signal</li>\n<li>gaussian noise</li>\n<li>volume gain</li>\n<li>mixup on spectrograms</li>\n</ul>\n<p>For mixup on spectrograms - I used constant alpha (0.5) and hard labels with clipping (0,1). Masks were also added. </p>\n<h2>Pseudolabeling stages</h2>\n<p>Sampled TP/FP with p=0.5 otherwise made a random crop from the full spectrogram. </p>\n<p>Without TP/FP sampling labels can become very soft and the score decreases after 2 or 3 rounds.<br>\nAfter training 4 folds of effnet/rexnet I generated OOF labels and ensembled their predictions. Then the training is repeated from scratch.</p>\n<p><strong>Augmentations</strong></p>\n<ul>\n<li>gaussian noise</li>\n<li>volume gain</li>\n<li>mixup </li>\n<li>time warping</li>\n<li>spec augment</li>\n<li>mixup on spectrograms</li>\n</ul>\n<p><strong>Mixup</strong></p>\n<p>I used constant alpha (0.5) and added soft labels from two samples. This hurts logloss on FP a bit but at the same time significantly increases recall on TP.</p>\n<h2>Validation</h2>\n<p>Local validation did not have high correlation with the public leaderboard. Logloss on TP was somehow correlated but still it was not robust. <br>\nSo without proper validation I decided to not select the best checkpoints and just trained 60 epochs (around 200 batches in each epoch) with CosineLR and AdamW optimizer.</p>\n<p>My best models on validation - auc 0.999, log loss 0.03 did not produce great results (0.95). After the competition though It turned out that they can be easily improved with  postprocessing to 97x-98x range.</p>\n<h2>Final ensemble</h2>\n<p>I used 4 models with 4 folds from Effnet and Rexnet (<a href=\"https://arxiv.org/abs/2007.00992\" target=\"_blank\">https://arxiv.org/abs/2007.00992</a>  lightweight models with great performance) families:</p>\n<ul>\n<li>Rexnet-200 (4 sec training/inference), EffnetB3 (4 sec training/inference)</li>\n<li>Rexnet-150 (8 sec training/inference), EffnetB1 (8 sec training/inference)</li>\n</ul>\n<p>Rexnet was much better than EfficientNet alone (less overfitting), but in ensemble they worked great.</p>\n<p>During inference I just used 0.5 second sliding window and took max probabilities for the full clip and then averaged predictions from different models.</p>\n<h2>Lessons learned</h2>\n<p>I did not know about the paper and lacked this useful information about the dataset.</p>\n<p>In my solutions I often rely on models alone but don’t explore the data deeply. <br>\nIn this case I understood that the relabeled train set has similar class distribution to the test set and decided that models would easily learn that. I was wrong and simple post-processing could significantly improve results (though this happened due to severe class imbalance). </p>",
  "messages": [
    {
      "id": "1210468",
      "postDate": "02/19/2021 13:07:16",
      "content": "<h2>Overview</h2>\n<p>I trained simple classification models (24 binary classes) with logmel spectrograms :</p>\n<ol>\n<li>bootstrap stage: models are trained on TP/FP with masked BCE loss</li>\n<li>generate soft pseudo labels with 0.5 second sliding window </li>\n<li>train models with pseudo labels and also sample (with p=0.5) places with TP/FP - this partially solves confirmation bias problem. </li>\n</ol>\n<p>Rounds of pseudo labeling and retraining (points 2,3) were repeated until the score on public LB didn't improve. Depending on the settings it took  around 4-10 rounds to converge.</p>\n<p>My initial models that gave 0.86 on TP/FP alone easily reached 0.96x with pseudo labeling . After this success I gave this challenge a 5 weeks break as I lost any motivation to improve my score :)<br>\nLater to my surprise it was extremely hard to beat 0.97 even with improved first stage models.</p>\n<h2>Melspectrogram parameters</h2>\n<ul>\n<li>256 mel bins</li>\n<li>512 hop length</li>\n<li>original SR</li>\n<li>4096 nfft</li>\n</ul>\n<h2>FreqConv (CoordConv for frequency)</h2>\n<p>After my first successful experiment with pseudo labeling that reached 0.969 on public LB  I tried to just swap encoders and blend models but this did not bring any improvements. <br>\nSo I visualised the data for different classes and understood that when working with mel spectrograms for this task we don’t need translation invariance and classes really depend on both frequency and patterns.<br>\nI added a channel to CNN input which contains the number of mel bin scaled to 0-1. This significantly improved validation metrics and after this change log loss on crops around TP after the first round of training with pseudo labels was around 0.04 (same for crops around FP). Though it only slightly improved results on the LB.</p>\n<h2>First stage</h2>\n<p>For the first stage I used all tp/fp information without any sampling and made crops around the center of the signal.</p>\n<p><strong>Augmentations</strong></p>\n<ul>\n<li>time warping</li>\n<li>random frequency masking below TP/FP signal</li>\n<li>random frequency masking above TP/FP signal</li>\n<li>gaussian noise</li>\n<li>volume gain</li>\n<li>mixup on spectrograms</li>\n</ul>\n<p>For mixup on spectrograms - I used constant alpha (0.5) and hard labels with clipping (0,1). Masks were also added. </p>\n<h2>Pseudolabeling stages</h2>\n<p>Sampled TP/FP with p=0.5 otherwise made a random crop from the full spectrogram. </p>\n<p>Without TP/FP sampling labels can become very soft and the score decreases after 2 or 3 rounds.<br>\nAfter training 4 folds of effnet/rexnet I generated OOF labels and ensembled their predictions. Then the training is repeated from scratch.</p>\n<p><strong>Augmentations</strong></p>\n<ul>\n<li>gaussian noise</li>\n<li>volume gain</li>\n<li>mixup </li>\n<li>time warping</li>\n<li>spec augment</li>\n<li>mixup on spectrograms</li>\n</ul>\n<p><strong>Mixup</strong></p>\n<p>I used constant alpha (0.5) and added soft labels from two samples. This hurts logloss on FP a bit but at the same time significantly increases recall on TP.</p>\n<h2>Validation</h2>\n<p>Local validation did not have high correlation with the public leaderboard. Logloss on TP was somehow correlated but still it was not robust. <br>\nSo without proper validation I decided to not select the best checkpoints and just trained 60 epochs (around 200 batches in each epoch) with CosineLR and AdamW optimizer.</p>\n<p>My best models on validation - auc 0.999, log loss 0.03 did not produce great results (0.95). After the competition though It turned out that they can be easily improved with  postprocessing to 97x-98x range.</p>\n<h2>Final ensemble</h2>\n<p>I used 4 models with 4 folds from Effnet and Rexnet (<a href=\"https://arxiv.org/abs/2007.00992\" target=\"_blank\">https://arxiv.org/abs/2007.00992</a>  lightweight models with great performance) families:</p>\n<ul>\n<li>Rexnet-200 (4 sec training/inference), EffnetB3 (4 sec training/inference)</li>\n<li>Rexnet-150 (8 sec training/inference), EffnetB1 (8 sec training/inference)</li>\n</ul>\n<p>Rexnet was much better than EfficientNet alone (less overfitting), but in ensemble they worked great.</p>\n<p>During inference I just used 0.5 second sliding window and took max probabilities for the full clip and then averaged predictions from different models.</p>\n<h2>Lessons learned</h2>\n<p>I did not know about the paper and lacked this useful information about the dataset.</p>\n<p>In my solutions I often rely on models alone but don’t explore the data deeply. <br>\nIn this case I understood that the relabeled train set has similar class distribution to the test set and decided that models would easily learn that. I was wrong and simple post-processing could significantly improve results (though this happened due to severe class imbalance). </p>",
      "rawMarkdown": "## Overview\n\nI trained simple classification models (24 binary classes) with logmel spectrograms :\n1. bootstrap stage: models are trained on TP/FP with masked BCE loss\n2. generate soft pseudo labels with 0.5 second sliding window \n3. train models with pseudo labels and also sample (with p=0.5) places with TP/FP - this partially solves confirmation bias problem. \n\nRounds of pseudo labeling and retraining (points 2,3) were repeated until the score on public LB didn't improve. Depending on the settings it took  around 4-10 rounds to converge.\n\nMy initial models that gave 0.86 on TP/FP alone easily reached 0.96x with pseudo labeling . After this success I gave this challenge a 5 weeks break as I lost any motivation to improve my score :)\nLater to my surprise it was extremely hard to beat 0.97 even with improved first stage models.\n\n\n## Melspectrogram parameters\n- 256 mel bins\n- 512 hop length\n- original SR\n- 4096 nfft\n\n## FreqConv (CoordConv for frequency)\nAfter my first successful experiment with pseudo labeling that reached 0.969 on public LB  I tried to just swap encoders and blend models but this did not bring any improvements. \nSo I visualised the data for different classes and understood that when working with mel spectrograms for this task we don’t need translation invariance and classes really depend on both frequency and patterns.\nI added a channel to CNN input which contains the number of mel bin scaled to 0-1. This significantly improved validation metrics and after this change log loss on crops around TP after the first round of training with pseudo labels was around 0.04 (same for crops around FP). Though it only slightly improved results on the LB.\n\n\n## First stage \n\nFor the first stage I used all tp/fp information without any sampling and made crops around the center of the signal.\n\n**Augmentations**\n- time warping\n- random frequency masking below TP/FP signal\n- random frequency masking above TP/FP signal\n- gaussian noise\n- volume gain\n- mixup on spectrograms\n\nFor mixup on spectrograms - I used constant alpha (0.5) and hard labels with clipping (0,1). Masks were also added. \n\n## Pseudolabeling stages\nSampled TP/FP with p=0.5 otherwise made a random crop from the full spectrogram. \n\nWithout TP/FP sampling labels can become very soft and the score decreases after 2 or 3 rounds.\nAfter training 4 folds of effnet/rexnet I generated OOF labels and ensembled their predictions. Then the training is repeated from scratch.\n\n**Augmentations**\n- gaussian noise\n- volume gain\n- mixup \n- time warping\n- spec augment\n- mixup on spectrograms\n\n**Mixup**\n\nI used constant alpha (0.5) and added soft labels from two samples. This hurts logloss on FP a bit but at the same time significantly increases recall on TP.\n\n## Validation\n\t\nLocal validation did not have high correlation with the public leaderboard. Logloss on TP was somehow correlated but still it was not robust. \nSo without proper validation I decided to not select the best checkpoints and just trained 60 epochs (around 200 batches in each epoch) with CosineLR and AdamW optimizer.\n\nMy best models on validation - auc 0.999, log loss 0.03 did not produce great results (0.95). After the competition though It turned out that they can be easily improved with  postprocessing to 97x-98x range.\n\n\n## Final ensemble\n\nI used 4 models with 4 folds from Effnet and Rexnet (https://arxiv.org/abs/2007.00992  lightweight models with great performance) families:\n- Rexnet-200 (4 sec training/inference), EffnetB3 (4 sec training/inference)\n- Rexnet-150 (8 sec training/inference), EffnetB1 (8 sec training/inference)\n\nRexnet was much better than EfficientNet alone (less overfitting), but in ensemble they worked great.\n\nDuring inference I just used 0.5 second sliding window and took max probabilities for the full clip and then averaged predictions from different models.\n\n## Lessons learned\nI did not know about the paper and lacked this useful information about the dataset.\n\nIn my solutions I often rely on models alone but don’t explore the data deeply. \nIn this case I understood that the relabeled train set has similar class distribution to the test set and decided that models would easily learn that. I was wrong and simple post-processing could significantly improve results (though this happened due to severe class imbalance).",
      "votes": null
    },
    {
      "id": "1210541",
      "postDate": "02/19/2021 14:02:48",
      "content": "<p>Rexnet-200 for 60 epochs.. wow) </p>",
      "rawMarkdown": "Rexnet-200 for 60 epochs.. wow)",
      "votes": null
    },
    {
      "id": "1210546",
      "postDate": "02/19/2021 14:04:30",
      "content": "<p>Thanks for sharing.  and congrats on the solo prize!</p>\n<p>You were first in finding that frequency masking per species and adding frequency location info was key.  And you made pseudo labeling work (not like nme). You deserve this great resuit.</p>",
      "rawMarkdown": "Thanks for sharing.  and congrats on the solo prize!\n\nYou were first in finding that frequency masking per species and adding frequency location info was key.  And you made pseudo labeling work (not like nme). You deserve this great resuit.",
      "votes": null
    },
    {
      "id": "1210553",
      "postDate": "02/19/2021 14:08:45",
      "content": "<p><a href=\"https://arxiv.org/abs/2007.00992\" target=\"_blank\">https://arxiv.org/abs/2007.00992</a> Rexnet-200 (200 = 2.0 scale, 150 = 1.5)  is a lightweight model. It is not related to resnet, it is a modification of MobileNet that performs like EffNet. <br>\nSo it was around 1 minute per epoch </p>",
      "rawMarkdown": "https://arxiv.org/abs/2007.00992 Rexnet-200 (200 = 2.0 scale, 150 = 1.5)  is a lightweight model. It is not related to resnet, it is a modification of MobileNet that performs like EffNet. \nSo it was around 1 minute per epoch",
      "votes": null
    },
    {
      "id": "1210560",
      "postDate": "02/19/2021 14:12:02",
      "content": "<p>Thanks! Actually pseudo-labeling did not work that well for SED models (tried framewise/clipwise pseudo labels) and I switched to standard image classification</p>",
      "rawMarkdown": "Thanks! Actually pseudo-labeling did not work that well for SED models (tried framewise/clipwise pseudo labels) and I switched to standard image classification",
      "votes": null
    },
    {
      "id": "1210562",
      "postDate": "02/19/2021 14:13:44",
      "content": "<p>Great work! Congrats on solo 2nd place and thanks for the writeup <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a> </p>",
      "rawMarkdown": "Great work! Congrats on solo 2nd place and thanks for the writeup @selimsef",
      "votes": null
    },
    {
      "id": "1210571",
      "postDate": "02/19/2021 14:24:22",
      "content": "<p>I never used SED.</p>\n<p>But I think I now get why my pseudo labeling did not work.  Running experiemnt to check.</p>",
      "rawMarkdown": "I never used SED.\n\nBut I think I now get why my pseudo labeling did not work.  Running experiemnt to check.",
      "votes": null
    },
    {
      "id": "1210574",
      "postDate": "02/19/2021 14:25:36",
      "content": "<p>Thanks for sharing and congrats on getting the global 12th rank!</p>",
      "rawMarkdown": "Thanks for sharing and congrats on getting the global 12th rank!",
      "votes": null
    },
    {
      "id": "1210577",
      "postDate": "02/19/2021 14:29:13",
      "content": "<blockquote>\n  <p>Sampled TP/FP with p=0.5 otherwise made a random crop from the full spectrogram.</p>\n</blockquote>\n<p>a question on this, do you train on the same random sampled audio crop and pseudo labels after or you re-sample tp/fp or made a random crop again in stage 2?</p>",
      "rawMarkdown": "> Sampled TP/FP with p=0.5 otherwise made a random crop from the full spectrogram.\n\na question on this, do you train on the same random sampled audio crop and pseudo labels after or you re-sample tp/fp or made a random crop again in stage 2?",
      "votes": null
    },
    {
      "id": "1210587",
      "postDate": "02/19/2021 14:35:29",
      "content": "<p>Thanks! I was exactly there some time ago😄Looks like I need to participate in a few competitions in a row (it is not possible to get high global rank otherwise due to Kaggle points decay)</p>",
      "rawMarkdown": "Thanks! I was exactly there some time ago😄Looks like I need to participate in a few competitions in a row (it is not possible to get high global rank otherwise due to Kaggle points decay)",
      "votes": null
    },
    {
      "id": "1210592",
      "postDate": "02/19/2021 14:41:59",
      "content": "<p>Thanks for sharing, great work!</p>\n<p>Did you try to add freqs to the loss function (it could force the network to use the FreqConv)? <br>\nI used CoorConv in some tasks and without adding the same coordinates to the loss function it didn't improve the network.<br>\nI have a simple check for it - pass 0 or noise to the channel with coordinates, when I didn't add coordinates to the loss function, it performs the same as with proper data in the CoordConv channel. When I trained with coordinates in the loss, 0 or noise in the CoordConv - network inferences much worse.</p>",
      "rawMarkdown": "Thanks for sharing, great work!\n\nDid you try to add freqs to the loss function (it could force the network to use the FreqConv)? \nI used CoorConv in some tasks and without adding the same coordinates to the loss function it didn't improve the network.\nI have a simple check for it - pass 0 or noise to the channel with coordinates, when I didn't add coordinates to the loss function, it performs the same as with proper data in the CoordConv channel. When I trained with coordinates in the loss, 0 or noise in the CoordConv - network inferences much worse.",
      "votes": null
    },
    {
      "id": "1210593",
      "postDate": "02/19/2021 14:42:04",
      "content": "<blockquote>\n  <p>a question on this, do you train on the same random sampled audio crop and pseudo labels after or you re-sample tp/fp or made a random crop again in stage 2?</p>\n</blockquote>\n<p>My pseudolabels (for 4 second models) had 113 frames per audio clip (0.5 sec sliding window). <br>\nIf TP/FP is sampled - I made a random crop around the center and then found nearest frame from pseudolabels. Pseudo labels were fixed using tp/fp. <br>\nOtherwise I just took random frame from 113 and postprocessed labels if they overlap with tp/fp data.</p>",
      "rawMarkdown": ">  a question on this, do you train on the same random sampled audio crop and pseudo labels after or you re-sample tp/fp or made a random crop again in stage 2?\n\nMy pseudolabels (for 4 second models) had 113 frames per audio clip (0.5 sec sliding window). \nIf TP/FP is sampled - I made a random crop around the center and then found nearest frame from pseudolabels. Pseudo labels were fixed using tp/fp. \nOtherwise I just took random frame from 113 and postprocessed labels if they overlap with tp/fp data.",
      "votes": null
    },
    {
      "id": "1210600",
      "postDate": "02/19/2021 14:47:59",
      "content": "<p>interesting, I assume that the distance from the <code>random crop around the center</code> to the nearest frame from pseudo labels is so close that you can ignore the distance and just assign the same labels, I tried just using the same crops without random cropping again. I think the key probably was that you trained a fixed 60 epochs 😂 I couldn't get it to work properly when I didn't have extra hand labels yet. </p>",
      "rawMarkdown": "interesting, I assume that the distance from the `random crop around the center` to the nearest frame from pseudo labels is so close that you can ignore the distance and just assign the same labels, I tried just using the same crops without random cropping again. I think the key probably was that you trained a fixed 60 epochs 😂 I couldn't get it to work properly when I didn't have extra hand labels yet.",
      "votes": null
    },
    {
      "id": "1210604",
      "postDate": "02/19/2021 14:50:50",
      "content": "<p>Good point! Just checked some checkpoint</p>\n<pre><code>With proper frequencies\nneg_logloss: 0.1588028629548308\npos_logloss: 0.05171160377776966\n\nWith zeros\nneg_logloss: 0.19649322897329777\npos_logloss: 0.3455986050609499\n</code></pre>\n<p>So my models really use it somehow</p>",
      "rawMarkdown": "Good point! Just checked some checkpoint\n```\nWith proper frequencies\nneg_logloss: 0.1588028629548308\npos_logloss: 0.05171160377776966\n\nWith zeros\nneg_logloss: 0.19649322897329777\npos_logloss: 0.3455986050609499\n\n```\nSo my models really use it somehow",
      "votes": null
    },
    {
      "id": "1210608",
      "postDate": "02/19/2021 14:53:50",
      "content": "<p>max distance = 0.25 seconds, so it should not have big impact on model's quality. More frames would be better but it would take much more time to generate pseudolabels.</p>",
      "rawMarkdown": "max distance = 0.25 seconds, so it should not have big impact on model's quality. More frames would be better but it would take much more time to generate pseudolabels.",
      "votes": null
    },
    {
      "id": "1210610",
      "postDate": "02/19/2021 14:55:00",
      "content": "<p>Indeed, looks like the network managed to use it without loss modifications.</p>",
      "rawMarkdown": "Indeed, looks like the network managed to use it without loss modifications.",
      "votes": null
    },
    {
      "id": "1210612",
      "postDate": "02/19/2021 14:57:53",
      "content": "<p>Thanks for sharing your work. i am curious to learn new thing from your work<br>\ni have a question on masked BCE loss. how is this implemented. when you say mask, are you masking the loss for other classes that are not in the label or masking the time frames(say 4sec input where the label is only for sec1-2) where there is no label.  </p>\n<p>second question . how are your using both TP and FP. how this loss fn will look like.</p>",
      "rawMarkdown": "Thanks for sharing your work. i am curious to learn new thing from your work\ni have a question on masked BCE loss. how is this implemented. when you say mask, are you masking the loss for other classes that are not in the label or masking the time frames(say 4sec input where the label is only for sec1-2) where there is no label.  \n\nsecond question . how are your using both TP and FP. how this loss fn will look like.",
      "votes": null
    },
    {
      "id": "1210629",
      "postDate": "02/19/2021 15:14:09",
      "content": "<p>Masks for loss function have the same shape as labels, 1 for the classes we know (TP/FP) 0 for others</p>\n<pre><code>class BCEMasked(nn.Module):\n   def forward(self, inputs, targets, mask=None):\n       bce_loss = binary_cross_entropy_with_logits(inputs, targets, reduction='none')\n       if mask is not None:\n           bce_loss = bce_loss[mask &gt; 0]\n       return bce_loss.mean()\n</code></pre>",
      "rawMarkdown": "Masks for loss function have the same shape as labels, 1 for the classes we know (TP/FP) 0 for others\n```\nclass BCEMasked(nn.Module):\n    def forward(self, inputs, targets, mask=None):\n        bce_loss = binary_cross_entropy_with_logits(inputs, targets, reduction='none')\n        if mask is not None:\n            bce_loss = bce_loss[mask > 0]\n        return bce_loss.mean()\n```",
      "votes": null
    },
    {
      "id": "1210652",
      "postDate": "02/19/2021 15:30:16",
      "content": "<p>thanks for answering . about TP/FP , how are you using FP for training the model. with sigmoid you cant have FP = -1. </p>",
      "rawMarkdown": "thanks for answering . about TP/FP , how are you using FP for training the model. with sigmoid you cant have FP = -1.",
      "votes": null
    },
    {
      "id": "1210662",
      "postDate": "02/19/2021 15:35:22",
      "content": "<p>example - FP for s0<br>\nmask=[1, 0, 0, …,0], targets=[0, 0, 0, …, 0] - only the first element from the targets and output will be considered by the loss function</p>",
      "rawMarkdown": "example - FP for s0\nmask=[1, 0, 0, ...,0], targets=[0, 0, 0, ..., 0] - only the first element from the targets and output will be considered by the loss function",
      "votes": null
    },
    {
      "id": "1210665",
      "postDate": "02/19/2021 15:37:44",
      "content": "<p>ahh. now i understand . … thanks for explaining .</p>",
      "rawMarkdown": "ahh. now i understand . ... thanks for explaining .",
      "votes": null
    },
    {
      "id": "1211187",
      "postDate": "02/20/2021 03:31:29",
      "content": "<p>\"why my pseudo labeling did not work. Running experiemnt to check.\"</p>\n<p>one simple way to check pseudo label is:</p>\n<ul>\n<li>make a mean image of the pseudo label cropped melspec and see if it looks familiar to the train tp (or fp)</li>\n<li>use tsne to show your train tp,fp and pseudo label </li>\n<li>measure purity based on the validation set</li>\n</ul>\n<p>for image application, I usually do a manual inspection to estimate purity. e.g. test random sample of 100 pseudo label samples</p>",
      "rawMarkdown": "\"why my pseudo labeling did not work. Running experiemnt to check.\"\n\none simple way to check pseudo label is:\n- make a mean image of the pseudo label cropped melspec and see if it looks familiar to the train tp (or fp)\n- use tsne to show your train tp,fp and pseudo label \n- measure purity based on the validation set\n\nfor image application, I usually do a manual inspection to estimate purity. e.g. test random sample of 100 pseudo label samples",
      "votes": null
    },
    {
      "id": "1211189",
      "postDate": "02/20/2021 03:36:50",
      "content": "<p><a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a> <br>\nthanks for the write-up and congrats on your winning.</p>\n<p>So it confirms that 0.960~0.970 is the limit with the naive pesudo label.</p>",
      "rawMarkdown": "selimsef \nthanks for the write-up and congrats on your winning.\n\nSo it confirms that 0.960~0.970 is the limit with the naive pesudo label.",
      "votes": null
    },
    {
      "id": "1211535",
      "postDate": "02/20/2021 10:06:11",
      "content": "<p>Actually I tried different  SSL approaches, pseudo labeling with mining confident labels and loss masking was much worse than soft labels.  Postprocessing could fix that I guess. The main reason for than is comp metric. Something like F1 per class would better assess predictive performance and  would not be affected by class imbalance.</p>",
      "rawMarkdown": "Actually I tried different  SSL approaches, pseudo labeling with mining confident labels and loss masking was much worse than soft labels.  Postprocessing could fix that I guess. The main reason for than is comp metric. Something like F1 per class would better assess predictive performance and  would not be affected by class imbalance.",
      "votes": null
    },
    {
      "id": "1211717",
      "postDate": "02/20/2021 13:28:20",
      "content": "<p>Selim, thank you for the write up. I love how elegant/simple your solution is.</p>",
      "rawMarkdown": "Selim, thank you for the write up. I love how elegant/simple your solution is.",
      "votes": null
    },
    {
      "id": "1211802",
      "postDate": "02/20/2021 15:14:13",
      "content": "<p>Thanks for sharing!<br>\nSeems that the score benefits most from training 24 models separately right? I tried training 24 models and some has reached 99.9% accuracy locally(some just around 60% 80%) but still got no-good public leaderboard score and I just suspect that there was something wrong in my models validation procedure.</p>",
      "rawMarkdown": "Thanks for sharing!\nSeems that the score benefits most from training 24 models separately right? I tried training 24 models and some has reached 99.9% accuracy locally(some just around 60% 80%) but still got no-good public leaderboard score and I just suspect that there was something wrong in my models validation procedure.",
      "votes": null
    },
    {
      "id": "1211843",
      "postDate": "02/20/2021 15:54:21",
      "content": "<p>No, single 4 fold model scores 0.97 on public LB. Single checkpoint something like 0.968 (Public). Where did you see 24 models?) </p>",
      "rawMarkdown": "No, single 4 fold model scores 0.97 on public LB. Single checkpoint something like 0.968 (Public). Where did you see 24 models?)",
      "votes": null
    },
    {
      "id": "1211851",
      "postDate": "02/20/2021 16:05:00",
      "content": "<p>i tried to include the masking for loss fn. i am facing issue with having many predictions set to 1.0 and few classes are having 0.99 in the test dataset. i am still using my SED model to try this implementation. do you see any issue. i am using 6sec clip around TP.</p>\n<p><code>1.0\n0.99999994\n0.90000004\n1.0\n1.0\n1.0\n1.0</code></p>",
      "rawMarkdown": "i tried to include the masking for loss fn. i am facing issue with having many predictions set to 1.0 and few classes are having 0.99 in the test dataset. i am still using my SED model to try this implementation. do you see any issue. i am using 6sec clip around TP.\n\n`1.0\n0.99999994\n0.90000004\n1.0\n1.0\n1.0\n1.0`",
      "votes": null
    },
    {
      "id": "1211932",
      "postDate": "02/20/2021 17:13:16",
      "content": "<p>you are using only TP samples, of course if you have all targets = 1 NN will predict 1.0</p>",
      "rawMarkdown": "you are using only TP samples, of course if you have all targets = 1 NN will predict 1.0",
      "votes": null
    },
    {
      "id": "1211958",
      "postDate": "02/20/2021 17:42:09",
      "content": "<p>i am doing this in the following way</p>\n<pre><code>prediction = model output\ntarget = true target\nloss = loss(prediction, target, mask = target)\n</code></pre>\n<p>here target will act as mask as it has 1 for the class present and rest are 0</p>",
      "rawMarkdown": "i am doing this in the following way\n```\nprediction = model output\ntarget = true target\nloss = loss(prediction, target, mask = target)\n```\n\nhere target will act as mask as it has 1 for the class present and rest are 0",
      "votes": null
    },
    {
      "id": "1212213",
      "postDate": "02/21/2021 02:36:40",
      "content": "<p>I agree that frequency location is the key to this competition according to this <img src=\"https://i.imgur.com/D3g3vJc.png\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> your solution is a perfect way to capture such information.</p>",
      "rawMarkdown": "I agree that frequency location is the key to this competition according to this ![](https://i.imgur.com/D3g3vJc.png)\n\n@cpmpml your solution is a perfect way to capture such information.",
      "votes": null
    },
    {
      "id": "1212214",
      "postDate": "02/21/2021 02:38:04",
      "content": "<p>Oh.. Thanks so top3 are training on a single model, I was just reading too fast and misunderstood that, sorry for that</p>",
      "rawMarkdown": "Oh.. Thanks so top3 are training on a single model, I was just reading too fast and misunderstood that, sorry for that",
      "votes": null
    },
    {
      "id": "1212802",
      "postDate": "02/21/2021 16:02:34",
      "content": "<p>Thanks for the writeup and congratulations on your winning. I thought that you would not make any other submission after that 0.97 😄</p>\n<p>I wonder that what would be your final score with the mentioned postprocessing methods like Chris's method <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Thanks for the writeup and congratulations on your winning. I thought that you would not make any other submission after that 0.97 😄\n\nI wonder that what would be your final score with the mentioned postprocessing methods like Chris's method [here](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389)",
      "votes": null
    },
    {
      "id": "1212840",
      "postDate": "02/21/2021 16:50:19",
      "content": "<p>982-983 depending on submission</p>",
      "rawMarkdown": "982-983 depending on submission",
      "votes": null
    },
    {
      "id": "1212889",
      "postDate": "02/21/2021 17:38:21",
      "content": "<p>Woow, it's great. Thanks 👍</p>\n<p>Did you have any idea that you wanted to try but didn't?</p>",
      "rawMarkdown": "Woow, it's great. Thanks 👍\n\nDid you have any idea that you wanted to try but didn't?",
      "votes": null
    },
    {
      "id": "1219843",
      "postDate": "02/27/2021 09:43:31",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/xmpgek\" target=\"_blank\">@xmpgek</a> , could you talk more details about how to add the coordinates to the loss function?</p>",
      "rawMarkdown": "Hi @xmpgek , could you talk more details about how to add the coordinates to the loss function?",
      "votes": null
    },
    {
      "id": "1226031",
      "postDate": "03/04/2021 07:03:22",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a>, congratulations and thank you very much for the write up. I would like to ask for models that you trained with FreqConv, did you use pretrained models? If yes, how did you modify the initial layer of the pretrained models to accommodate one more input channel?  Thank you very much.</p>",
      "rawMarkdown": "Hi @selimsef, congratulations and thank you very much for the write up. I would like to ask for models that you trained with FreqConv, did you use pretrained models? If yes, how did you modify the initial layer of the pretrained models to accommodate one more input channel?  Thank you very much.",
      "votes": null
    },
    {
      "id": "1227081",
      "postDate": "03/05/2021 07:13:33",
      "content": "<p>With one input channel instead of RGB it is quite easy to use pretrain, just sum weights from 3 channels. With 2 channels it’s a bit different. Timm library supports that out of the box, just pass in_chans=2.</p>",
      "rawMarkdown": "With one input channel instead of RGB it is quite easy to use pretrain, just sum weights from 3 channels. With 2 channels it’s a bit different. Timm library supports that out of the box, just pass in_chans=2.",
      "votes": null
    },
    {
      "id": "1286808",
      "postDate": "04/28/2021 12:25:33",
      "content": "<p>Hello, I have just re-read your post; it just looks like the first sentence may be a bit misleading to me, though I am not sure if it is my own problem. \"I trained simple classification models (24 binary classes) with logmel spectrograms :\" looks to me that it is 24 binary classification model? May you clarify on this? Thanks.</p>",
      "rawMarkdown": "Hello, I have just re-read your post; it just looks like the first sentence may be a bit misleading to me, though I am not sure if it is my own problem. \"I trained simple classification models (24 binary classes) with logmel spectrograms :\" looks to me that it is 24 binary classification model? May you clarify on this? Thanks.",
      "votes": null
    },
    {
      "id": "1335050",
      "postDate": "06/04/2021 01:44:07",
      "content": "<p>Congratulations team!<br>\nCan you provide a link to github？</p>",
      "rawMarkdown": "Congratulations team!\nCan you provide a link to github？",
      "votes": null
    },
    {
      "id": "1358732",
      "postDate": "06/20/2021 17:31:05",
      "content": "<p>Thanks for the excellent write-up and congrats for the 2nd place! <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a> <br>\nAnd I hope I am not too late to ask a few questions to further understand your solution. I believe this could help me reproduce the solution:</p>\n<ol>\n<li>What was the cropped duration that you used for training?</li>\n<li>How did mixup work together with masked BCE loss? coz I can see there are some nuances here. (e.g. how did u define the target when u mix TP with FP? which labels did u mask for target that is mixup-ed?)<br>\n3a. in pseudo labeling stage, u mentioned you have special sampling strategy (i.e. sample TP/FP at p=0.5, otherwise randomly crop from a full clip) to train models that are responsible for generating pseudo labels, because <code>Without TP/FP sampling labels can become very soft and the score decreases after 2 or 3 rounds</code>. Could you elaborate more on this issue? <br>\n3b. Follow up on 3a. how did u define the target for sample that is randomly cropped from a full clip? Additionally, was this sampling strategy only applied in first round? (i.e. the round without any pseudo labels in training set). </li>\n<li>When you train your model with pseudo labels in your training set, did you handle pseudo labels specially? (e.g. apply penalty proportional to prediction confidence in the masked BCE loss). Also, how did u sample pseudo labels among TP/ FP? </li>\n</ol>\n<p>Look forward to your reply! </p>",
      "rawMarkdown": "Thanks for the excellent write-up and congrats for the 2nd place! @selimsef \nAnd I hope I am not too late to ask a few questions to further understand your solution. I believe this could help me reproduce the solution:\n1. What was the cropped duration that you used for training?\n2. How did mixup work together with masked BCE loss? coz I can see there are some nuances here. (e.g. how did u define the target when u mix TP with FP? which labels did u mask for target that is mixup-ed?)\n3a. in pseudo labeling stage, u mentioned you have special sampling strategy (i.e. sample TP/FP at p=0.5, otherwise randomly crop from a full clip) to train models that are responsible for generating pseudo labels, because `Without TP/FP sampling labels can become very soft and the score decreases after 2 or 3 rounds`. Could you elaborate more on this issue? \n3b. Follow up on 3a. how did u define the target for sample that is randomly cropped from a full clip? Additionally, was this sampling strategy only applied in first round? (i.e. the round without any pseudo labels in training set). \n4. When you train your model with pseudo labels in your training set, did you handle pseudo labels specially? (e.g. apply penalty proportional to prediction confidence in the masked BCE loss). Also, how did u sample pseudo labels among TP/ FP? \n\nLook forward to your reply!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1210541,
      "author_name": "kupchanski",
      "author_url": "",
      "post_date": "02/19/2021 14:02:48",
      "content": "<p>Rexnet-200 for 60 epochs.. wow) </p>",
      "votes": null,
      "replies": [
        {
          "id": 1210553,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "02/19/2021 14:08:45",
          "content": "<p><a href=\"https://arxiv.org/abs/2007.00992\" target=\"_blank\">https://arxiv.org/abs/2007.00992</a> Rexnet-200 (200 = 2.0 scale, 150 = 1.5)  is a lightweight model. It is not related to resnet, it is a modification of MobileNet that performs like EffNet. <br>\nSo it was around 1 minute per epoch </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1210546,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/19/2021 14:04:30",
      "content": "<p>Thanks for sharing.  and congrats on the solo prize!</p>\n<p>You were first in finding that frequency masking per species and adding frequency location info was key.  And you made pseudo labeling work (not like nme). You deserve this great resuit.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210560,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "02/19/2021 14:12:02",
          "content": "<p>Thanks! Actually pseudo-labeling did not work that well for SED models (tried framewise/clipwise pseudo labels) and I switched to standard image classification</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210571,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/19/2021 14:24:22",
          "content": "<p>I never used SED.</p>\n<p>But I think I now get why my pseudo labeling did not work.  Running experiemnt to check.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1211187,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/20/2021 03:31:29",
          "content": "<p>\"why my pseudo labeling did not work. Running experiemnt to check.\"</p>\n<p>one simple way to check pseudo label is:</p>\n<ul>\n<li>make a mean image of the pseudo label cropped melspec and see if it looks familiar to the train tp (or fp)</li>\n<li>use tsne to show your train tp,fp and pseudo label </li>\n<li>measure purity based on the validation set</li>\n</ul>\n<p>for image application, I usually do a manual inspection to estimate purity. e.g. test random sample of 100 pseudo label samples</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1212213,
          "author_name": "jihangz",
          "author_url": "",
          "post_date": "02/21/2021 02:36:40",
          "content": "<p>I agree that frequency location is the key to this competition according to this <img src=\"https://i.imgur.com/D3g3vJc.png\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> your solution is a perfect way to capture such information.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1210562,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "02/19/2021 14:13:44",
      "content": "<p>Great work! Congrats on solo 2nd place and thanks for the writeup <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1210574,
      "author_name": "dicksonchin93",
      "author_url": "",
      "post_date": "02/19/2021 14:25:36",
      "content": "<p>Thanks for sharing and congrats on getting the global 12th rank!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210577,
          "author_name": "dicksonchin93",
          "author_url": "",
          "post_date": "02/19/2021 14:29:13",
          "content": "<blockquote>\n  <p>Sampled TP/FP with p=0.5 otherwise made a random crop from the full spectrogram.</p>\n</blockquote>\n<p>a question on this, do you train on the same random sampled audio crop and pseudo labels after or you re-sample tp/fp or made a random crop again in stage 2?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210587,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "02/19/2021 14:35:29",
          "content": "<p>Thanks! I was exactly there some time ago😄Looks like I need to participate in a few competitions in a row (it is not possible to get high global rank otherwise due to Kaggle points decay)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210593,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "02/19/2021 14:42:04",
          "content": "<blockquote>\n  <p>a question on this, do you train on the same random sampled audio crop and pseudo labels after or you re-sample tp/fp or made a random crop again in stage 2?</p>\n</blockquote>\n<p>My pseudolabels (for 4 second models) had 113 frames per audio clip (0.5 sec sliding window). <br>\nIf TP/FP is sampled - I made a random crop around the center and then found nearest frame from pseudolabels. Pseudo labels were fixed using tp/fp. <br>\nOtherwise I just took random frame from 113 and postprocessed labels if they overlap with tp/fp data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210600,
          "author_name": "dicksonchin93",
          "author_url": "",
          "post_date": "02/19/2021 14:47:59",
          "content": "<p>interesting, I assume that the distance from the <code>random crop around the center</code> to the nearest frame from pseudo labels is so close that you can ignore the distance and just assign the same labels, I tried just using the same crops without random cropping again. I think the key probably was that you trained a fixed 60 epochs 😂 I couldn't get it to work properly when I didn't have extra hand labels yet. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210608,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "02/19/2021 14:53:50",
          "content": "<p>max distance = 0.25 seconds, so it should not have big impact on model's quality. More frames would be better but it would take much more time to generate pseudolabels.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1210592,
      "author_name": "xmpgek",
      "author_url": "",
      "post_date": "02/19/2021 14:41:59",
      "content": "<p>Thanks for sharing, great work!</p>\n<p>Did you try to add freqs to the loss function (it could force the network to use the FreqConv)? <br>\nI used CoorConv in some tasks and without adding the same coordinates to the loss function it didn't improve the network.<br>\nI have a simple check for it - pass 0 or noise to the channel with coordinates, when I didn't add coordinates to the loss function, it performs the same as with proper data in the CoordConv channel. When I trained with coordinates in the loss, 0 or noise in the CoordConv - network inferences much worse.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210604,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "02/19/2021 14:50:50",
          "content": "<p>Good point! Just checked some checkpoint</p>\n<pre><code>With proper frequencies\nneg_logloss: 0.1588028629548308\npos_logloss: 0.05171160377776966\n\nWith zeros\nneg_logloss: 0.19649322897329777\npos_logloss: 0.3455986050609499\n</code></pre>\n<p>So my models really use it somehow</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210610,
          "author_name": "xmpgek",
          "author_url": "",
          "post_date": "02/19/2021 14:55:00",
          "content": "<p>Indeed, looks like the network managed to use it without loss modifications.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1219843,
          "author_name": "karlyukang",
          "author_url": "",
          "post_date": "02/27/2021 09:43:31",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/xmpgek\" target=\"_blank\">@xmpgek</a> , could you talk more details about how to add the coordinates to the loss function?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1210612,
      "author_name": "yuvaramsingh",
      "author_url": "",
      "post_date": "02/19/2021 14:57:53",
      "content": "<p>Thanks for sharing your work. i am curious to learn new thing from your work<br>\ni have a question on masked BCE loss. how is this implemented. when you say mask, are you masking the loss for other classes that are not in the label or masking the time frames(say 4sec input where the label is only for sec1-2) where there is no label.  </p>\n<p>second question . how are your using both TP and FP. how this loss fn will look like.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210629,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "02/19/2021 15:14:09",
          "content": "<p>Masks for loss function have the same shape as labels, 1 for the classes we know (TP/FP) 0 for others</p>\n<pre><code>class BCEMasked(nn.Module):\n   def forward(self, inputs, targets, mask=None):\n       bce_loss = binary_cross_entropy_with_logits(inputs, targets, reduction='none')\n       if mask is not None:\n           bce_loss = bce_loss[mask &gt; 0]\n       return bce_loss.mean()\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210652,
          "author_name": "yuvaramsingh",
          "author_url": "",
          "post_date": "02/19/2021 15:30:16",
          "content": "<p>thanks for answering . about TP/FP , how are you using FP for training the model. with sigmoid you cant have FP = -1. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210662,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "02/19/2021 15:35:22",
          "content": "<p>example - FP for s0<br>\nmask=[1, 0, 0, …,0], targets=[0, 0, 0, …, 0] - only the first element from the targets and output will be considered by the loss function</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210665,
          "author_name": "yuvaramsingh",
          "author_url": "",
          "post_date": "02/19/2021 15:37:44",
          "content": "<p>ahh. now i understand . … thanks for explaining .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1211851,
          "author_name": "yuvaramsingh",
          "author_url": "",
          "post_date": "02/20/2021 16:05:00",
          "content": "<p>i tried to include the masking for loss fn. i am facing issue with having many predictions set to 1.0 and few classes are having 0.99 in the test dataset. i am still using my SED model to try this implementation. do you see any issue. i am using 6sec clip around TP.</p>\n<p><code>1.0\n0.99999994\n0.90000004\n1.0\n1.0\n1.0\n1.0</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1211932,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "02/20/2021 17:13:16",
          "content": "<p>you are using only TP samples, of course if you have all targets = 1 NN will predict 1.0</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1211958,
          "author_name": "yuvaramsingh",
          "author_url": "",
          "post_date": "02/20/2021 17:42:09",
          "content": "<p>i am doing this in the following way</p>\n<pre><code>prediction = model output\ntarget = true target\nloss = loss(prediction, target, mask = target)\n</code></pre>\n<p>here target will act as mask as it has 1 for the class present and rest are 0</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1211189,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/20/2021 03:36:50",
      "content": "<p><a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a> <br>\nthanks for the write-up and congrats on your winning.</p>\n<p>So it confirms that 0.960~0.970 is the limit with the naive pesudo label.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1211535,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "02/20/2021 10:06:11",
          "content": "<p>Actually I tried different  SSL approaches, pseudo labeling with mining confident labels and loss masking was much worse than soft labels.  Postprocessing could fix that I guess. The main reason for than is comp metric. Something like F1 per class would better assess predictive performance and  would not be affected by class imbalance.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1211717,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "02/20/2021 13:28:20",
      "content": "<p>Selim, thank you for the write up. I love how elegant/simple your solution is.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1211802,
      "author_name": "wubinbai",
      "author_url": "",
      "post_date": "02/20/2021 15:14:13",
      "content": "<p>Thanks for sharing!<br>\nSeems that the score benefits most from training 24 models separately right? I tried training 24 models and some has reached 99.9% accuracy locally(some just around 60% 80%) but still got no-good public leaderboard score and I just suspect that there was something wrong in my models validation procedure.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1211843,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "02/20/2021 15:54:21",
          "content": "<p>No, single 4 fold model scores 0.97 on public LB. Single checkpoint something like 0.968 (Public). Where did you see 24 models?) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1212214,
          "author_name": "wubinbai",
          "author_url": "",
          "post_date": "02/21/2021 02:38:04",
          "content": "<p>Oh.. Thanks so top3 are training on a single model, I was just reading too fast and misunderstood that, sorry for that</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1286808,
          "author_name": "wubinbai",
          "author_url": "",
          "post_date": "04/28/2021 12:25:33",
          "content": "<p>Hello, I have just re-read your post; it just looks like the first sentence may be a bit misleading to me, though I am not sure if it is my own problem. \"I trained simple classification models (24 binary classes) with logmel spectrograms :\" looks to me that it is 24 binary classification model? May you clarify on this? Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1212802,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "02/21/2021 16:02:34",
      "content": "<p>Thanks for the writeup and congratulations on your winning. I thought that you would not make any other submission after that 0.97 😄</p>\n<p>I wonder that what would be your final score with the mentioned postprocessing methods like Chris's method <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389\" target=\"_blank\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1212840,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "02/21/2021 16:50:19",
          "content": "<p>982-983 depending on submission</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1212889,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "02/21/2021 17:38:21",
          "content": "<p>Woow, it's great. Thanks 👍</p>\n<p>Did you have any idea that you wanted to try but didn't?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1226031,
      "author_name": "thomeou",
      "author_url": "",
      "post_date": "03/04/2021 07:03:22",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a>, congratulations and thank you very much for the write up. I would like to ask for models that you trained with FreqConv, did you use pretrained models? If yes, how did you modify the initial layer of the pretrained models to accommodate one more input channel?  Thank you very much.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1227081,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "03/05/2021 07:13:33",
          "content": "<p>With one input channel instead of RGB it is quite easy to use pretrain, just sum weights from 3 channels. With 2 channels it’s a bit different. Timm library supports that out of the box, just pass in_chans=2.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335050,
      "author_name": "zhaijianyang",
      "author_url": "",
      "post_date": "06/04/2021 01:44:07",
      "content": "<p>Congratulations team!<br>\nCan you provide a link to github？</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1358732,
      "author_name": "alexlwh",
      "author_url": "",
      "post_date": "06/20/2021 17:31:05",
      "content": "<p>Thanks for the excellent write-up and congrats for the 2nd place! <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a> <br>\nAnd I hope I am not too late to ask a few questions to further understand your solution. I believe this could help me reproduce the solution:</p>\n<ol>\n<li>What was the cropped duration that you used for training?</li>\n<li>How did mixup work together with masked BCE loss? coz I can see there are some nuances here. (e.g. how did u define the target when u mix TP with FP? which labels did u mask for target that is mixup-ed?)<br>\n3a. in pseudo labeling stage, u mentioned you have special sampling strategy (i.e. sample TP/FP at p=0.5, otherwise randomly crop from a full clip) to train models that are responsible for generating pseudo labels, because <code>Without TP/FP sampling labels can become very soft and the score decreases after 2 or 3 rounds</code>. Could you elaborate more on this issue? <br>\n3b. Follow up on 3a. how did u define the target for sample that is randomly cropped from a full clip? Additionally, was this sampling strategy only applied in first round? (i.e. the round without any pseudo labels in training set). </li>\n<li>When you train your model with pseudo labels in your training set, did you handle pseudo labels specially? (e.g. apply penalty proportional to prediction confidence in the masked BCE loss). Also, how did u sample pseudo labels among TP/ FP? </li>\n</ol>\n<p>Look forward to your reply! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1210468": "## Overview\n\nI trained simple classification models (24 binary classes) with logmel spectrograms :\n1. bootstrap stage: models are trained on TP/FP with masked BCE loss\n2. generate soft pseudo labels with 0.5 second sliding window \n3. train models with pseudo labels and also sample (with p=0.5) places with TP/FP - this partially solves confirmation bias problem. \n\nRounds of pseudo labeling and retraining (points 2,3) were repeated until the score on public LB didn't improve. Depending on the settings it took  around 4-10 rounds to converge.\n\nMy initial models that gave 0.86 on TP/FP alone easily reached 0.96x with pseudo labeling . After this success I gave this challenge a 5 weeks break as I lost any motivation to improve my score :)\nLater to my surprise it was extremely hard to beat 0.97 even with improved first stage models.\n\n\n## Melspectrogram parameters\n- 256 mel bins\n- 512 hop length\n- original SR\n- 4096 nfft\n\n## FreqConv (CoordConv for frequency)\nAfter my first successful experiment with pseudo labeling that reached 0.969 on public LB  I tried to just swap encoders and blend models but this did not bring any improvements. \nSo I visualised the data for different classes and understood that when working with mel spectrograms for this task we don’t need translation invariance and classes really depend on both frequency and patterns.\nI added a channel to CNN input which contains the number of mel bin scaled to 0-1. This significantly improved validation metrics and after this change log loss on crops around TP after the first round of training with pseudo labels was around 0.04 (same for crops around FP). Though it only slightly improved results on the LB.\n\n\n## First stage \n\nFor the first stage I used all tp/fp information without any sampling and made crops around the center of the signal.\n\n**Augmentations**\n- time warping\n- random frequency masking below TP/FP signal\n- random frequency masking above TP/FP signal\n- gaussian noise\n- volume gain\n- mixup on spectrograms\n\nFor mixup on spectrograms - I used constant alpha (0.5) and hard labels with clipping (0,1). Masks were also added. \n\n## Pseudolabeling stages\nSampled TP/FP with p=0.5 otherwise made a random crop from the full spectrogram. \n\nWithout TP/FP sampling labels can become very soft and the score decreases after 2 or 3 rounds.\nAfter training 4 folds of effnet/rexnet I generated OOF labels and ensembled their predictions. Then the training is repeated from scratch.\n\n**Augmentations**\n- gaussian noise\n- volume gain\n- mixup \n- time warping\n- spec augment\n- mixup on spectrograms\n\n**Mixup**\n\nI used constant alpha (0.5) and added soft labels from two samples. This hurts logloss on FP a bit but at the same time significantly increases recall on TP.\n\n## Validation\n\t\nLocal validation did not have high correlation with the public leaderboard. Logloss on TP was somehow correlated but still it was not robust. \nSo without proper validation I decided to not select the best checkpoints and just trained 60 epochs (around 200 batches in each epoch) with CosineLR and AdamW optimizer.\n\nMy best models on validation - auc 0.999, log loss 0.03 did not produce great results (0.95). After the competition though It turned out that they can be easily improved with  postprocessing to 97x-98x range.\n\n\n## Final ensemble\n\nI used 4 models with 4 folds from Effnet and Rexnet (https://arxiv.org/abs/2007.00992  lightweight models with great performance) families:\n- Rexnet-200 (4 sec training/inference), EffnetB3 (4 sec training/inference)\n- Rexnet-150 (8 sec training/inference), EffnetB1 (8 sec training/inference)\n\nRexnet was much better than EfficientNet alone (less overfitting), but in ensemble they worked great.\n\nDuring inference I just used 0.5 second sliding window and took max probabilities for the full clip and then averaged predictions from different models.\n\n## Lessons learned\nI did not know about the paper and lacked this useful information about the dataset.\n\nIn my solutions I often rely on models alone but don’t explore the data deeply. \nIn this case I understood that the relabeled train set has similar class distribution to the test set and decided that models would easily learn that. I was wrong and simple post-processing could significantly improve results (though this happened due to severe class imbalance).",
    "1210541": "Rexnet-200 for 60 epochs.. wow)",
    "1210546": "Thanks for sharing.  and congrats on the solo prize!\n\nYou were first in finding that frequency masking per species and adding frequency location info was key.  And you made pseudo labeling work (not like nme). You deserve this great resuit.",
    "1210553": "https://arxiv.org/abs/2007.00992 Rexnet-200 (200 = 2.0 scale, 150 = 1.5)  is a lightweight model. It is not related to resnet, it is a modification of MobileNet that performs like EffNet. \nSo it was around 1 minute per epoch",
    "1210560": "Thanks! Actually pseudo-labeling did not work that well for SED models (tried framewise/clipwise pseudo labels) and I switched to standard image classification",
    "1210562": "Great work! Congrats on solo 2nd place and thanks for the writeup @selimsef",
    "1210571": "I never used SED.\n\nBut I think I now get why my pseudo labeling did not work.  Running experiemnt to check.",
    "1210574": "Thanks for sharing and congrats on getting the global 12th rank!",
    "1210577": "> Sampled TP/FP with p=0.5 otherwise made a random crop from the full spectrogram.\n\na question on this, do you train on the same random sampled audio crop and pseudo labels after or you re-sample tp/fp or made a random crop again in stage 2?",
    "1210587": "Thanks! I was exactly there some time ago😄Looks like I need to participate in a few competitions in a row (it is not possible to get high global rank otherwise due to Kaggle points decay)",
    "1210592": "Thanks for sharing, great work!\n\nDid you try to add freqs to the loss function (it could force the network to use the FreqConv)? \nI used CoorConv in some tasks and without adding the same coordinates to the loss function it didn't improve the network.\nI have a simple check for it - pass 0 or noise to the channel with coordinates, when I didn't add coordinates to the loss function, it performs the same as with proper data in the CoordConv channel. When I trained with coordinates in the loss, 0 or noise in the CoordConv - network inferences much worse.",
    "1210593": ">  a question on this, do you train on the same random sampled audio crop and pseudo labels after or you re-sample tp/fp or made a random crop again in stage 2?\n\nMy pseudolabels (for 4 second models) had 113 frames per audio clip (0.5 sec sliding window). \nIf TP/FP is sampled - I made a random crop around the center and then found nearest frame from pseudolabels. Pseudo labels were fixed using tp/fp. \nOtherwise I just took random frame from 113 and postprocessed labels if they overlap with tp/fp data.",
    "1210600": "interesting, I assume that the distance from the `random crop around the center` to the nearest frame from pseudo labels is so close that you can ignore the distance and just assign the same labels, I tried just using the same crops without random cropping again. I think the key probably was that you trained a fixed 60 epochs 😂 I couldn't get it to work properly when I didn't have extra hand labels yet.",
    "1210604": "Good point! Just checked some checkpoint\n```\nWith proper frequencies\nneg_logloss: 0.1588028629548308\npos_logloss: 0.05171160377776966\n\nWith zeros\nneg_logloss: 0.19649322897329777\npos_logloss: 0.3455986050609499\n\n```\nSo my models really use it somehow",
    "1210608": "max distance = 0.25 seconds, so it should not have big impact on model's quality. More frames would be better but it would take much more time to generate pseudolabels.",
    "1210610": "Indeed, looks like the network managed to use it without loss modifications.",
    "1210612": "Thanks for sharing your work. i am curious to learn new thing from your work\ni have a question on masked BCE loss. how is this implemented. when you say mask, are you masking the loss for other classes that are not in the label or masking the time frames(say 4sec input where the label is only for sec1-2) where there is no label.  \n\nsecond question . how are your using both TP and FP. how this loss fn will look like.",
    "1210629": "Masks for loss function have the same shape as labels, 1 for the classes we know (TP/FP) 0 for others\n```\nclass BCEMasked(nn.Module):\n    def forward(self, inputs, targets, mask=None):\n        bce_loss = binary_cross_entropy_with_logits(inputs, targets, reduction='none')\n        if mask is not None:\n            bce_loss = bce_loss[mask > 0]\n        return bce_loss.mean()\n```",
    "1210652": "thanks for answering . about TP/FP , how are you using FP for training the model. with sigmoid you cant have FP = -1.",
    "1210662": "example - FP for s0\nmask=[1, 0, 0, ...,0], targets=[0, 0, 0, ..., 0] - only the first element from the targets and output will be considered by the loss function",
    "1210665": "ahh. now i understand . ... thanks for explaining .",
    "1211187": "\"why my pseudo labeling did not work. Running experiemnt to check.\"\n\none simple way to check pseudo label is:\n- make a mean image of the pseudo label cropped melspec and see if it looks familiar to the train tp (or fp)\n- use tsne to show your train tp,fp and pseudo label \n- measure purity based on the validation set\n\nfor image application, I usually do a manual inspection to estimate purity. e.g. test random sample of 100 pseudo label samples",
    "1211189": "selimsef \nthanks for the write-up and congrats on your winning.\n\nSo it confirms that 0.960~0.970 is the limit with the naive pesudo label.",
    "1211535": "Actually I tried different  SSL approaches, pseudo labeling with mining confident labels and loss masking was much worse than soft labels.  Postprocessing could fix that I guess. The main reason for than is comp metric. Something like F1 per class would better assess predictive performance and  would not be affected by class imbalance.",
    "1211717": "Selim, thank you for the write up. I love how elegant/simple your solution is.",
    "1211802": "Thanks for sharing!\nSeems that the score benefits most from training 24 models separately right? I tried training 24 models and some has reached 99.9% accuracy locally(some just around 60% 80%) but still got no-good public leaderboard score and I just suspect that there was something wrong in my models validation procedure.",
    "1211843": "No, single 4 fold model scores 0.97 on public LB. Single checkpoint something like 0.968 (Public). Where did you see 24 models?)",
    "1211851": "i tried to include the masking for loss fn. i am facing issue with having many predictions set to 1.0 and few classes are having 0.99 in the test dataset. i am still using my SED model to try this implementation. do you see any issue. i am using 6sec clip around TP.\n\n`1.0\n0.99999994\n0.90000004\n1.0\n1.0\n1.0\n1.0`",
    "1211932": "you are using only TP samples, of course if you have all targets = 1 NN will predict 1.0",
    "1211958": "i am doing this in the following way\n```\nprediction = model output\ntarget = true target\nloss = loss(prediction, target, mask = target)\n```\n\nhere target will act as mask as it has 1 for the class present and rest are 0",
    "1212213": "I agree that frequency location is the key to this competition according to this ![](https://i.imgur.com/D3g3vJc.png)\n\n@cpmpml your solution is a perfect way to capture such information.",
    "1212214": "Oh.. Thanks so top3 are training on a single model, I was just reading too fast and misunderstood that, sorry for that",
    "1212802": "Thanks for the writeup and congratulations on your winning. I thought that you would not make any other submission after that 0.97 😄\n\nI wonder that what would be your final score with the mentioned postprocessing methods like Chris's method [here](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389)",
    "1212840": "982-983 depending on submission",
    "1212889": "Woow, it's great. Thanks 👍\n\nDid you have any idea that you wanted to try but didn't?",
    "1219843": "Hi @xmpgek , could you talk more details about how to add the coordinates to the loss function?",
    "1226031": "Hi @selimsef, congratulations and thank you very much for the write up. I would like to ask for models that you trained with FreqConv, did you use pretrained models? If yes, how did you modify the initial layer of the pretrained models to accommodate one more input channel?  Thank you very much.",
    "1227081": "With one input channel instead of RGB it is quite easy to use pretrain, just sum weights from 3 channels. With 2 channels it’s a bit different. Timm library supports that out of the box, just pass in_chans=2.",
    "1286808": "Hello, I have just re-read your post; it just looks like the first sentence may be a bit misleading to me, though I am not sure if it is my own problem. \"I trained simple classification models (24 binary classes) with logmel spectrograms :\" looks to me that it is 24 binary classification model? May you clarify on this? Thanks.",
    "1335050": "Congratulations team!\nCan you provide a link to github？",
    "1358732": "Thanks for the excellent write-up and congrats for the 2nd place! @selimsef \nAnd I hope I am not too late to ask a few questions to further understand your solution. I believe this could help me reproduce the solution:\n1. What was the cropped duration that you used for training?\n2. How did mixup work together with masked BCE loss? coz I can see there are some nuances here. (e.g. how did u define the target when u mix TP with FP? which labels did u mask for target that is mixup-ed?)\n3a. in pseudo labeling stage, u mentioned you have special sampling strategy (i.e. sample TP/FP at p=0.5, otherwise randomly crop from a full clip) to train models that are responsible for generating pseudo labels, because `Without TP/FP sampling labels can become very soft and the score decreases after 2 or 3 rounds`. Could you elaborate more on this issue? \n3b. Follow up on 3a. how did u define the target for sample that is randomly cropped from a full clip? Additionally, was this sampling strategy only applied in first round? (i.e. the round without any pseudo labels in training set). \n4. When you train your model with pseudo labels in your training set, did you handle pseudo labels specially? (e.g. apply penalty proportional to prediction confidence in the masked BCE loss). Also, how did u sample pseudo labels among TP/ FP? \n\nLook forward to your reply!"
  },
  "source": "meta"
}