{
  "id": 220319,
  "title": "10th-place solution - Agnostic loss for semi-supervised learning",
  "url": "/competitions/rfcx-species-audio-detection/writeups/rna-10th-place-solution-agnostic-loss-for-semi-sup",
  "author_name": "",
  "post_date": "2021-02-18T21:56:06.803Z",
  "votes": 52,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Thank you to the organisers, Kaggle, and to everyone who shared ideas and code for this competition. I learned a lot, as I'm sure many of you have, and I thought I would break down my approach since I know many competitors couldn't find a way to use the False Positive labels. I'm thrilled to have secured my first gold (and solo gold) in a competition, so it's the least I could do!</p>\n<h3>Summary</h3>\n<p>On a high-level, my approach was as follows:</p>\n<ul>\n<li>train models using existing labels</li>\n<li>generate pseudo-labels on train/test</li>\n<li>isolate the frames which had a very high/low prob across an ensemble of models</li>\n<li>conservatively threshold these and use as new TP/FP values for the relevant class</li>\n<li>repeat</li>\n</ul>\n<p>This gradually increased the amount of data I had to work with, until I had at least 2 frames from each recording with an identified TP and/or FP. The growing data diversity allowed successive models to generalise better.</p>\n<h3>What Worked</h3>\n<p><strong>Fixed Windows</strong></p>\n<p>I used a window of 5 seconds centred on the TP/FP. For predicting on test, I used overlapping windows with a step of 2.5 seconds. This ensured that train/test had identical preprocessing for their inputs. The maximum value was taken for each class across all of these windows. For some submissions, I used the average of the top 3 predictions but this didn't seem to notably change the LB score.</p>\n<p><strong>Agnostic Loss</strong></p>\n<p>Perhaps a better term already exists for this, but this was what I called my method for using both the TPs and FPs in a semi-supervised fashion. The problem we face with unlabelled data is that any spectrogram can contain multiple classes, so setting the target as 0 for everything apart from the <em>given</em> label will penalise true positives for other classes. We can only know for sure that one class is present or absent, and the loss needs to reflect this. So I excluded all non-definite targets from the loss calculation. In the target tensor, a TP is 1 while an FP is 0. Unlabelled classes are given as 0.5. These values are then excluded from the loss calculation. So if we had 5 classes (we have 24, but I'm saving room here) and this time window contained a TP for class 0 and an FP for class 3:</p>\n<pre><code>y = torch.Tensor([1., 0.5, 0.5, 0., 0.5]) \n</code></pre>\n<p>And in the loss calculation:</p>\n<pre><code>preds = model(inputs)\npreds[targets==0.5] = 0                    \nloss = BCEWithLogitsLoss(preds, targets)\nloss.backward()\n</code></pre>\n<p>Thus the model is 'agnostic' to the majority of the potential labels. This allows the model to build a guided feature representation of the different classes without being inadvertently given false negatives. This approach gave me substantially better LB scores.</p>\n<p>The figure of 0.5 is arbitrary and could've been any value apart from 0 or 1: the salient point is that the loss resulting from unlabelled classes is always constant. Note that this kind of inplace operation is incompatible with <code>nn.Sigmoid</code> or its functional equivalent when performing backprop so you need to use the raw logits via <code>torch.nn.BCEWithLogitsLoss()</code>.</p>\n<p><strong>ResNeSt</strong></p>\n<p>I found EfficientNet to be surprisingly poor in this competition, and all of my best scores came from using variants of ResNeSt (<a href=\"https://github.com/zhanghang1989/ResNeSt\" target=\"_blank\">https://github.com/zhanghang1989/ResNeSt</a>) paper available <a href=\"https://arxiv.org/abs/2004.08955\" target=\"_blank\">here</a>.</p>\n<p>For 3-channel input I used the <code>librosa</code> mel-spectrogram with power 1, power 2 and the <code>delta</code> function to capture temporal information. With some models I experimented with a single power-1 spectrogram, delta and delta-delta features instead. While quicker to preprocess, I noticed no impact on scores.</p>\n<p>I also incorporated it into the SED architecture as the encoder. This showed very promising metrics during training, and while sadly I didn't have time to run a fully-convergent example its inclusion still helped my score. In future competitions this could be a very useful model. ResNeSt itself only takes a 3-channel input and has no inbuilt function to extract features, so I had to rejig it to work properly: I'll be uploading a script with that model shortly in case anyone is interested.</p>\n<p><strong>Augmentations</strong></p>\n<p>From Hidehisa Arai's excellent kernel <a href=\"https://www.kaggle.com/hidehisaarai1213/rfcx-audio-data-augmentation-japanese-english\" target=\"_blank\">here</a>, I selected <code>GaussianNoiseSNR()</code>, <code>PinkNoiseSNR()</code>,  <code>TimeShift()</code> and <code>VolumeControl()</code>. I was wary of augmentation methods that blank out time windows or frequency bands like <a href=\"https://arxiv.org/abs/1904.08779\" target=\"_blank\">SpecAugment</a>. Some of the sounds occur in a very narrow frequency range (e.g. species 21) or in a very narrow time window (e.g. species 9) and I didn't want to make 'empty', positive samples that would coerce the model into learning spurious features. I also added some trivial augmentations of my own:</p>\n<ul>\n<li>swapping the first and last half of the audio vector</li>\n<li>adding a random constant before spectrogram normalisation (occluding the relevant features)</li>\n<li>'jiggling' the time window around the centre of <code>t_mid</code>, at a maximum of 1 second offset in either direction</li>\n</ul>\n<p><strong>Eliminating species 19</strong></p>\n<p>A minor point, but this class was so rare that setting all of its predictions to zero usually improved the LB score by 0.001. There were many unrelated sounds that would presumably cause the model to produce false positives. My best submission didn't do this however; it was a simple blend of models including ResNeSt-50, ResNeSt-101, EfficientNet-b1 and the SED architecture described above. I used weighted averaging roughly in proportion to the individual models' performance.</p>\n<h3>What didn't work</h3>\n<ul>\n<li>Using separate models for different frequency ranges (these models never gained an adequate feature representation, and produced many false positives).</li>\n<li>EfficientNet alone gave poor results, but helped as part of an ensemble.</li>\n<li>Larger models (ResNeSt101, EfficientNet b3) didn't improve scores.</li>\n<li>TP-only training.</li>\n<li>Models that worked on all windows for a single clip - these were slow and produced inferior results.</li>\n</ul>\n<p>Otherwise I was quite lucky - I thought about my methodology for a while and most of what I tried worked well on the first attempt. If I'd had more time, I would have liked to try:</p>\n<ul>\n<li>automatically labelling some species as an FP of a similar class (e.g species 9 &amp; 17)</li>\n<li>probing the LB for class distribution (I suspect you could get +0.9 by only predicting the most common half of the classes and ignoring everything else) - I realised the importance of this too close to the deadline</li>\n<li>experimenting with different encoders for the SED architecture.</li>\n<li>using a smaller window size (&lt;=3 seconds) for greater fidelity. </li>\n</ul>\n<p>The overall class prediction histograms for my final submission were as follows:</p>\n<p><img src=\"https://i.imgur.com/GmJ9BkI.png\" alt=\"\"></p>\n<p>Some classes gave me particular trouble. I used my own, simple scoring metric during training that recorded the proportion of positive cases that were predicted above a certain threshold. I never satisfactorily made a model that could detect the rarer classes like 6, 19 or 20 in a reliable fashion.</p>\n<p>Overall I had an interesting time exploring how to work with the application of CNNs to spectrograms, and with large amounts of unlabelled data. In the next audio competition, perhaps I'll aim a little higher! I'm looking forward to seeing how those who scored &gt; 0.97 managed to achieve their results.</p>\n<p>If you have any questions I'll do my best to answer them!</p>",
  "messages": [
    {
      "id": "1207705",
      "postDate": "02/18/2021 01:31:08",
      "content": "<p>Thank you to the organisers, Kaggle, and to everyone who shared ideas and code for this competition. I learned a lot, as I'm sure many of you have, and I thought I would break down my approach since I know many competitors couldn't find a way to use the False Positive labels. I'm thrilled to have secured my first gold (and solo gold) in a competition, so it's the least I could do!</p>\n<h3>Summary</h3>\n<p>On a high-level, my approach was as follows:</p>\n<ul>\n<li>train models using existing labels</li>\n<li>generate pseudo-labels on train/test</li>\n<li>isolate the frames which had a very high/low prob across an ensemble of models</li>\n<li>conservatively threshold these and use as new TP/FP values for the relevant class</li>\n<li>repeat</li>\n</ul>\n<p>This gradually increased the amount of data I had to work with, until I had at least 2 frames from each recording with an identified TP and/or FP. The growing data diversity allowed successive models to generalise better.</p>\n<h3>What Worked</h3>\n<p><strong>Fixed Windows</strong></p>\n<p>I used a window of 5 seconds centred on the TP/FP. For predicting on test, I used overlapping windows with a step of 2.5 seconds. This ensured that train/test had identical preprocessing for their inputs. The maximum value was taken for each class across all of these windows. For some submissions, I used the average of the top 3 predictions but this didn't seem to notably change the LB score.</p>\n<p><strong>Agnostic Loss</strong></p>\n<p>Perhaps a better term already exists for this, but this was what I called my method for using both the TPs and FPs in a semi-supervised fashion. The problem we face with unlabelled data is that any spectrogram can contain multiple classes, so setting the target as 0 for everything apart from the <em>given</em> label will penalise true positives for other classes. We can only know for sure that one class is present or absent, and the loss needs to reflect this. So I excluded all non-definite targets from the loss calculation. In the target tensor, a TP is 1 while an FP is 0. Unlabelled classes are given as 0.5. These values are then excluded from the loss calculation. So if we had 5 classes (we have 24, but I'm saving room here) and this time window contained a TP for class 0 and an FP for class 3:</p>\n<pre><code>y = torch.Tensor([1., 0.5, 0.5, 0., 0.5]) \n</code></pre>\n<p>And in the loss calculation:</p>\n<pre><code>preds = model(inputs)\npreds[targets==0.5] = 0                    \nloss = BCEWithLogitsLoss(preds, targets)\nloss.backward()\n</code></pre>\n<p>Thus the model is 'agnostic' to the majority of the potential labels. This allows the model to build a guided feature representation of the different classes without being inadvertently given false negatives. This approach gave me substantially better LB scores.</p>\n<p>The figure of 0.5 is arbitrary and could've been any value apart from 0 or 1: the salient point is that the loss resulting from unlabelled classes is always constant. Note that this kind of inplace operation is incompatible with <code>nn.Sigmoid</code> or its functional equivalent when performing backprop so you need to use the raw logits via <code>torch.nn.BCEWithLogitsLoss()</code>.</p>\n<p><strong>ResNeSt</strong></p>\n<p>I found EfficientNet to be surprisingly poor in this competition, and all of my best scores came from using variants of ResNeSt (<a href=\"https://github.com/zhanghang1989/ResNeSt\" target=\"_blank\">https://github.com/zhanghang1989/ResNeSt</a>) paper available <a href=\"https://arxiv.org/abs/2004.08955\" target=\"_blank\">here</a>.</p>\n<p>For 3-channel input I used the <code>librosa</code> mel-spectrogram with power 1, power 2 and the <code>delta</code> function to capture temporal information. With some models I experimented with a single power-1 spectrogram, delta and delta-delta features instead. While quicker to preprocess, I noticed no impact on scores.</p>\n<p>I also incorporated it into the SED architecture as the encoder. This showed very promising metrics during training, and while sadly I didn't have time to run a fully-convergent example its inclusion still helped my score. In future competitions this could be a very useful model. ResNeSt itself only takes a 3-channel input and has no inbuilt function to extract features, so I had to rejig it to work properly: I'll be uploading a script with that model shortly in case anyone is interested.</p>\n<p><strong>Augmentations</strong></p>\n<p>From Hidehisa Arai's excellent kernel <a href=\"https://www.kaggle.com/hidehisaarai1213/rfcx-audio-data-augmentation-japanese-english\" target=\"_blank\">here</a>, I selected <code>GaussianNoiseSNR()</code>, <code>PinkNoiseSNR()</code>,  <code>TimeShift()</code> and <code>VolumeControl()</code>. I was wary of augmentation methods that blank out time windows or frequency bands like <a href=\"https://arxiv.org/abs/1904.08779\" target=\"_blank\">SpecAugment</a>. Some of the sounds occur in a very narrow frequency range (e.g. species 21) or in a very narrow time window (e.g. species 9) and I didn't want to make 'empty', positive samples that would coerce the model into learning spurious features. I also added some trivial augmentations of my own:</p>\n<ul>\n<li>swapping the first and last half of the audio vector</li>\n<li>adding a random constant before spectrogram normalisation (occluding the relevant features)</li>\n<li>'jiggling' the time window around the centre of <code>t_mid</code>, at a maximum of 1 second offset in either direction</li>\n</ul>\n<p><strong>Eliminating species 19</strong></p>\n<p>A minor point, but this class was so rare that setting all of its predictions to zero usually improved the LB score by 0.001. There were many unrelated sounds that would presumably cause the model to produce false positives. My best submission didn't do this however; it was a simple blend of models including ResNeSt-50, ResNeSt-101, EfficientNet-b1 and the SED architecture described above. I used weighted averaging roughly in proportion to the individual models' performance.</p>\n<h3>What didn't work</h3>\n<ul>\n<li>Using separate models for different frequency ranges (these models never gained an adequate feature representation, and produced many false positives).</li>\n<li>EfficientNet alone gave poor results, but helped as part of an ensemble.</li>\n<li>Larger models (ResNeSt101, EfficientNet b3) didn't improve scores.</li>\n<li>TP-only training.</li>\n<li>Models that worked on all windows for a single clip - these were slow and produced inferior results.</li>\n</ul>\n<p>Otherwise I was quite lucky - I thought about my methodology for a while and most of what I tried worked well on the first attempt. If I'd had more time, I would have liked to try:</p>\n<ul>\n<li>automatically labelling some species as an FP of a similar class (e.g species 9 &amp; 17)</li>\n<li>probing the LB for class distribution (I suspect you could get +0.9 by only predicting the most common half of the classes and ignoring everything else) - I realised the importance of this too close to the deadline</li>\n<li>experimenting with different encoders for the SED architecture.</li>\n<li>using a smaller window size (&lt;=3 seconds) for greater fidelity. </li>\n</ul>\n<p>The overall class prediction histograms for my final submission were as follows:</p>\n<p><img src=\"https://i.imgur.com/GmJ9BkI.png\" alt=\"\"></p>\n<p>Some classes gave me particular trouble. I used my own, simple scoring metric during training that recorded the proportion of positive cases that were predicted above a certain threshold. I never satisfactorily made a model that could detect the rarer classes like 6, 19 or 20 in a reliable fashion.</p>\n<p>Overall I had an interesting time exploring how to work with the application of CNNs to spectrograms, and with large amounts of unlabelled data. In the next audio competition, perhaps I'll aim a little higher! I'm looking forward to seeing how those who scored &gt; 0.97 managed to achieve their results.</p>\n<p>If you have any questions I'll do my best to answer them!</p>",
      "rawMarkdown": "Thank you to the organisers, Kaggle, and to everyone who shared ideas and code for this competition. I learned a lot, as I'm sure many of you have, and I thought I would break down my approach since I know many competitors couldn't find a way to use the False Positive labels. I'm thrilled to have secured my first gold (and solo gold) in a competition, so it's the least I could do!\n\n### Summary\n\nOn a high-level, my approach was as follows:\n\n* train models using existing labels\n* generate pseudo-labels on train/test\n* isolate the frames which had a very high/low prob across an ensemble of models\n* conservatively threshold these and use as new TP/FP values for the relevant class\n* repeat\n\nThis gradually increased the amount of data I had to work with, until I had at least 2 frames from each recording with an identified TP and/or FP. The growing data diversity allowed successive models to generalise better.\n\n### What Worked\n\n**Fixed Windows**\n\nI used a window of 5 seconds centred on the TP/FP. For predicting on test, I used overlapping windows with a step of 2.5 seconds. This ensured that train/test had identical preprocessing for their inputs. The maximum value was taken for each class across all of these windows. For some submissions, I used the average of the top 3 predictions but this didn't seem to notably change the LB score.\n\n**Agnostic Loss**\n\nPerhaps a better term already exists for this, but this was what I called my method for using both the TPs and FPs in a semi-supervised fashion. The problem we face with unlabelled data is that any spectrogram can contain multiple classes, so setting the target as 0 for everything apart from the *given* label will penalise true positives for other classes. We can only know for sure that one class is present or absent, and the loss needs to reflect this. So I excluded all non-definite targets from the loss calculation. In the target tensor, a TP is 1 while an FP is 0. Unlabelled classes are given as 0.5. These values are then excluded from the loss calculation. So if we had 5 classes (we have 24, but I'm saving room here) and this time window contained a TP for class 0 and an FP for class 3:\n\n```\ny = torch.Tensor([1., 0.5, 0.5, 0., 0.5]) \n\n```\nAnd in the loss calculation:\n```\npreds = model(inputs)\npreds[targets==0.5] = 0                    \nloss = BCEWithLogitsLoss(preds, targets)\nloss.backward()\n```\nThus the model is 'agnostic' to the majority of the potential labels. This allows the model to build a guided feature representation of the different classes without being inadvertently given false negatives. This approach gave me substantially better LB scores.\n\nThe figure of 0.5 is arbitrary and could've been any value apart from 0 or 1: the salient point is that the loss resulting from unlabelled classes is always constant. Note that this kind of inplace operation is incompatible with `nn.Sigmoid` or its functional equivalent when performing backprop so you need to use the raw logits via `torch.nn.BCEWithLogitsLoss()`.\n\n\n**ResNeSt**\n\nI found EfficientNet to be surprisingly poor in this competition, and all of my best scores came from using variants of ResNeSt (https://github.com/zhanghang1989/ResNeSt) paper available [here](https://arxiv.org/abs/2004.08955).\n\nFor 3-channel input I used the `librosa` mel-spectrogram with power 1, power 2 and the `delta` function to capture temporal information. With some models I experimented with a single power-1 spectrogram, delta and delta-delta features instead. While quicker to preprocess, I noticed no impact on scores.\n\nI also incorporated it into the SED architecture as the encoder. This showed very promising metrics during training, and while sadly I didn't have time to run a fully-convergent example its inclusion still helped my score. In future competitions this could be a very useful model. ResNeSt itself only takes a 3-channel input and has no inbuilt function to extract features, so I had to rejig it to work properly: I'll be uploading a script with that model shortly in case anyone is interested.\n\n**Augmentations**\n\nFrom Hidehisa Arai's excellent kernel [here](https://www.kaggle.com/hidehisaarai1213/rfcx-audio-data-augmentation-japanese-english), I selected `GaussianNoiseSNR()`, `PinkNoiseSNR()`,  `TimeShift()` and `VolumeControl()`. I was wary of augmentation methods that blank out time windows or frequency bands like [SpecAugment](https://arxiv.org/abs/1904.08779). Some of the sounds occur in a very narrow frequency range (e.g. species 21) or in a very narrow time window (e.g. species 9) and I didn't want to make 'empty', positive samples that would coerce the model into learning spurious features. I also added some trivial augmentations of my own:\n\n* swapping the first and last half of the audio vector\n* adding a random constant before spectrogram normalisation (occluding the relevant features)\n* 'jiggling' the time window around the centre of `t_mid`, at a maximum of 1 second offset in either direction\n\n**Eliminating species 19**\n\nA minor point, but this class was so rare that setting all of its predictions to zero usually improved the LB score by 0.001. There were many unrelated sounds that would presumably cause the model to produce false positives. My best submission didn't do this however; it was a simple blend of models including ResNeSt-50, ResNeSt-101, EfficientNet-b1 and the SED architecture described above. I used weighted averaging roughly in proportion to the individual models' performance.\n\n### What didn't work\n\n* Using separate models for different frequency ranges (these models never gained an adequate feature representation, and produced many false positives).\n* EfficientNet alone gave poor results, but helped as part of an ensemble.\n* Larger models (ResNeSt101, EfficientNet b3) didn't improve scores.\n* TP-only training.\n* Models that worked on all windows for a single clip - these were slow and produced inferior results.\n\nOtherwise I was quite lucky - I thought about my methodology for a while and most of what I tried worked well on the first attempt. If I'd had more time, I would have liked to try:\n\n* automatically labelling some species as an FP of a similar class (e.g species 9 & 17)\n* probing the LB for class distribution (I suspect you could get +0.9 by only predicting the most common half of the classes and ignoring everything else) - I realised the importance of this too close to the deadline\n* experimenting with different encoders for the SED architecture.\n* using a smaller window size (<=3 seconds) for greater fidelity. \n\nThe overall class prediction histograms for my final submission were as follows:\n\n![](https://i.imgur.com/GmJ9BkI.png)\n\nSome classes gave me particular trouble. I used my own, simple scoring metric during training that recorded the proportion of positive cases that were predicted above a certain threshold. I never satisfactorily made a model that could detect the rarer classes like 6, 19 or 20 in a reliable fashion.\n\nOverall I had an interesting time exploring how to work with the application of CNNs to spectrograms, and with large amounts of unlabelled data. In the next audio competition, perhaps I'll aim a little higher! I'm looking forward to seeing how those who scored > 0.97 managed to achieve their results.\n\nIf you have any questions I'll do my best to answer them!",
      "votes": null
    },
    {
      "id": "1207709",
      "postDate": "02/18/2021 01:35:21",
      "content": "<p>Congrats on 11th place and solo gold medal  <a href=\"https://www.kaggle.com/bigironsphere\" target=\"_blank\">@bigironsphere</a>. Thanks for the writeup. </p>",
      "rawMarkdown": "Congrats on 11th place and solo gold medal  @bigironsphere. Thanks for the writeup.",
      "votes": null
    },
    {
      "id": "1207742",
      "postDate": "02/18/2021 02:34:45",
      "content": "<p>I've provided the code for single-channel ResNeSt and its incorporation into the PANNs SED model here:</p>\n<p><a href=\"https://www.kaggle.com/bigironsphere/single-channel-resnest-for-panns-sed-architecture\" target=\"_blank\">https://www.kaggle.com/bigironsphere/single-channel-resnest-for-panns-sed-architecture</a></p>\n<p>I hope someone finds it useful!</p>",
      "rawMarkdown": "I've provided the code for single-channel ResNeSt and its incorporation into the PANNs SED model here:\n\nhttps://www.kaggle.com/bigironsphere/single-channel-resnest-for-panns-sed-architecture\n\nI hope someone finds it useful!",
      "votes": null
    },
    {
      "id": "1207775",
      "postDate": "02/18/2021 03:17:16",
      "content": "<p>Congratulations on a strong solo finish! You used few submissions and the LB was filled with grandmasters and masters. Very awesome, the loss masking seems key!</p>",
      "rawMarkdown": "Congratulations on a strong solo finish! You used few submissions and the LB was filled with grandmasters and masters. Very awesome, the loss masking seems key!",
      "votes": null
    },
    {
      "id": "1207999",
      "postDate": "02/18/2021 05:35:42",
      "content": "<p>Great result. Congratulations. I went down a lot of the same avenues, just never really put them all together. The process of iteratively making more and more labeled data from pseudolabeling the training data was also mentioned in the birdcall competition but I never got around to it. How important do you think that process was to your result?</p>\n<p>And you mention the 3 different input channels, do you know how significant this was vs just copying the same spectogram representation across all 3 channels?</p>",
      "rawMarkdown": "Great result. Congratulations. I went down a lot of the same avenues, just never really put them all together. The process of iteratively making more and more labeled data from pseudolabeling the training data was also mentioned in the birdcall competition but I never got around to it. How important do you think that process was to your result?\n\nAnd you mention the 3 different input channels, do you know how significant this was vs just copying the same spectogram representation across all 3 channels?",
      "votes": null
    },
    {
      "id": "1208471",
      "postDate": "02/18/2021 09:49:01",
      "content": "<p>Congratz ! I knew that you had came up with a nice pipeline when you asked about reproductibility. Glad to see you at the top :)</p>",
      "rawMarkdown": "Congratz ! I knew that you had came up with a nice pipeline when you asked about reproductibility. Glad to see you at the top :)",
      "votes": null
    },
    {
      "id": "1208900",
      "postDate": "02/18/2021 14:51:34",
      "content": "<p>I believe it was very important, but I've been through some of my predictions and the pseudolabels were indeed poor for some classes - species 20 looks to be the worst. The most common classes in the data seem to be ones that were quite easy for models to identify, which makes me wonder if this competition would've been more interesting if we were evaluated on class-averaged metrics.</p>\n<p>Good question about the input channels - I never even tested your alternative. Since non-SED models were a significant part of my submission, I just assumed it would be important to include some temporal information via delta and/or delta-delta.</p>",
      "rawMarkdown": "I believe it was very important, but I've been through some of my predictions and the pseudolabels were indeed poor for some classes - species 20 looks to be the worst. The most common classes in the data seem to be ones that were quite easy for models to identify, which makes me wonder if this competition would've been more interesting if we were evaluated on class-averaged metrics.\n\nGood question about the input channels - I never even tested your alternative. Since non-SED models were a significant part of my submission, I just assumed it would be important to include some temporal information via delta and/or delta-delta.",
      "votes": null
    },
    {
      "id": "1208901",
      "postDate": "02/18/2021 14:51:47",
      "content": "<p>Cheers Corey!</p>",
      "rawMarkdown": "Cheers Corey!",
      "votes": null
    },
    {
      "id": "1209100",
      "postDate": "02/18/2021 17:22:37",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/bigironsphere\" target=\"_blank\">@bigironsphere</a> , very smart \"Agnostic Loss\" which fixed partial label problem and leverage fp data meanwhile. </p>\n<blockquote>\n  <p>The figure of 0.5 is arbitrary and could've been any value apart from 0 or 1: the salient point is that the loss resulting from unlabelled classes is always constant. </p>\n</blockquote>\n<p>Regarding above description, after study your \"Agnostic Loss\" and did some calcalation, I feel you have to set target value(y) of unlablled class to 0.5, as definintion of BCEWITHLOGITSLOSS <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.BCEWithLogitsLoss.html#bcewithlogitsloss\" target=\"_blank\">here</a>, for A = sigmoid(x_n), Z = x_n = WX+b, here X stands for input and W stands for weights, then the gradient of weights is(sorry I don`t know how to insert LaTex here): <br>\ndZ = A - Y<br>\ndW = dZ * (A^(L-1))^T</p>\n<p>IMHO your trick is to let <code>$x_n$</code> be 0 thus A = sigmoid(<code>$x_n$</code>) = 0.5, and you set Y of unlabelld classes to 0.5 as well, thus <code>$dZ=0.5-0.5=0$</code>, so <code>$dW$</code>(gradient for unlabelled classes) is always 0, then your model`s weights are only updated for labelled classes when gradient descent. But if you set Y of unlabelld classes to other value than 0.5, then the gradient for unlabelled classes will not be zero, thus your training will affected by unlabelled classes. Please correct me if I am wrong.</p>",
      "rawMarkdown": "Congratulations @bigironsphere , very smart \"Agnostic Loss\" which fixed partial label problem and leverage fp data meanwhile. \n> The figure of 0.5 is arbitrary and could've been any value apart from 0 or 1: the salient point is that the loss resulting from unlabelled classes is always constant. \n\nRegarding above description, after study your \"Agnostic Loss\" and did some calcalation, I feel you have to set target value(y) of unlablled class to 0.5, as definintion of BCEWITHLOGITSLOSS [here](https://pytorch.org/docs/stable/generated/torch.nn.BCEWithLogitsLoss.html#bcewithlogitsloss), for A = sigmoid(x_n), Z = x_n = WX+b, here X stands for input and W stands for weights, then the gradient of weights is(sorry I don`t know how to insert LaTex here): \ndZ = A - Y\ndW = dZ * (A^(L-1))^T\n\nIMHO your trick is to let `$x_n$` be 0 thus A = sigmoid(`$x_n$`) = 0.5, and you set Y of unlabelld classes to 0.5 as well, thus `$dZ=0.5-0.5=0$`, so `$dW$`(gradient for unlabelled classes) is always 0, then your model`s weights are only updated for labelled classes when gradient descent. But if you set Y of unlabelld classes to other value than 0.5, then the gradient for unlabelled classes will not be zero, thus your training will affected by unlabelled classes. Please correct me if I am wrong.",
      "votes": null
    },
    {
      "id": "1209118",
      "postDate": "02/18/2021 17:35:01",
      "content": "<p>0.5 is just a placeholder value. It is just a specific distinction so that later on he can zero out the loss for any sample that had a label of 0.5. He could have set the value to 1000 instead and it would not make a difference because loss of those samples for those classes are never backpropogated </p>",
      "rawMarkdown": "0.5 is just a placeholder value. It is just a specific distinction so that later on he can zero out the loss for any sample that had a label of 0.5. He could have set the value to 1000 instead and it would not make a difference because loss of those samples for those classes are never backpropogated",
      "votes": null
    },
    {
      "id": "1209392",
      "postDate": "02/18/2021 21:32:06",
      "content": "<p>Thanks for sharing.  We did something quite similar but you made it work better.  I'll reread more carefully to see where you made a difference, maybe the model (I used efficientnet) or the pseudo labels.</p>\n<p>Congrats on the end result!</p>",
      "rawMarkdown": "Thanks for sharing.  We did something quite similar but you made it work better.  I'll reread more carefully to see where you made a difference, maybe the model (I used efficientnet) or the pseudo labels.\n\nCongrats on the end result!",
      "votes": null
    },
    {
      "id": "1209416",
      "postDate": "02/18/2021 21:58:29",
      "content": "<p>Thanks! Using ResNeSt gave me substantially better results. If you want the SED variant, it's <a href=\"https://www.kaggle.com/bigironsphere/single-channel-resnest-for-panns-sed-architecture\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Thanks! Using ResNeSt gave me substantially better results. If you want the SED variant, it's [here](https://www.kaggle.com/bigironsphere/single-channel-resnest-for-panns-sed-architecture)",
      "votes": null
    },
    {
      "id": "1228436",
      "postDate": "03/06/2021 11:30:07",
      "content": "<p>Congrats on the solo gold <a href=\"https://www.kaggle.com/bigironsphere\" target=\"_blank\">@bigironsphere</a> ! <br>\nI have a question:</p>\n<blockquote>\n  <p>isolate the frames which had a very high/low prob across an ensemble of models.</p>\n</blockquote>\n<p>Why you remove those frames with very high/low probs first? Is it because you think these are caused by the uncertainty of the model?</p>",
      "rawMarkdown": "Congrats on the solo gold @bigironsphere ! \nI have a question:\n\n> isolate the frames which had a very high/low prob across an ensemble of models.\n\nWhy you remove those frames with very high/low probs first? Is it because you think these are caused by the uncertainty of the model?",
      "votes": null
    },
    {
      "id": "1241110",
      "postDate": "03/16/2021 22:24:00",
      "content": "<p>To clarify, those are the frames which I kept - the scores indicated that the model was very confident a class was/wasn't present in it. I would then use this as a definite label.</p>",
      "rawMarkdown": "To clarify, those are the frames which I kept - the scores indicated that the model was very confident a class was/wasn't present in it. I would then use this as a definite label.",
      "votes": null
    },
    {
      "id": "1360239",
      "postDate": "06/21/2021 23:55:54",
      "content": "<p>Congrats for the incredible solo gold! I hope I am not too late to have a few questions regarding your solution:</p>\n<ol>\n<li>would like to know more about how you did training with pseudo labels: e.g. did you apply any soft penalty on your pseudo labels? what was the threshold u set for TPFP? did u apply any sampling scheme to balance TP, FP, psuedo labels (coz in my case I noticed pseudo labels are oversized compared to TP+FP)? How did u validate the yielded pseudo labels make sense? </li>\n<li>for your ResNest approach, did u find 3-channel is significantly different from 1-channel? did u do any normalization on your input (if so, what normalization did u use)? </li>\n</ol>\n<p>Many thanks!</p>",
      "rawMarkdown": "Congrats for the incredible solo gold! I hope I am not too late to have a few questions regarding your solution:\n1. would like to know more about how you did training with pseudo labels: e.g. did you apply any soft penalty on your pseudo labels? what was the threshold u set for TPFP? did u apply any sampling scheme to balance TP, FP, psuedo labels (coz in my case I noticed pseudo labels are oversized compared to TP+FP)? How did u validate the yielded pseudo labels make sense? \n2. for your ResNest approach, did u find 3-channel is significantly different from 1-channel? did u do any normalization on your input (if so, what normalization did u use)? \n\nMany thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1207709,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "02/18/2021 01:35:21",
      "content": "<p>Congrats on 11th place and solo gold medal  <a href=\"https://www.kaggle.com/bigironsphere\" target=\"_blank\">@bigironsphere</a>. Thanks for the writeup. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1207742,
      "author_name": "bigironsphere",
      "author_url": "",
      "post_date": "02/18/2021 02:34:45",
      "content": "<p>I've provided the code for single-channel ResNeSt and its incorporation into the PANNs SED model here:</p>\n<p><a href=\"https://www.kaggle.com/bigironsphere/single-channel-resnest-for-panns-sed-architecture\" target=\"_blank\">https://www.kaggle.com/bigironsphere/single-channel-resnest-for-panns-sed-architecture</a></p>\n<p>I hope someone finds it useful!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1207775,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "02/18/2021 03:17:16",
      "content": "<p>Congratulations on a strong solo finish! You used few submissions and the LB was filled with grandmasters and masters. Very awesome, the loss masking seems key!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1208901,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "02/18/2021 14:51:47",
          "content": "<p>Cheers Corey!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207999,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "02/18/2021 05:35:42",
      "content": "<p>Great result. Congratulations. I went down a lot of the same avenues, just never really put them all together. The process of iteratively making more and more labeled data from pseudolabeling the training data was also mentioned in the birdcall competition but I never got around to it. How important do you think that process was to your result?</p>\n<p>And you mention the 3 different input channels, do you know how significant this was vs just copying the same spectogram representation across all 3 channels?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1208900,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "02/18/2021 14:51:34",
          "content": "<p>I believe it was very important, but I've been through some of my predictions and the pseudolabels were indeed poor for some classes - species 20 looks to be the worst. The most common classes in the data seem to be ones that were quite easy for models to identify, which makes me wonder if this competition would've been more interesting if we were evaluated on class-averaged metrics.</p>\n<p>Good question about the input channels - I never even tested your alternative. Since non-SED models were a significant part of my submission, I just assumed it would be important to include some temporal information via delta and/or delta-delta.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208471,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "02/18/2021 09:49:01",
      "content": "<p>Congratz ! I knew that you had came up with a nice pipeline when you asked about reproductibility. Glad to see you at the top :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209100,
      "author_name": "superchenhao",
      "author_url": "",
      "post_date": "02/18/2021 17:22:37",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/bigironsphere\" target=\"_blank\">@bigironsphere</a> , very smart \"Agnostic Loss\" which fixed partial label problem and leverage fp data meanwhile. </p>\n<blockquote>\n  <p>The figure of 0.5 is arbitrary and could've been any value apart from 0 or 1: the salient point is that the loss resulting from unlabelled classes is always constant. </p>\n</blockquote>\n<p>Regarding above description, after study your \"Agnostic Loss\" and did some calcalation, I feel you have to set target value(y) of unlablled class to 0.5, as definintion of BCEWITHLOGITSLOSS <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.BCEWithLogitsLoss.html#bcewithlogitsloss\" target=\"_blank\">here</a>, for A = sigmoid(x_n), Z = x_n = WX+b, here X stands for input and W stands for weights, then the gradient of weights is(sorry I don`t know how to insert LaTex here): <br>\ndZ = A - Y<br>\ndW = dZ * (A^(L-1))^T</p>\n<p>IMHO your trick is to let <code>$x_n$</code> be 0 thus A = sigmoid(<code>$x_n$</code>) = 0.5, and you set Y of unlabelld classes to 0.5 as well, thus <code>$dZ=0.5-0.5=0$</code>, so <code>$dW$</code>(gradient for unlabelled classes) is always 0, then your model`s weights are only updated for labelled classes when gradient descent. But if you set Y of unlabelld classes to other value than 0.5, then the gradient for unlabelled classes will not be zero, thus your training will affected by unlabelled classes. Please correct me if I am wrong.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209118,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "02/18/2021 17:35:01",
          "content": "<p>0.5 is just a placeholder value. It is just a specific distinction so that later on he can zero out the loss for any sample that had a label of 0.5. He could have set the value to 1000 instead and it would not make a difference because loss of those samples for those classes are never backpropogated </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209392,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/18/2021 21:32:06",
      "content": "<p>Thanks for sharing.  We did something quite similar but you made it work better.  I'll reread more carefully to see where you made a difference, maybe the model (I used efficientnet) or the pseudo labels.</p>\n<p>Congrats on the end result!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209416,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "02/18/2021 21:58:29",
          "content": "<p>Thanks! Using ResNeSt gave me substantially better results. If you want the SED variant, it's <a href=\"https://www.kaggle.com/bigironsphere/single-channel-resnest-for-panns-sed-architecture\" target=\"_blank\">here</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1228436,
      "author_name": "karlyukang",
      "author_url": "",
      "post_date": "03/06/2021 11:30:07",
      "content": "<p>Congrats on the solo gold <a href=\"https://www.kaggle.com/bigironsphere\" target=\"_blank\">@bigironsphere</a> ! <br>\nI have a question:</p>\n<blockquote>\n  <p>isolate the frames which had a very high/low prob across an ensemble of models.</p>\n</blockquote>\n<p>Why you remove those frames with very high/low probs first? Is it because you think these are caused by the uncertainty of the model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1241110,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "03/16/2021 22:24:00",
          "content": "<p>To clarify, those are the frames which I kept - the scores indicated that the model was very confident a class was/wasn't present in it. I would then use this as a definite label.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1360239,
      "author_name": "alexlwh",
      "author_url": "",
      "post_date": "06/21/2021 23:55:54",
      "content": "<p>Congrats for the incredible solo gold! I hope I am not too late to have a few questions regarding your solution:</p>\n<ol>\n<li>would like to know more about how you did training with pseudo labels: e.g. did you apply any soft penalty on your pseudo labels? what was the threshold u set for TPFP? did u apply any sampling scheme to balance TP, FP, psuedo labels (coz in my case I noticed pseudo labels are oversized compared to TP+FP)? How did u validate the yielded pseudo labels make sense? </li>\n<li>for your ResNest approach, did u find 3-channel is significantly different from 1-channel? did u do any normalization on your input (if so, what normalization did u use)? </li>\n</ol>\n<p>Many thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1207705": "Thank you to the organisers, Kaggle, and to everyone who shared ideas and code for this competition. I learned a lot, as I'm sure many of you have, and I thought I would break down my approach since I know many competitors couldn't find a way to use the False Positive labels. I'm thrilled to have secured my first gold (and solo gold) in a competition, so it's the least I could do!\n\n### Summary\n\nOn a high-level, my approach was as follows:\n\n* train models using existing labels\n* generate pseudo-labels on train/test\n* isolate the frames which had a very high/low prob across an ensemble of models\n* conservatively threshold these and use as new TP/FP values for the relevant class\n* repeat\n\nThis gradually increased the amount of data I had to work with, until I had at least 2 frames from each recording with an identified TP and/or FP. The growing data diversity allowed successive models to generalise better.\n\n### What Worked\n\n**Fixed Windows**\n\nI used a window of 5 seconds centred on the TP/FP. For predicting on test, I used overlapping windows with a step of 2.5 seconds. This ensured that train/test had identical preprocessing for their inputs. The maximum value was taken for each class across all of these windows. For some submissions, I used the average of the top 3 predictions but this didn't seem to notably change the LB score.\n\n**Agnostic Loss**\n\nPerhaps a better term already exists for this, but this was what I called my method for using both the TPs and FPs in a semi-supervised fashion. The problem we face with unlabelled data is that any spectrogram can contain multiple classes, so setting the target as 0 for everything apart from the *given* label will penalise true positives for other classes. We can only know for sure that one class is present or absent, and the loss needs to reflect this. So I excluded all non-definite targets from the loss calculation. In the target tensor, a TP is 1 while an FP is 0. Unlabelled classes are given as 0.5. These values are then excluded from the loss calculation. So if we had 5 classes (we have 24, but I'm saving room here) and this time window contained a TP for class 0 and an FP for class 3:\n\n```\ny = torch.Tensor([1., 0.5, 0.5, 0., 0.5]) \n\n```\nAnd in the loss calculation:\n```\npreds = model(inputs)\npreds[targets==0.5] = 0                    \nloss = BCEWithLogitsLoss(preds, targets)\nloss.backward()\n```\nThus the model is 'agnostic' to the majority of the potential labels. This allows the model to build a guided feature representation of the different classes without being inadvertently given false negatives. This approach gave me substantially better LB scores.\n\nThe figure of 0.5 is arbitrary and could've been any value apart from 0 or 1: the salient point is that the loss resulting from unlabelled classes is always constant. Note that this kind of inplace operation is incompatible with `nn.Sigmoid` or its functional equivalent when performing backprop so you need to use the raw logits via `torch.nn.BCEWithLogitsLoss()`.\n\n\n**ResNeSt**\n\nI found EfficientNet to be surprisingly poor in this competition, and all of my best scores came from using variants of ResNeSt (https://github.com/zhanghang1989/ResNeSt) paper available [here](https://arxiv.org/abs/2004.08955).\n\nFor 3-channel input I used the `librosa` mel-spectrogram with power 1, power 2 and the `delta` function to capture temporal information. With some models I experimented with a single power-1 spectrogram, delta and delta-delta features instead. While quicker to preprocess, I noticed no impact on scores.\n\nI also incorporated it into the SED architecture as the encoder. This showed very promising metrics during training, and while sadly I didn't have time to run a fully-convergent example its inclusion still helped my score. In future competitions this could be a very useful model. ResNeSt itself only takes a 3-channel input and has no inbuilt function to extract features, so I had to rejig it to work properly: I'll be uploading a script with that model shortly in case anyone is interested.\n\n**Augmentations**\n\nFrom Hidehisa Arai's excellent kernel [here](https://www.kaggle.com/hidehisaarai1213/rfcx-audio-data-augmentation-japanese-english), I selected `GaussianNoiseSNR()`, `PinkNoiseSNR()`,  `TimeShift()` and `VolumeControl()`. I was wary of augmentation methods that blank out time windows or frequency bands like [SpecAugment](https://arxiv.org/abs/1904.08779). Some of the sounds occur in a very narrow frequency range (e.g. species 21) or in a very narrow time window (e.g. species 9) and I didn't want to make 'empty', positive samples that would coerce the model into learning spurious features. I also added some trivial augmentations of my own:\n\n* swapping the first and last half of the audio vector\n* adding a random constant before spectrogram normalisation (occluding the relevant features)\n* 'jiggling' the time window around the centre of `t_mid`, at a maximum of 1 second offset in either direction\n\n**Eliminating species 19**\n\nA minor point, but this class was so rare that setting all of its predictions to zero usually improved the LB score by 0.001. There were many unrelated sounds that would presumably cause the model to produce false positives. My best submission didn't do this however; it was a simple blend of models including ResNeSt-50, ResNeSt-101, EfficientNet-b1 and the SED architecture described above. I used weighted averaging roughly in proportion to the individual models' performance.\n\n### What didn't work\n\n* Using separate models for different frequency ranges (these models never gained an adequate feature representation, and produced many false positives).\n* EfficientNet alone gave poor results, but helped as part of an ensemble.\n* Larger models (ResNeSt101, EfficientNet b3) didn't improve scores.\n* TP-only training.\n* Models that worked on all windows for a single clip - these were slow and produced inferior results.\n\nOtherwise I was quite lucky - I thought about my methodology for a while and most of what I tried worked well on the first attempt. If I'd had more time, I would have liked to try:\n\n* automatically labelling some species as an FP of a similar class (e.g species 9 & 17)\n* probing the LB for class distribution (I suspect you could get +0.9 by only predicting the most common half of the classes and ignoring everything else) - I realised the importance of this too close to the deadline\n* experimenting with different encoders for the SED architecture.\n* using a smaller window size (<=3 seconds) for greater fidelity. \n\nThe overall class prediction histograms for my final submission were as follows:\n\n![](https://i.imgur.com/GmJ9BkI.png)\n\nSome classes gave me particular trouble. I used my own, simple scoring metric during training that recorded the proportion of positive cases that were predicted above a certain threshold. I never satisfactorily made a model that could detect the rarer classes like 6, 19 or 20 in a reliable fashion.\n\nOverall I had an interesting time exploring how to work with the application of CNNs to spectrograms, and with large amounts of unlabelled data. In the next audio competition, perhaps I'll aim a little higher! I'm looking forward to seeing how those who scored > 0.97 managed to achieve their results.\n\nIf you have any questions I'll do my best to answer them!",
    "1207709": "Congrats on 11th place and solo gold medal  @bigironsphere. Thanks for the writeup.",
    "1207742": "I've provided the code for single-channel ResNeSt and its incorporation into the PANNs SED model here:\n\nhttps://www.kaggle.com/bigironsphere/single-channel-resnest-for-panns-sed-architecture\n\nI hope someone finds it useful!",
    "1207775": "Congratulations on a strong solo finish! You used few submissions and the LB was filled with grandmasters and masters. Very awesome, the loss masking seems key!",
    "1207999": "Great result. Congratulations. I went down a lot of the same avenues, just never really put them all together. The process of iteratively making more and more labeled data from pseudolabeling the training data was also mentioned in the birdcall competition but I never got around to it. How important do you think that process was to your result?\n\nAnd you mention the 3 different input channels, do you know how significant this was vs just copying the same spectogram representation across all 3 channels?",
    "1208471": "Congratz ! I knew that you had came up with a nice pipeline when you asked about reproductibility. Glad to see you at the top :)",
    "1208900": "I believe it was very important, but I've been through some of my predictions and the pseudolabels were indeed poor for some classes - species 20 looks to be the worst. The most common classes in the data seem to be ones that were quite easy for models to identify, which makes me wonder if this competition would've been more interesting if we were evaluated on class-averaged metrics.\n\nGood question about the input channels - I never even tested your alternative. Since non-SED models were a significant part of my submission, I just assumed it would be important to include some temporal information via delta and/or delta-delta.",
    "1208901": "Cheers Corey!",
    "1209100": "Congratulations @bigironsphere , very smart \"Agnostic Loss\" which fixed partial label problem and leverage fp data meanwhile. \n> The figure of 0.5 is arbitrary and could've been any value apart from 0 or 1: the salient point is that the loss resulting from unlabelled classes is always constant. \n\nRegarding above description, after study your \"Agnostic Loss\" and did some calcalation, I feel you have to set target value(y) of unlablled class to 0.5, as definintion of BCEWITHLOGITSLOSS [here](https://pytorch.org/docs/stable/generated/torch.nn.BCEWithLogitsLoss.html#bcewithlogitsloss), for A = sigmoid(x_n), Z = x_n = WX+b, here X stands for input and W stands for weights, then the gradient of weights is(sorry I don`t know how to insert LaTex here): \ndZ = A - Y\ndW = dZ * (A^(L-1))^T\n\nIMHO your trick is to let `$x_n$` be 0 thus A = sigmoid(`$x_n$`) = 0.5, and you set Y of unlabelld classes to 0.5 as well, thus `$dZ=0.5-0.5=0$`, so `$dW$`(gradient for unlabelled classes) is always 0, then your model`s weights are only updated for labelled classes when gradient descent. But if you set Y of unlabelld classes to other value than 0.5, then the gradient for unlabelled classes will not be zero, thus your training will affected by unlabelled classes. Please correct me if I am wrong.",
    "1209118": "0.5 is just a placeholder value. It is just a specific distinction so that later on he can zero out the loss for any sample that had a label of 0.5. He could have set the value to 1000 instead and it would not make a difference because loss of those samples for those classes are never backpropogated",
    "1209392": "Thanks for sharing.  We did something quite similar but you made it work better.  I'll reread more carefully to see where you made a difference, maybe the model (I used efficientnet) or the pseudo labels.\n\nCongrats on the end result!",
    "1209416": "Thanks! Using ResNeSt gave me substantially better results. If you want the SED variant, it's [here](https://www.kaggle.com/bigironsphere/single-channel-resnest-for-panns-sed-architecture)",
    "1228436": "Congrats on the solo gold @bigironsphere ! \nI have a question:\n\n> isolate the frames which had a very high/low prob across an ensemble of models.\n\nWhy you remove those frames with very high/low probs first? Is it because you think these are caused by the uncertainty of the model?",
    "1241110": "To clarify, those are the frames which I kept - the scores indicated that the model was very confident a class was/wasn't present in it. I would then use this as a definite label.",
    "1360239": "Congrats for the incredible solo gold! I hope I am not too late to have a few questions regarding your solution:\n1. would like to know more about how you did training with pseudo labels: e.g. did you apply any soft penalty on your pseudo labels? what was the threshold u set for TPFP? did u apply any sampling scheme to balance TP, FP, psuedo labels (coz in my case I noticed pseudo labels are oversized compared to TP+FP)? How did u validate the yielded pseudo labels make sense? \n2. for your ResNest approach, did u find 3-channel is significantly different from 1-channel? did u do any normalization on your input (if so, what normalization did u use)? \n\nMany thanks!"
  },
  "source": "meta"
}