{
  "id": 220308,
  "title": "13th Place Solution – Mean Co-Teachers and Noisy Students",
  "url": "/competitions/rfcx-species-audio-detection/discussion/220308",
  "author_name": "Ryan Epp",
  "post_date": "2021-02-18T00:12:20.127000",
  "votes": 42,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hey Everybody, I wanted to dump my solution real quick in case anyone was interested. </p>\n<p>It seemed to me that the critical issue is that there are a <strong>TON</strong> of missing labels. The provided positive examples data (train_tp.csv) has ~1.2k labels. The lb probing that <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> did suggest 4-5  labels per clip on average. If the train data follows the same distribution we should expect ~21k labels,  and that's just at the <em>clip</em> level. We'd expect to see multiple calls from the same bird per clip, i.e. multiple frame labels per clip label. My best models seemed to think there were closer to 40k labels.</p>\n<p>So my idea was to do something along the lines of <a href=\"https://arxiv.org/abs/1911.04252\" target=\"_blank\">Noisy Student</a>, where the general idea is to do progressive pseudo labeling where each successive model is larger and there's more noise applied to the training data. On its own, Noisy Student doesn't work very well, so I used a few other tricks. </p>\n<h3>1. Mean Teacher</h3>\n<p>My first setup looks super similar to what's going on in  <a href=\"https://www.kaggle.com/reppic/mean-teachers-find-more-birds\" target=\"_blank\">Mean Teachers Find More Birds</a>. I train on a combo of centered positive examples and random unlabeled samples using consistency and BCE loss. Here, I'm using SED + resnet34 and some light augmentation: gaussian noise, frame/frequency dropout. This gets me to <strong>0.865</strong>on the public lb.</p>\n<p>Using 5-fold mean-teacher models, I do OOF prediction to get pseudo labels over the entire training dataset. </p>\n<h3>2. Co-Teaching</h3>\n<p>Now I want to train on my pseudo labels, but it's safe to assume they're pretty noisy. To deal with the bias introduced by my new, noisy labels, I do something along the lines of <a href=\"https://arxiv.org/abs/1804.06872\" target=\"_blank\">Co-Teaching</a>. Briefly, the idea is to train 2 models simultaneously on the same data, but with different augmentations applied to each. Then the samples with the highest loss from Model A are ignored when doing backprop in Model B and vice versa. The % of ignored samples gets ramped up slowly. The theory is that the models will learn the correct labels early in training and start to overfit to noise later on. By dropping potentially noisy labels, we avoid introducing a bad bias from our pseudo labels. </p>\n<p>I modified the authors idea slightly for the competition. In my setup, it's impossible for either model to ignore the good labels from train_tp or train_fp. Only pseudo labels can be ignored. I believe this helps with class imbalance issues. </p>\n<p>Using this setup with more aggressive augmentation and densenet 121, I'm able to get to <strong>0.906</strong> on the public lb. </p>\n<h3>3. Heavy Mixup</h3>\n<p>Finally, using my second round of pseudo labels, I train on randomly sampled segments from all the training data. Here I apply even more aggressive augmentations and add mixup 60% of the time with a mixing weight sampled from <code>Beta(5,5)</code> (typically around 0.5). For mixup, any label present in either clip gets set to 1.0. I run this for 80 epochs. The prev 2 models were run for around 32 epochs. A 5 fold ensemble with this setup using densenet 121 gets me up to <strong>0.940</strong> on the public lb.</p>\n<p>I’m able to get to 0.943 by ensembling ~90 models taking the geometric mean.</p>\n<h3>Other Tricks</h3>\n<ul>\n<li>Centering the labels from train_tp in the sampled clip segment early on seemed to help.</li>\n<li>When making predictions I’m averaging 4 metrics: average and max clip-wise and frame-wise predictions.</li>\n<li>Mixup only worked for me when it was done on the log mel spectrograms. Doing it on the audio didn't work.</li>\n<li>Augmentations (intensities varied) (excluding mixup): </li>\n</ul>\n<pre><code>augmenter = A.Compose([\n    A.AddGaussianNoise(p=0.5, max_amplitude=0.033),\n    A.AddGaussianSNR(p=0.5),\n    A.FrequencyMask(min_frequency_band=0.01, max_frequency_band=0.5, p=0.5), \n    A.TimeMask(min_band_part=0.01, max_band_part=0.5, p=0.5),\n    A.Gain(p=0.5)\n])\n</code></pre>\n<p>Let me know if you have any questions! </p>",
  "messages": [
    {
      "id": 1207605,
      "postDate": "2021-02-18T00:12:20.127Z",
      "content": "<p>Hey Everybody, I wanted to dump my solution real quick in case anyone was interested. </p>\n<p>It seemed to me that the critical issue is that there are a <strong>TON</strong> of missing labels. The provided positive examples data (train_tp.csv) has ~1.2k labels. The lb probing that <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> did suggest 4-5  labels per clip on average. If the train data follows the same distribution we should expect ~21k labels,  and that's just at the <em>clip</em> level. We'd expect to see multiple calls from the same bird per clip, i.e. multiple frame labels per clip label. My best models seemed to think there were closer to 40k labels.</p>\n<p>So my idea was to do something along the lines of <a href=\"https://arxiv.org/abs/1911.04252\" target=\"_blank\">Noisy Student</a>, where the general idea is to do progressive pseudo labeling where each successive model is larger and there's more noise applied to the training data. On its own, Noisy Student doesn't work very well, so I used a few other tricks. </p>\n<h3>1. Mean Teacher</h3>\n<p>My first setup looks super similar to what's going on in  <a href=\"https://www.kaggle.com/reppic/mean-teachers-find-more-birds\" target=\"_blank\">Mean Teachers Find More Birds</a>. I train on a combo of centered positive examples and random unlabeled samples using consistency and BCE loss. Here, I'm using SED + resnet34 and some light augmentation: gaussian noise, frame/frequency dropout. This gets me to <strong>0.865</strong>on the public lb.</p>\n<p>Using 5-fold mean-teacher models, I do OOF prediction to get pseudo labels over the entire training dataset. </p>\n<h3>2. Co-Teaching</h3>\n<p>Now I want to train on my pseudo labels, but it's safe to assume they're pretty noisy. To deal with the bias introduced by my new, noisy labels, I do something along the lines of <a href=\"https://arxiv.org/abs/1804.06872\" target=\"_blank\">Co-Teaching</a>. Briefly, the idea is to train 2 models simultaneously on the same data, but with different augmentations applied to each. Then the samples with the highest loss from Model A are ignored when doing backprop in Model B and vice versa. The % of ignored samples gets ramped up slowly. The theory is that the models will learn the correct labels early in training and start to overfit to noise later on. By dropping potentially noisy labels, we avoid introducing a bad bias from our pseudo labels. </p>\n<p>I modified the authors idea slightly for the competition. In my setup, it's impossible for either model to ignore the good labels from train_tp or train_fp. Only pseudo labels can be ignored. I believe this helps with class imbalance issues. </p>\n<p>Using this setup with more aggressive augmentation and densenet 121, I'm able to get to <strong>0.906</strong> on the public lb. </p>\n<h3>3. Heavy Mixup</h3>\n<p>Finally, using my second round of pseudo labels, I train on randomly sampled segments from all the training data. Here I apply even more aggressive augmentations and add mixup 60% of the time with a mixing weight sampled from <code>Beta(5,5)</code> (typically around 0.5). For mixup, any label present in either clip gets set to 1.0. I run this for 80 epochs. The prev 2 models were run for around 32 epochs. A 5 fold ensemble with this setup using densenet 121 gets me up to <strong>0.940</strong> on the public lb.</p>\n<p>I’m able to get to 0.943 by ensembling ~90 models taking the geometric mean.</p>\n<h3>Other Tricks</h3>\n<ul>\n<li>Centering the labels from train_tp in the sampled clip segment early on seemed to help.</li>\n<li>When making predictions I’m averaging 4 metrics: average and max clip-wise and frame-wise predictions.</li>\n<li>Mixup only worked for me when it was done on the log mel spectrograms. Doing it on the audio didn't work.</li>\n<li>Augmentations (intensities varied) (excluding mixup): </li>\n</ul>\n<pre><code>augmenter = A.Compose([\n    A.AddGaussianNoise(p=0.5, max_amplitude=0.033),\n    A.AddGaussianSNR(p=0.5),\n    A.FrequencyMask(min_frequency_band=0.01, max_frequency_band=0.5, p=0.5), \n    A.TimeMask(min_band_part=0.01, max_band_part=0.5, p=0.5),\n    A.Gain(p=0.5)\n])\n</code></pre>\n<p>Let me know if you have any questions! </p>",
      "rawMarkdown": "Hey Everybody, I wanted to dump my solution real quick in case anyone was interested. \n\nIt seemed to me that the critical issue is that there are a **TON** of missing labels. The provided positive examples data (train_tp.csv) has ~1.2k labels. The lb probing that @cpmpml did suggest 4-5  labels per clip on average. If the train data follows the same distribution we should expect ~21k labels,  and that's just at the *clip* level. We'd expect to see multiple calls from the same bird per clip, i.e. multiple frame labels per clip label. My best models seemed to think there were closer to 40k labels.\n\nSo my idea was to do something along the lines of [Noisy Student](https://arxiv.org/abs/1911.04252), where the general idea is to do progressive pseudo labeling where each successive model is larger and there's more noise applied to the training data. On its own, Noisy Student doesn't work very well, so I used a few other tricks. \n\n### 1. Mean Teacher\nMy first setup looks super similar to what's going on in  [Mean Teachers Find More Birds](https://www.kaggle.com/reppic/mean-teachers-find-more-birds). I train on a combo of centered positive examples and random unlabeled samples using consistency and BCE loss. Here, I'm using SED + resnet34 and some light augmentation: gaussian noise, frame/frequency dropout. This gets me to **0.865**on the public lb.\n\nUsing 5-fold mean-teacher models, I do OOF prediction to get pseudo labels over the entire training dataset. \n\n### 2. Co-Teaching\nNow I want to train on my pseudo labels, but it's safe to assume they're pretty noisy. To deal with the bias introduced by my new, noisy labels, I do something along the lines of [Co-Teaching](https://arxiv.org/abs/1804.06872). Briefly, the idea is to train 2 models simultaneously on the same data, but with different augmentations applied to each. Then the samples with the highest loss from Model A are ignored when doing backprop in Model B and vice versa. The % of ignored samples gets ramped up slowly. The theory is that the models will learn the correct labels early in training and start to overfit to noise later on. By dropping potentially noisy labels, we avoid introducing a bad bias from our pseudo labels. \n\nI modified the authors idea slightly for the competition. In my setup, it's impossible for either model to ignore the good labels from train_tp or train_fp. Only pseudo labels can be ignored. I believe this helps with class imbalance issues. \n\nUsing this setup with more aggressive augmentation and densenet 121, I'm able to get to **0.906** on the public lb. \n\n### 3. Heavy Mixup\nFinally, using my second round of pseudo labels, I train on randomly sampled segments from all the training data. Here I apply even more aggressive augmentations and add mixup 60% of the time with a mixing weight sampled from `Beta(5,5)` (typically around 0.5). For mixup, any label present in either clip gets set to 1.0. I run this for 80 epochs. The prev 2 models were run for around 32 epochs. A 5 fold ensemble with this setup using densenet 121 gets me up to **0.940** on the public lb.\n\nI’m able to get to 0.943 by ensembling ~90 models taking the geometric mean.\n\n### Other Tricks\n* Centering the labels from train_tp in the sampled clip segment early on seemed to help.\n* When making predictions I’m averaging 4 metrics: average and max clip-wise and frame-wise predictions.\n* Mixup only worked for me when it was done on the log mel spectrograms. Doing it on the audio didn't work.\n* Augmentations (intensities varied) (excluding mixup): \n```\naugmenter = A.Compose([\n    A.AddGaussianNoise(p=0.5, max_amplitude=0.033),\n    A.AddGaussianSNR(p=0.5),\n    A.FrequencyMask(min_frequency_band=0.01, max_frequency_band=0.5, p=0.5), \n    A.TimeMask(min_band_part=0.01, max_band_part=0.5, p=0.5),\n    A.Gain(p=0.5)\n])\n```\n\nLet me know if you have any questions! ",
      "votes": 40
    },
    {
      "id": 1208484,
      "postDate": "2021-02-18T09:51:38.183Z",
      "content": "<p>Very nice solution, congratz ! I find it really cool to see research papers work in practice.</p>",
      "rawMarkdown": "Very nice solution, congratz ! I find it really cool to see research papers work in practice.",
      "votes": 1
    },
    {
      "id": 1207668,
      "postDate": "2021-02-18T00:52:59.577Z",
      "content": "<p>Congrats on strong finish <a href=\"https://www.kaggle.com/reppic\" target=\"_blank\">@reppic</a> and thanks for sharing details solution </p>",
      "rawMarkdown": "Congrats on strong finish @reppic and thanks for sharing details solution ",
      "votes": 1
    },
    {
      "id": 1207665,
      "postDate": "2021-02-18T00:51:05.537Z",
      "content": "<p>Thanks for your write-up. I have tried your notebook. I was able to get 0.891 from a mixnet model with mixup and grid-cutout augmentation added.</p>",
      "rawMarkdown": "Thanks for your write-up. I have tried your notebook. I was able to get 0.891 from a mixnet model with mixup and grid-cutout augmentation added.",
      "votes": 1
    },
    {
      "id": 1210379,
      "postDate": "2021-02-19T11:39:56.910Z",
      "content": "<p>Well done!  You are probably disappointed to miss gold by one rank. I've been there, it is mixed feelings.  Keep on, I'm sure you'll get solid gold soon.</p>\n<p>And thanks for the mean teacher reference.</p>",
      "rawMarkdown": "Well done!  You are probably disappointed to miss gold by one rank. I've been there, it is mixed feelings.  Keep on, I'm sure you'll get solid gold soon.\n\nAnd thanks for the mean teacher reference."
    },
    {
      "id": 1208006,
      "postDate": "2021-02-18T05:42:12.533Z",
      "content": "<p>Thanks for the write-up the coteaching approach is new to me. We ended up using the mean teacher approach you shared pretty extensively so can give a +1 that it seems to work well. Surprised that technique does so well. People have been doing stuff like that in the ranzcr catheter competition as well. </p>\n<p>It is kind of similar to contrastive learning it seems. </p>\n<p>Was mixup really enough to boost you from .906 to .940? That is a surprising amount. I agree with the idea of making it so any label present is treated as a 1 rather than blending. The blended spectrogram input is just like one of the clips is slightly quieter than the other so doesn't make sense to make it only a partial label. </p>",
      "rawMarkdown": "Thanks for the write-up the coteaching approach is new to me. We ended up using the mean teacher approach you shared pretty extensively so can give a +1 that it seems to work well. Surprised that technique does so well. People have been doing stuff like that in the ranzcr catheter competition as well. \n\nIt is kind of similar to contrastive learning it seems. \n\nWas mixup really enough to boost you from .906 to .940? That is a surprising amount. I agree with the idea of making it so any label present is treated as a 1 rather than blending. The blended spectrogram input is just like one of the clips is slightly quieter than the other so doesn't make sense to make it only a partial label. "
    },
    {
      "id": 1207814,
      "postDate": "2021-02-18T03:50:56.943Z",
      "content": "<p>Thank you a lot. I also somewhat used ideas from noisy student which was inspired by your mean-teacher post (I read your approach and jumped to that paper).</p>",
      "rawMarkdown": "Thank you a lot. I also somewhat used ideas from noisy student which was inspired by your mean-teacher post (I read your approach and jumped to that paper)."
    },
    {
      "id": 1207768,
      "postDate": "2021-02-18T03:03:53.997Z",
      "content": "<p>Great solution!</p>\n<p>I also tried Lq and <a href=\"https://arxiv.org/abs/2007.00151\" target=\"_blank\">ELR</a> for missing labels. But these did not worked.<br>\nCo-teaching seem to be nice.</p>",
      "rawMarkdown": "Great solution!\n\nI also tried Lq and [ELR](https://arxiv.org/abs/2007.00151) for missing labels. But these did not worked.\nCo-teaching seem to be nice."
    },
    {
      "id": 1207724,
      "postDate": "2021-02-18T02:04:19.957Z",
      "content": "<p>Congratz <a href=\"https://www.kaggle.com/reppic\" target=\"_blank\">@reppic</a>!<br>\nI have tried your SED notebook which helped me to figure out my problems! Thank you so much and well deserved!</p>",
      "rawMarkdown": "Congratz @reppic!\nI have tried your SED notebook which helped me to figure out my problems! Thank you so much and well deserved!"
    },
    {
      "id": 1207706,
      "postDate": "2021-02-18T01:32:57.390Z",
      "content": "<p>\"Mixup only worked for me when it was done on the log mel spectrograms. Doing it on the audio didn't work\"</p>\n<p>apply mixup on fp annotation (negative samples only). this creates more negative samples by mixing negative samples. you can do this in spectrograms.</p>",
      "rawMarkdown": "\"Mixup only worked for me when it was done on the log mel spectrograms. Doing it on the audio didn't work\"\n\napply mixup on fp annotation (negative samples only). this creates more negative samples by mixing negative samples. you can do this in spectrograms."
    },
    {
      "id": 1207703,
      "postDate": "2021-02-18T01:30:54.167Z",
      "content": "<p>actually one should look at LWLRAP as top-K accuracy metric.<br>\nwe know that each clip on average have at least 4 (or 3 from the paper) labels.</p>\n<p>if you are having LWLRAP of about 0.9, you are having better than 0.90 for your top-1 guess.<br>\ntop-1, top-2 predictions are very good for pesudo labels.</p>",
      "rawMarkdown": "actually one should look at LWLRAP as top-K accuracy metric.\nwe know that each clip on average have at least 4 (or 3 from the paper) labels.\n\nif you are having LWLRAP of about 0.9, you are having better than 0.90 for your top-1 guess.\ntop-1, top-2 predictions are very good for pesudo labels."
    },
    {
      "id": 1208107,
      "postDate": "2021-02-18T06:56:16.373Z",
      "content": "<p>Thanks for sharing! I tried similar mixup approach borrowed from the Cornell 3rd solution. But for me it performed worse than no mixup and the vanilla mixup. </p>",
      "rawMarkdown": "Thanks for sharing! I tried similar mixup approach borrowed from the Cornell 3rd solution. But for me it performed worse than no mixup and the vanilla mixup. ",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1208484,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2021-02-18T09:51:38.183000",
      "content": "<p>Very nice solution, congratz ! I find it really cool to see research papers work in practice.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1207668,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2021-02-18T00:52:59.577000",
      "content": "<p>Congrats on strong finish <a href=\"https://www.kaggle.com/reppic\" target=\"_blank\">@reppic</a> and thanks for sharing details solution </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1207665,
      "author_name": "Shai",
      "author_url": "",
      "post_date": "2021-02-18T00:51:05.537000",
      "content": "<p>Thanks for your write-up. I have tried your notebook. I was able to get 0.891 from a mixnet model with mixup and grid-cutout augmentation added.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1210379,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-02-19T11:39:56.910000",
      "content": "<p>Well done!  You are probably disappointed to miss gold by one rank. I've been there, it is mixed feelings.  Keep on, I'm sure you'll get solid gold soon.</p>\n<p>And thanks for the mean teacher reference.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1208006,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2021-02-18T05:42:12.533000",
      "content": "<p>Thanks for the write-up the coteaching approach is new to me. We ended up using the mean teacher approach you shared pretty extensively so can give a +1 that it seems to work well. Surprised that technique does so well. People have been doing stuff like that in the ranzcr catheter competition as well. </p>\n<p>It is kind of similar to contrastive learning it seems. </p>\n<p>Was mixup really enough to boost you from .906 to .940? That is a surprising amount. I agree with the idea of making it so any label present is treated as a 1 rather than blending. The blended spectrogram input is just like one of the clips is slightly quieter than the other so doesn't make sense to make it only a partial label. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1207814,
      "author_name": "dan",
      "author_url": "",
      "post_date": "2021-02-18T03:50:56.943000",
      "content": "<p>Thank you a lot. I also somewhat used ideas from noisy student which was inspired by your mean-teacher post (I read your approach and jumped to that paper).</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1207768,
      "author_name": "shinmura0",
      "author_url": "",
      "post_date": "2021-02-18T03:03:53.997000",
      "content": "<p>Great solution!</p>\n<p>I also tried Lq and <a href=\"https://arxiv.org/abs/2007.00151\" target=\"_blank\">ELR</a> for missing labels. But these did not worked.<br>\nCo-teaching seem to be nice.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1207724,
      "author_name": "Bayartsogt Yadamsuren",
      "author_url": "",
      "post_date": "2021-02-18T02:04:19.957000",
      "content": "<p>Congratz <a href=\"https://www.kaggle.com/reppic\" target=\"_blank\">@reppic</a>!<br>\nI have tried your SED notebook which helped me to figure out my problems! Thank you so much and well deserved!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1207706,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-02-18T01:32:57.390000",
      "content": "<p>\"Mixup only worked for me when it was done on the log mel spectrograms. Doing it on the audio didn't work\"</p>\n<p>apply mixup on fp annotation (negative samples only). this creates more negative samples by mixing negative samples. you can do this in spectrograms.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1207703,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-02-18T01:30:54.167000",
      "content": "<p>actually one should look at LWLRAP as top-K accuracy metric.<br>\nwe know that each clip on average have at least 4 (or 3 from the paper) labels.</p>\n<p>if you are having LWLRAP of about 0.9, you are having better than 0.90 for your top-1 guess.<br>\ntop-1, top-2 predictions are very good for pesudo labels.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1208107,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-18T06:56:16.373000",
      "content": "<p>Thanks for sharing! I tried similar mixup approach borrowed from the Cornell 3rd solution. But for me it performed worse than no mixup and the vanilla mixup. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1207605": "Hey Everybody, I wanted to dump my solution real quick in case anyone was interested. \n\nIt seemed to me that the critical issue is that there are a **TON** of missing labels. The provided positive examples data (train_tp.csv) has ~1.2k labels. The lb probing that @cpmpml did suggest 4-5  labels per clip on average. If the train data follows the same distribution we should expect ~21k labels,  and that's just at the *clip* level. We'd expect to see multiple calls from the same bird per clip, i.e. multiple frame labels per clip label. My best models seemed to think there were closer to 40k labels.\n\nSo my idea was to do something along the lines of [Noisy Student](https://arxiv.org/abs/1911.04252), where the general idea is to do progressive pseudo labeling where each successive model is larger and there's more noise applied to the training data. On its own, Noisy Student doesn't work very well, so I used a few other tricks. \n\n### 1. Mean Teacher\nMy first setup looks super similar to what's going on in  [Mean Teachers Find More Birds](https://www.kaggle.com/reppic/mean-teachers-find-more-birds). I train on a combo of centered positive examples and random unlabeled samples using consistency and BCE loss. Here, I'm using SED + resnet34 and some light augmentation: gaussian noise, frame/frequency dropout. This gets me to **0.865**on the public lb.\n\nUsing 5-fold mean-teacher models, I do OOF prediction to get pseudo labels over the entire training dataset. \n\n### 2. Co-Teaching\nNow I want to train on my pseudo labels, but it's safe to assume they're pretty noisy. To deal with the bias introduced by my new, noisy labels, I do something along the lines of [Co-Teaching](https://arxiv.org/abs/1804.06872). Briefly, the idea is to train 2 models simultaneously on the same data, but with different augmentations applied to each. Then the samples with the highest loss from Model A are ignored when doing backprop in Model B and vice versa. The % of ignored samples gets ramped up slowly. The theory is that the models will learn the correct labels early in training and start to overfit to noise later on. By dropping potentially noisy labels, we avoid introducing a bad bias from our pseudo labels. \n\nI modified the authors idea slightly for the competition. In my setup, it's impossible for either model to ignore the good labels from train_tp or train_fp. Only pseudo labels can be ignored. I believe this helps with class imbalance issues. \n\nUsing this setup with more aggressive augmentation and densenet 121, I'm able to get to **0.906** on the public lb. \n\n### 3. Heavy Mixup\nFinally, using my second round of pseudo labels, I train on randomly sampled segments from all the training data. Here I apply even more aggressive augmentations and add mixup 60% of the time with a mixing weight sampled from `Beta(5,5)` (typically around 0.5). For mixup, any label present in either clip gets set to 1.0. I run this for 80 epochs. The prev 2 models were run for around 32 epochs. A 5 fold ensemble with this setup using densenet 121 gets me up to **0.940** on the public lb.\n\nI’m able to get to 0.943 by ensembling ~90 models taking the geometric mean.\n\n### Other Tricks\n* Centering the labels from train_tp in the sampled clip segment early on seemed to help.\n* When making predictions I’m averaging 4 metrics: average and max clip-wise and frame-wise predictions.\n* Mixup only worked for me when it was done on the log mel spectrograms. Doing it on the audio didn't work.\n* Augmentations (intensities varied) (excluding mixup): \n```\naugmenter = A.Compose([\n    A.AddGaussianNoise(p=0.5, max_amplitude=0.033),\n    A.AddGaussianSNR(p=0.5),\n    A.FrequencyMask(min_frequency_band=0.01, max_frequency_band=0.5, p=0.5), \n    A.TimeMask(min_band_part=0.01, max_band_part=0.5, p=0.5),\n    A.Gain(p=0.5)\n])\n```\n\nLet me know if you have any questions! ",
    "1208484": "Very nice solution, congratz ! I find it really cool to see research papers work in practice.",
    "1207668": "Congrats on strong finish @reppic and thanks for sharing details solution ",
    "1207665": "Thanks for your write-up. I have tried your notebook. I was able to get 0.891 from a mixnet model with mixup and grid-cutout augmentation added.",
    "1210379": "Well done!  You are probably disappointed to miss gold by one rank. I've been there, it is mixed feelings.  Keep on, I'm sure you'll get solid gold soon.\n\nAnd thanks for the mean teacher reference.",
    "1208006": "Thanks for the write-up the coteaching approach is new to me. We ended up using the mean teacher approach you shared pretty extensively so can give a +1 that it seems to work well. Surprised that technique does so well. People have been doing stuff like that in the ranzcr catheter competition as well. \n\nIt is kind of similar to contrastive learning it seems. \n\nWas mixup really enough to boost you from .906 to .940? That is a surprising amount. I agree with the idea of making it so any label present is treated as a 1 rather than blending. The blended spectrogram input is just like one of the clips is slightly quieter than the other so doesn't make sense to make it only a partial label. ",
    "1207814": "Thank you a lot. I also somewhat used ideas from noisy student which was inspired by your mean-teacher post (I read your approach and jumped to that paper).",
    "1207768": "Great solution!\n\nI also tried Lq and [ELR](https://arxiv.org/abs/2007.00151) for missing labels. But these did not worked.\nCo-teaching seem to be nice.",
    "1207724": "Congratz @reppic!\nI have tried your SED notebook which helped me to figure out my problems! Thank you so much and well deserved!",
    "1207706": "\"Mixup only worked for me when it was done on the log mel spectrograms. Doing it on the audio didn't work\"\n\napply mixup on fp annotation (negative samples only). this creates more negative samples by mixing negative samples. you can do this in spectrograms.",
    "1207703": "actually one should look at LWLRAP as top-K accuracy metric.\nwe know that each clip on average have at least 4 (or 3 from the paper) labels.\n\nif you are having LWLRAP of about 0.9, you are having better than 0.90 for your top-1 guess.\ntop-1, top-2 predictions are very good for pesudo labels.",
    "1208107": "Thanks for sharing! I tried similar mixup approach borrowed from the Cornell 3rd solution. But for me it performed worse than no mixup and the vanilla mixup. "
  }
}