{
  "id": 220322,
  "title": "32nd place solution - what worked for me",
  "url": "/competitions/rfcx-species-audio-detection/writeups/yoonsoo-32nd-place-solution-what-worked-for-me",
  "author_name": "",
  "post_date": "2021-02-19T00:26:27.187Z",
  "votes": 21,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Congratulations to the top finishers!</p>\n<p>This was my first encounter to audio competition, so I tried a lot of maybe implausible ideas and learned a lot. Especially, solutions from Cornell Birdcall Identification and Freesound Audio Tagging 2019 were helpful.</p>\n<p>My finish is not strong, but I wanted to share some of the things that I <em>believe(not sure since my score is not sufficiently high)</em> worked for me (increased cv or public score), and hear what other kagglers experienced.</p>\n<ul>\n<li><p>Frequency Crop</p>\n<ul>\n<li>For one audio clip, crop 24 different crops according to fmin&amp;fmax of each species.</li>\n<li>I believe it is similar to what <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> did and <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> did(without repeating convolution computations)</li></ul></li>\n<li><p>LSoft</p>\n<ul>\n<li><p>I used only TP and FP crops as labels. For example, for each row in TP or FP, only 1 out of 24 labels is present.</p></li>\n<li><p>Based on BCELoss, I used 'lsoft' for unknown labels. LSoft is introduced kindly by <a href=\"https://www.kaggle.com/romul0212\" target=\"_blank\">@romul0212</a> at <a href=\"https://github.com/lRomul/argus-freesound/blob/master/src/losses.py\" target=\"_blank\">https://github.com/lRomul/argus-freesound/blob/master/src/losses.py</a></p></li>\n<li><p>This is my loss computation with LSoft. <code>mask</code> indicates where the label is known. <code>true</code> for unknown labels is initialized with 0.</p></li></ul>\n<pre><code>tmp_true = (1 - lsoft) * true + lsoft * torch.sigmoid(pred)\ntrue = torch.where(masks == 0, tmp_true, true)\nloss = nn.BCEWithLogitsLoss()(pred, true)\n</code></pre></li>\n<li><p>Iterative Pseudo Training</p>\n<ul>\n<li>Since train set is only very sparsely annotated, I thought re-labeling with the model then re-training will help, and it indeed helped. I pseudo trained for 3 stages.</li>\n<li>When pseudo training, I didn't use LSoft and used vanilla BCE.</li></ul></li>\n<li><p>LSEPLoss</p>\n<ul>\n<li><p>Our metric is LWLRAP, so it is important to focus on rank between labels for each row. I used LSepLoss which fits this purpose, which is introduced kindly by <a href=\"https://www.kaggle.com/ddanevskyi\" target=\"_blank\">@ddanevskyi</a> at <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97926\" target=\"_blank\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97926</a></p></li>\n<li><p>After stage3 of BCE pseudo training, I pseudo trained extra 2 stages with LSEPLoss</p></li>\n<li><p>I fixed original code a bit to allow soft labels.</p></li></ul>\n<pre><code>def lsep_loss(input, target):\n    input_differences = input.unsqueeze(1) - input.unsqueeze(2)\n    target_differences = target.unsqueeze(2) - target.unsqueeze(1)\n    target_differences = torch.maximum(torch.tensor(0).to(input.device), target_differences)\n    exps = input_differences.exp() * target_differences\n    lsep = torch.log(1 + exps.sum(2).sum(1))\n    return lsep.mean()\n</code></pre></li>\n\n<li><p>Global Average Pooling on only positive values</p>\n<ul>\n<li><p>We need to know if the species is present or not. We don't care if it appears frequently or not. I thought doing global average pooling on whole last feature map of CNN will yield high probabilities for frequent occurrences of birdcall in one clip and low probabilities for infrequent occurrences, which doesn't match our goal. So I took mean of only positive values from the last feature map of CNN.</p></li>\n<li><p>Following code is attached at the end of CNN's extracted feature map</p></li></ul>\n<pre><code>mask = (x &gt; 0).float()\nfeatures = (x*mask).sum(dim=(2, 3))/(torch.maximum(mask.sum(dim=(2, 3)), torch.tensor(1e-8).to(mask.device)))\n</code></pre></li>\n<li><p>Augmentations</p>\n<ul>\n<li>Gaussian/Pink NoiseSNR, PitchShift, TimeStretch, TimeShift, VolumeControl, Mixup(take union of the labels), SpecAugment</li>\n<li>Thanks to <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> for kindly sharing <a href=\"https://www.kaggle.com/hidehisaarai1213/rfcx-audio-data-augmentation-japanese-english\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/rfcx-audio-data-augmentation-japanese-english</a></li></ul></li>\n<li><p>One inference on full clip</p>\n<ul>\n<li>I didn't resize the spectrogram, so I was able to train on crops and infer on full image.</li>\n<li>When we don't resize, due to the property of CNN, I believe doing sliding windows prediction on small crops is just an approximation for doing one inference on the full image.</li></ul></li>\n<li><p>Validation - only use known labels</p>\n<ul>\n<li>I did validation clip-wise, only on TP and FP labels. From the prediction, I removed all values corresponding to unknown labels, flattened, then calculated LWLRAP. It correlated with LB quite well on my fold0</li></ul></li>\n</ul>\n<p>My baseline was not so strong(~0.8), so I might had fundamental mistakes in my baseline.<br>\nI achieved 0.927 public with efficientnet-b0 fold0 3seed average, but my score worsened when doing 5fold ensembling. I tried average, rank mean, scaling on axis1 then taking mean, calculating mean of pairwise differences taking average, but it didn't help.<br>\nI'm planning to study top solutions to find out what I missed</p>\n<p>I'd really appreciate it if you share some opinions with my approaches and things that I missed.</p>",
  "messages": [
    {
      "id": "1207727",
      "postDate": "02/18/2021 02:08:56",
      "content": "<p>Congratulations to the top finishers!</p>\n<p>This was my first encounter to audio competition, so I tried a lot of maybe implausible ideas and learned a lot. Especially, solutions from Cornell Birdcall Identification and Freesound Audio Tagging 2019 were helpful.</p>\n<p>My finish is not strong, but I wanted to share some of the things that I <em>believe(not sure since my score is not sufficiently high)</em> worked for me (increased cv or public score), and hear what other kagglers experienced.</p>\n<ul>\n<li><p>Frequency Crop</p>\n<ul>\n<li>For one audio clip, crop 24 different crops according to fmin&amp;fmax of each species.</li>\n<li>I believe it is similar to what <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> did and <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> did(without repeating convolution computations)</li></ul></li>\n<li><p>LSoft</p>\n<ul>\n<li><p>I used only TP and FP crops as labels. For example, for each row in TP or FP, only 1 out of 24 labels is present.</p></li>\n<li><p>Based on BCELoss, I used 'lsoft' for unknown labels. LSoft is introduced kindly by <a href=\"https://www.kaggle.com/romul0212\" target=\"_blank\">@romul0212</a> at <a href=\"https://github.com/lRomul/argus-freesound/blob/master/src/losses.py\" target=\"_blank\">https://github.com/lRomul/argus-freesound/blob/master/src/losses.py</a></p></li>\n<li><p>This is my loss computation with LSoft. <code>mask</code> indicates where the label is known. <code>true</code> for unknown labels is initialized with 0.</p></li></ul>\n<pre><code>tmp_true = (1 - lsoft) * true + lsoft * torch.sigmoid(pred)\ntrue = torch.where(masks == 0, tmp_true, true)\nloss = nn.BCEWithLogitsLoss()(pred, true)\n</code></pre></li>\n<li><p>Iterative Pseudo Training</p>\n<ul>\n<li>Since train set is only very sparsely annotated, I thought re-labeling with the model then re-training will help, and it indeed helped. I pseudo trained for 3 stages.</li>\n<li>When pseudo training, I didn't use LSoft and used vanilla BCE.</li></ul></li>\n<li><p>LSEPLoss</p>\n<ul>\n<li><p>Our metric is LWLRAP, so it is important to focus on rank between labels for each row. I used LSepLoss which fits this purpose, which is introduced kindly by <a href=\"https://www.kaggle.com/ddanevskyi\" target=\"_blank\">@ddanevskyi</a> at <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97926\" target=\"_blank\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97926</a></p></li>\n<li><p>After stage3 of BCE pseudo training, I pseudo trained extra 2 stages with LSEPLoss</p></li>\n<li><p>I fixed original code a bit to allow soft labels.</p></li></ul>\n<pre><code>def lsep_loss(input, target):\n    input_differences = input.unsqueeze(1) - input.unsqueeze(2)\n    target_differences = target.unsqueeze(2) - target.unsqueeze(1)\n    target_differences = torch.maximum(torch.tensor(0).to(input.device), target_differences)\n    exps = input_differences.exp() * target_differences\n    lsep = torch.log(1 + exps.sum(2).sum(1))\n    return lsep.mean()\n</code></pre></li>\n\n<li><p>Global Average Pooling on only positive values</p>\n<ul>\n<li><p>We need to know if the species is present or not. We don't care if it appears frequently or not. I thought doing global average pooling on whole last feature map of CNN will yield high probabilities for frequent occurrences of birdcall in one clip and low probabilities for infrequent occurrences, which doesn't match our goal. So I took mean of only positive values from the last feature map of CNN.</p></li>\n<li><p>Following code is attached at the end of CNN's extracted feature map</p></li></ul>\n<pre><code>mask = (x &gt; 0).float()\nfeatures = (x*mask).sum(dim=(2, 3))/(torch.maximum(mask.sum(dim=(2, 3)), torch.tensor(1e-8).to(mask.device)))\n</code></pre></li>\n<li><p>Augmentations</p>\n<ul>\n<li>Gaussian/Pink NoiseSNR, PitchShift, TimeStretch, TimeShift, VolumeControl, Mixup(take union of the labels), SpecAugment</li>\n<li>Thanks to <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> for kindly sharing <a href=\"https://www.kaggle.com/hidehisaarai1213/rfcx-audio-data-augmentation-japanese-english\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/rfcx-audio-data-augmentation-japanese-english</a></li></ul></li>\n<li><p>One inference on full clip</p>\n<ul>\n<li>I didn't resize the spectrogram, so I was able to train on crops and infer on full image.</li>\n<li>When we don't resize, due to the property of CNN, I believe doing sliding windows prediction on small crops is just an approximation for doing one inference on the full image.</li></ul></li>\n<li><p>Validation - only use known labels</p>\n<ul>\n<li>I did validation clip-wise, only on TP and FP labels. From the prediction, I removed all values corresponding to unknown labels, flattened, then calculated LWLRAP. It correlated with LB quite well on my fold0</li></ul></li>\n</ul>\n<p>My baseline was not so strong(~0.8), so I might had fundamental mistakes in my baseline.<br>\nI achieved 0.927 public with efficientnet-b0 fold0 3seed average, but my score worsened when doing 5fold ensembling. I tried average, rank mean, scaling on axis1 then taking mean, calculating mean of pairwise differences taking average, but it didn't help.<br>\nI'm planning to study top solutions to find out what I missed</p>\n<p>I'd really appreciate it if you share some opinions with my approaches and things that I missed.</p>",
      "rawMarkdown": "Congratulations to the top finishers!\n\nThis was my first encounter to audio competition, so I tried a lot of maybe implausible ideas and learned a lot. Especially, solutions from Cornell Birdcall Identification and Freesound Audio Tagging 2019 were helpful.\n\nMy finish is not strong, but I wanted to share some of the things that I *believe(not sure since my score is not sufficiently high)* worked for me (increased cv or public score), and hear what other kagglers experienced.\n\n* Frequency Crop\n\n  * For one audio clip, crop 24 different crops according to fmin&fmax of each species.\n  * I believe it is similar to what @cpmpml did and @hengck23 did(without repeating convolution computations)\n\n* LSoft\n\n  * I used only TP and FP crops as labels. For example, for each row in TP or FP, only 1 out of 24 labels is present.\n\n  * Based on BCELoss, I used 'lsoft' for unknown labels. LSoft is introduced kindly by @romul0212 at https://github.com/lRomul/argus-freesound/blob/master/src/losses.py\n\n  * This is my loss computation with LSoft. `mask` indicates where the label is known. `true` for unknown labels is initialized with 0.\n\n    ```python\n    tmp_true = (1 - lsoft) * true + lsoft * torch.sigmoid(pred)\n    true = torch.where(masks == 0, tmp_true, true)\n    loss = nn.BCEWithLogitsLoss()(pred, true)\n    ```\n\n* Iterative Pseudo Training\n\n  * Since train set is only very sparsely annotated, I thought re-labeling with the model then re-training will help, and it indeed helped. I pseudo trained for 3 stages.\n  * When pseudo training, I didn't use LSoft and used vanilla BCE.\n\n* LSEPLoss\n\n  * Our metric is LWLRAP, so it is important to focus on rank between labels for each row. I used LSepLoss which fits this purpose, which is introduced kindly by @ddanevskyi at https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97926\n\n  * After stage3 of BCE pseudo training, I pseudo trained extra 2 stages with LSEPLoss\n\n  * I fixed original code a bit to allow soft labels.\n\n    ```python\n    def lsep_loss(input, target):\n        input_differences = input.unsqueeze(1) - input.unsqueeze(2)\n        target_differences = target.unsqueeze(2) - target.unsqueeze(1)\n        target_differences = torch.maximum(torch.tensor(0).to(input.device), target_differences)\n        exps = input_differences.exp() * target_differences\n        lsep = torch.log(1 + exps.sum(2).sum(1))\n        return lsep.mean()\n    ```\n\n    \n\n* Global Average Pooling on only positive values\n\n  * We need to know if the species is present or not. We don't care if it appears frequently or not. I thought doing global average pooling on whole last feature map of CNN will yield high probabilities for frequent occurrences of birdcall in one clip and low probabilities for infrequent occurrences, which doesn't match our goal. So I took mean of only positive values from the last feature map of CNN.\n\n  * Following code is attached at the end of CNN's extracted feature map\n\n    ```python\n    mask = (x > 0).float()\n    features = (x*mask).sum(dim=(2, 3))/(torch.maximum(mask.sum(dim=(2, 3)), torch.tensor(1e-8).to(mask.device)))\n    ```\n\n* Augmentations\n\n  * Gaussian/Pink NoiseSNR, PitchShift, TimeStretch, TimeShift, VolumeControl, Mixup(take union of the labels), SpecAugment\n  * Thanks to @hidehisaarai1213 for kindly sharing https://www.kaggle.com/hidehisaarai1213/rfcx-audio-data-augmentation-japanese-english\n\n* One inference on full clip\n\n  * I didn't resize the spectrogram, so I was able to train on crops and infer on full image.\n  * When we don't resize, due to the property of CNN, I believe doing sliding windows prediction on small crops is just an approximation for doing one inference on the full image.\n\n* Validation - only use known labels\n\n  * I did validation clip-wise, only on TP and FP labels. From the prediction, I removed all values corresponding to unknown labels, flattened, then calculated LWLRAP. It correlated with LB quite well on my fold0\n\nMy baseline was not so strong(~0.8), so I might had fundamental mistakes in my baseline.\nI achieved 0.927 public with efficientnet-b0 fold0 3seed average, but my score worsened when doing 5fold ensembling. I tried average, rank mean, scaling on axis1 then taking mean, calculating mean of pairwise differences taking average, but it didn't help.\nI'm planning to study top solutions to find out what I missed\n\nI'd really appreciate it if you share some opinions with my approaches and things that I missed.",
      "votes": null
    },
    {
      "id": "1207747",
      "postDate": "02/18/2021 02:43:10",
      "content": "<p>Congrats on results and thanks for sharing solution. We also used LSepLoss <a href=\"https://www.kaggle.com/harangdev\" target=\"_blank\">@harangdev</a> </p>",
      "rawMarkdown": "Congrats on results and thanks for sharing solution. We also used LSepLoss @harangdev",
      "votes": null
    },
    {
      "id": "1207748",
      "postDate": "02/18/2021 02:45:35",
      "content": "<p>Thanks! Nice to hear that LSep worked for others too.</p>",
      "rawMarkdown": "Thanks! Nice to hear that LSep worked for others too.",
      "votes": null
    },
    {
      "id": "1207770",
      "postDate": "02/18/2021 03:06:26",
      "content": "<p>here is the trick:</p>\n<ul>\n<li><p>you are given a partial label, i.e. you are told if a class is present in the tp. but you are not told if other classes are present. The opposite is for fp annotation, where we are told if a class is absent.</p></li>\n<li><p>but we can create very confident negative samples. you can use external data or mix negative samples, etc.<br>\nthe trick is to create a very large pool of negative samples.</p></li>\n</ul>\n<p>because it is a ranking metric, we can score very well, if there is no negative sample in the top-3 or top-5.<br>\nThe score of the positive samples can be low, but it should not be lower than the positive samples.</p>\n<p>i think if you create many more negative samples, your score should improve (also for ensemble)</p>",
      "rawMarkdown": "here is the trick:\n\n- you are given a partial label, i.e. you are told if a class is present in the tp. but you are not told if other classes are present. The opposite is for fp annotation, where we are told if a class is absent.\n\n- but we can create very confident negative samples. you can use external data or mix negative samples, etc.\nthe trick is to create a very large pool of negative samples.\n\nbecause it is a ranking metric, we can score very well, if there is no negative sample in the top-3 or top-5.\nThe score of the positive samples can be low, but it should not be lower than the positive samples.\n\ni think if you create many more negative samples, your score should improve (also for ensemble)",
      "votes": null
    },
    {
      "id": "1207776",
      "postDate": "02/18/2021 03:18:43",
      "content": "<p>Thanks for your reply! I also agree that more negative samples will improve the performance. But from the solutions I've seen so far, external data is not used and I used mixup. Did you get performance boost from creating more negative samples by using external data or mixing negative samples?</p>",
      "rawMarkdown": "Thanks for your reply! I also agree that more negative samples will improve the performance. But from the solutions I've seen so far, external data is not used and I used mixup. Did you get performance boost from creating more negative samples by using external data or mixing negative samples?",
      "votes": null
    },
    {
      "id": "1208035",
      "postDate": "02/18/2021 06:02:21",
      "content": "<p>I am kind of surprised that the method of training on small crops and then inferring on the whole set works so well. Thanks for the write-up. Good to see all the diverse approaches</p>",
      "rawMarkdown": "I am kind of surprised that the method of training on small crops and then inferring on the whole set works so well. Thanks for the write-up. Good to see all the diverse approaches",
      "votes": null
    },
    {
      "id": "1208058",
      "postDate": "02/18/2021 06:22:15",
      "content": "<p>you can flip the melspec to make more negative samples.</p>\n<p>i did not use external data. <br>\ni did not do heavy augmentation yet. I did some shift from fp annotation, horizontal flip, mixing, etc</p>",
      "rawMarkdown": "you can flip the melspec to make more negative samples.\n\ni did not use external data. \ni did not do heavy augmentation yet. I did some shift from fp annotation, horizontal flip, mixing, etc",
      "votes": null
    },
    {
      "id": "1208059",
      "postDate": "02/18/2021 06:22:25",
      "content": "<p>Yeah, it worked better than taking max of crop predictions in my early experiments, but I might have done something wrong 🤔</p>",
      "rawMarkdown": "Yeah, it worked better than taking max of crop predictions in my early experiments, but I might have done something wrong 🤔",
      "votes": null
    },
    {
      "id": "1208126",
      "postDate": "02/18/2021 07:08:01",
      "content": "<p>Great job!</p>\n<p>btw, glad that my implementation of the LSEP loss worked well for you ;)</p>",
      "rawMarkdown": "Great job!\n\nbtw, glad that my implementation of the LSEP loss worked well for you ;)",
      "votes": null
    },
    {
      "id": "1208249",
      "postDate": "02/18/2021 08:13:18",
      "content": "<p>Thanks again for kindly sharing your code😀</p>",
      "rawMarkdown": "Thanks again for kindly sharing your code😀",
      "votes": null
    },
    {
      "id": "1208507",
      "postDate": "02/18/2021 10:03:05",
      "content": "<p>Congratz ! </p>\n<blockquote>\n  <p>I'd really appreciate it if you share some opinions with my approaches and things that I missed.</p>\n</blockquote>\n<p>Have a read at other's solutions, you'll find plenty of ideas there :)</p>",
      "rawMarkdown": "Congratz ! \n\n> I'd really appreciate it if you share some opinions with my approaches and things that I missed.\n\nHave a read at other's solutions, you'll find plenty of ideas there :)",
      "votes": null
    },
    {
      "id": "1208511",
      "postDate": "02/18/2021 10:06:27",
      "content": "<p>Thanks. Yeah, I see solutions are quite diverse in this competition😃</p>",
      "rawMarkdown": "Thanks. Yeah, I see solutions are quite diverse in this competition😃",
      "votes": null
    },
    {
      "id": "1208514",
      "postDate": "02/18/2021 10:07:59",
      "content": "<p>Nicely done. Congrats on your achievement. <br>\nCode on <code>lsoft_loss</code> is great.</p>",
      "rawMarkdown": "Nicely done. Congrats on your achievement. \nCode on `lsoft_loss` is great.",
      "votes": null
    },
    {
      "id": "1208663",
      "postDate": "02/18/2021 12:08:48",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations",
      "votes": null
    },
    {
      "id": "1220438",
      "postDate": "02/28/2021 02:01:17",
      "content": "<p>I am wondering what does that mean physically by \"flip the melspec\" (upside down?)? Why would that be useful? Thanks.</p>",
      "rawMarkdown": "I am wondering what does that mean physically by \"flip the melspec\" (upside down?)? Why would that be useful? Thanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1207747,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "02/18/2021 02:43:10",
      "content": "<p>Congrats on results and thanks for sharing solution. We also used LSepLoss <a href=\"https://www.kaggle.com/harangdev\" target=\"_blank\">@harangdev</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1207748,
          "author_name": "harangdev",
          "author_url": "",
          "post_date": "02/18/2021 02:45:35",
          "content": "<p>Thanks! Nice to hear that LSep worked for others too.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207770,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/18/2021 03:06:26",
      "content": "<p>here is the trick:</p>\n<ul>\n<li><p>you are given a partial label, i.e. you are told if a class is present in the tp. but you are not told if other classes are present. The opposite is for fp annotation, where we are told if a class is absent.</p></li>\n<li><p>but we can create very confident negative samples. you can use external data or mix negative samples, etc.<br>\nthe trick is to create a very large pool of negative samples.</p></li>\n</ul>\n<p>because it is a ranking metric, we can score very well, if there is no negative sample in the top-3 or top-5.<br>\nThe score of the positive samples can be low, but it should not be lower than the positive samples.</p>\n<p>i think if you create many more negative samples, your score should improve (also for ensemble)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1207776,
          "author_name": "harangdev",
          "author_url": "",
          "post_date": "02/18/2021 03:18:43",
          "content": "<p>Thanks for your reply! I also agree that more negative samples will improve the performance. But from the solutions I've seen so far, external data is not used and I used mixup. Did you get performance boost from creating more negative samples by using external data or mixing negative samples?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208058,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/18/2021 06:22:15",
          "content": "<p>you can flip the melspec to make more negative samples.</p>\n<p>i did not use external data. <br>\ni did not do heavy augmentation yet. I did some shift from fp annotation, horizontal flip, mixing, etc</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1220438,
          "author_name": "wubinbai",
          "author_url": "",
          "post_date": "02/28/2021 02:01:17",
          "content": "<p>I am wondering what does that mean physically by \"flip the melspec\" (upside down?)? Why would that be useful? Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208035,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "02/18/2021 06:02:21",
      "content": "<p>I am kind of surprised that the method of training on small crops and then inferring on the whole set works so well. Thanks for the write-up. Good to see all the diverse approaches</p>",
      "votes": null,
      "replies": [
        {
          "id": 1208059,
          "author_name": "harangdev",
          "author_url": "",
          "post_date": "02/18/2021 06:22:25",
          "content": "<p>Yeah, it worked better than taking max of crop predictions in my early experiments, but I might have done something wrong 🤔</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208126,
      "author_name": "ddanevskyi",
      "author_url": "",
      "post_date": "02/18/2021 07:08:01",
      "content": "<p>Great job!</p>\n<p>btw, glad that my implementation of the LSEP loss worked well for you ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1208249,
          "author_name": "harangdev",
          "author_url": "",
          "post_date": "02/18/2021 08:13:18",
          "content": "<p>Thanks again for kindly sharing your code😀</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208507,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "02/18/2021 10:03:05",
      "content": "<p>Congratz ! </p>\n<blockquote>\n  <p>I'd really appreciate it if you share some opinions with my approaches and things that I missed.</p>\n</blockquote>\n<p>Have a read at other's solutions, you'll find plenty of ideas there :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1208511,
          "author_name": "harangdev",
          "author_url": "",
          "post_date": "02/18/2021 10:06:27",
          "content": "<p>Thanks. Yeah, I see solutions are quite diverse in this competition😃</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208514,
      "author_name": "rajkumarl",
      "author_url": "",
      "post_date": "02/18/2021 10:07:59",
      "content": "<p>Nicely done. Congrats on your achievement. <br>\nCode on <code>lsoft_loss</code> is great.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208663,
      "author_name": "riadalmadani",
      "author_url": "",
      "post_date": "02/18/2021 12:08:48",
      "content": "<p>Congratulations</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1207727": "Congratulations to the top finishers!\n\nThis was my first encounter to audio competition, so I tried a lot of maybe implausible ideas and learned a lot. Especially, solutions from Cornell Birdcall Identification and Freesound Audio Tagging 2019 were helpful.\n\nMy finish is not strong, but I wanted to share some of the things that I *believe(not sure since my score is not sufficiently high)* worked for me (increased cv or public score), and hear what other kagglers experienced.\n\n* Frequency Crop\n\n  * For one audio clip, crop 24 different crops according to fmin&fmax of each species.\n  * I believe it is similar to what @cpmpml did and @hengck23 did(without repeating convolution computations)\n\n* LSoft\n\n  * I used only TP and FP crops as labels. For example, for each row in TP or FP, only 1 out of 24 labels is present.\n\n  * Based on BCELoss, I used 'lsoft' for unknown labels. LSoft is introduced kindly by @romul0212 at https://github.com/lRomul/argus-freesound/blob/master/src/losses.py\n\n  * This is my loss computation with LSoft. `mask` indicates where the label is known. `true` for unknown labels is initialized with 0.\n\n    ```python\n    tmp_true = (1 - lsoft) * true + lsoft * torch.sigmoid(pred)\n    true = torch.where(masks == 0, tmp_true, true)\n    loss = nn.BCEWithLogitsLoss()(pred, true)\n    ```\n\n* Iterative Pseudo Training\n\n  * Since train set is only very sparsely annotated, I thought re-labeling with the model then re-training will help, and it indeed helped. I pseudo trained for 3 stages.\n  * When pseudo training, I didn't use LSoft and used vanilla BCE.\n\n* LSEPLoss\n\n  * Our metric is LWLRAP, so it is important to focus on rank between labels for each row. I used LSepLoss which fits this purpose, which is introduced kindly by @ddanevskyi at https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97926\n\n  * After stage3 of BCE pseudo training, I pseudo trained extra 2 stages with LSEPLoss\n\n  * I fixed original code a bit to allow soft labels.\n\n    ```python\n    def lsep_loss(input, target):\n        input_differences = input.unsqueeze(1) - input.unsqueeze(2)\n        target_differences = target.unsqueeze(2) - target.unsqueeze(1)\n        target_differences = torch.maximum(torch.tensor(0).to(input.device), target_differences)\n        exps = input_differences.exp() * target_differences\n        lsep = torch.log(1 + exps.sum(2).sum(1))\n        return lsep.mean()\n    ```\n\n    \n\n* Global Average Pooling on only positive values\n\n  * We need to know if the species is present or not. We don't care if it appears frequently or not. I thought doing global average pooling on whole last feature map of CNN will yield high probabilities for frequent occurrences of birdcall in one clip and low probabilities for infrequent occurrences, which doesn't match our goal. So I took mean of only positive values from the last feature map of CNN.\n\n  * Following code is attached at the end of CNN's extracted feature map\n\n    ```python\n    mask = (x > 0).float()\n    features = (x*mask).sum(dim=(2, 3))/(torch.maximum(mask.sum(dim=(2, 3)), torch.tensor(1e-8).to(mask.device)))\n    ```\n\n* Augmentations\n\n  * Gaussian/Pink NoiseSNR, PitchShift, TimeStretch, TimeShift, VolumeControl, Mixup(take union of the labels), SpecAugment\n  * Thanks to @hidehisaarai1213 for kindly sharing https://www.kaggle.com/hidehisaarai1213/rfcx-audio-data-augmentation-japanese-english\n\n* One inference on full clip\n\n  * I didn't resize the spectrogram, so I was able to train on crops and infer on full image.\n  * When we don't resize, due to the property of CNN, I believe doing sliding windows prediction on small crops is just an approximation for doing one inference on the full image.\n\n* Validation - only use known labels\n\n  * I did validation clip-wise, only on TP and FP labels. From the prediction, I removed all values corresponding to unknown labels, flattened, then calculated LWLRAP. It correlated with LB quite well on my fold0\n\nMy baseline was not so strong(~0.8), so I might had fundamental mistakes in my baseline.\nI achieved 0.927 public with efficientnet-b0 fold0 3seed average, but my score worsened when doing 5fold ensembling. I tried average, rank mean, scaling on axis1 then taking mean, calculating mean of pairwise differences taking average, but it didn't help.\nI'm planning to study top solutions to find out what I missed\n\nI'd really appreciate it if you share some opinions with my approaches and things that I missed.",
    "1207747": "Congrats on results and thanks for sharing solution. We also used LSepLoss @harangdev",
    "1207748": "Thanks! Nice to hear that LSep worked for others too.",
    "1207770": "here is the trick:\n\n- you are given a partial label, i.e. you are told if a class is present in the tp. but you are not told if other classes are present. The opposite is for fp annotation, where we are told if a class is absent.\n\n- but we can create very confident negative samples. you can use external data or mix negative samples, etc.\nthe trick is to create a very large pool of negative samples.\n\nbecause it is a ranking metric, we can score very well, if there is no negative sample in the top-3 or top-5.\nThe score of the positive samples can be low, but it should not be lower than the positive samples.\n\ni think if you create many more negative samples, your score should improve (also for ensemble)",
    "1207776": "Thanks for your reply! I also agree that more negative samples will improve the performance. But from the solutions I've seen so far, external data is not used and I used mixup. Did you get performance boost from creating more negative samples by using external data or mixing negative samples?",
    "1208035": "I am kind of surprised that the method of training on small crops and then inferring on the whole set works so well. Thanks for the write-up. Good to see all the diverse approaches",
    "1208058": "you can flip the melspec to make more negative samples.\n\ni did not use external data. \ni did not do heavy augmentation yet. I did some shift from fp annotation, horizontal flip, mixing, etc",
    "1208059": "Yeah, it worked better than taking max of crop predictions in my early experiments, but I might have done something wrong 🤔",
    "1208126": "Great job!\n\nbtw, glad that my implementation of the LSEP loss worked well for you ;)",
    "1208249": "Thanks again for kindly sharing your code😀",
    "1208507": "Congratz ! \n\n> I'd really appreciate it if you share some opinions with my approaches and things that I missed.\n\nHave a read at other's solutions, you'll find plenty of ideas there :)",
    "1208511": "Thanks. Yeah, I see solutions are quite diverse in this competition😃",
    "1208514": "Nicely done. Congrats on your achievement. \nCode on `lsoft_loss` is great.",
    "1208663": "Congratulations",
    "1220438": "I am wondering what does that mean physically by \"flip the melspec\" (upside down?)? Why would that be useful? Thanks."
  },
  "source": "meta"
}