{
  "id": 208830,
  "title": "Methods that worked / did not work.",
  "url": "/competitions/rfcx-species-audio-detection/discussion/208830",
  "author_name": "shinmura0",
  "post_date": "2021-01-05T08:17:38.028000",
  "votes": 102,
  "comment_count": 73,
  "views": 0,
  "content": "<p>My main approach is <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">SED</a>.<br>\nThe following methods worked well(LB score is 0.9 over).</p>\n<ul>\n<li>PANNs SED architecture</li>\n<li>BCE Loss</li>\n<li>Random 10sec clip</li>\n<li>Using only tp file</li>\n<li>MixUp</li>\n<li>The base model is EfficientNet</li>\n</ul>\n<p>But the following methods didn't work.</p>\n<ul>\n<li>Dice Loss</li>\n<li>SpecAugment</li>\n<li>Using <strong>fp file</strong> as background noise sound(For example, mix 0.9tp + 0.1fp)</li>\n<li>Using <strong>fp file</strong> as no label sound (all zero label)</li>\n<li>Self Supervised Learning (<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197805\" target=\"_blank\">COLA</a>)</li>\n<li>Discover missing labels and retrain</li>\n</ul>\n<p>COLA has an optimal batch size of 1024 in the paper. However, in my environment, the limit of batch size was 64 and therefore it did not work.</p>\n<p>Also, as many participants may have noticed, there are many <strong>missing labels</strong> in the tp file sources. I discovered the missing labels and retrained using them, but it did not improve.</p>\n<p>Note: I showed missing label example in <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209040\" target=\"_blank\">this topic</a>.</p>",
  "messages": [
    {
      "id": 1139171,
      "postDate": "2021-01-05T08:17:38.030Z",
      "content": "<p>My main approach is <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">SED</a>.<br>\nThe following methods worked well(LB score is 0.9 over).</p>\n<ul>\n<li>PANNs SED architecture</li>\n<li>BCE Loss</li>\n<li>Random 10sec clip</li>\n<li>Using only tp file</li>\n<li>MixUp</li>\n<li>The base model is EfficientNet</li>\n</ul>\n<p>But the following methods didn't work.</p>\n<ul>\n<li>Dice Loss</li>\n<li>SpecAugment</li>\n<li>Using <strong>fp file</strong> as background noise sound(For example, mix 0.9tp + 0.1fp)</li>\n<li>Using <strong>fp file</strong> as no label sound (all zero label)</li>\n<li>Self Supervised Learning (<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197805\" target=\"_blank\">COLA</a>)</li>\n<li>Discover missing labels and retrain</li>\n</ul>\n<p>COLA has an optimal batch size of 1024 in the paper. However, in my environment, the limit of batch size was 64 and therefore it did not work.</p>\n<p>Also, as many participants may have noticed, there are many <strong>missing labels</strong> in the tp file sources. I discovered the missing labels and retrained using them, but it did not improve.</p>\n<p>Note: I showed missing label example in <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209040\" target=\"_blank\">this topic</a>.</p>",
      "rawMarkdown": "My main approach is [SED](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection).\nThe following methods worked well(LB score is 0.9 over).\n+ PANNs SED architecture\n+ BCE Loss\n+ Random 10sec clip\n+ Using only tp file\n+ MixUp\n+ The base model is EfficientNet\n\nBut the following methods didn't work.\n+ Dice Loss\n+ SpecAugment\n+ Using **fp file** as background noise sound(For example, mix 0.9tp + 0.1fp)\n+ Using **fp file** as no label sound (all zero label)\n+ Self Supervised Learning ([COLA](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197805))\n+ Discover missing labels and retrain\n\nCOLA has an optimal batch size of 1024 in the paper. However, in my environment, the limit of batch size was 64 and therefore it did not work.\n\nAlso, as many participants may have noticed, there are many **missing labels** in the tp file sources. I discovered the missing labels and retrained using them, but it did not improve.\n\nNote: I showed missing label example in [this topic](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209040).",
      "votes": 101
    },
    {
      "id": 1161714,
      "postDate": "2021-01-20T17:50:54.753Z",
      "content": "<blockquote>\n  <p>COLA has an optimal batch size of 1024 in the paper. However, in my environment, the limit of batch size was 64 and therefore it did not work.</p>\n</blockquote>\n<p>If you use pytorch then you can simulate larger batch size by one of these, or by combining them:</p>\n<ul>\n<li>gradient accumulation: only call optimizer step every n batches is almost the same as using batch size n times larger.  I say almost because batch norm is updated at each mini batch, not every n mini batches.</li>\n<li>use distributed data parallel, if you have several GPUs.</li>\n</ul>",
      "rawMarkdown": "> COLA has an optimal batch size of 1024 in the paper. However, in my environment, the limit of batch size was 64 and therefore it did not work.\n\nIf you use pytorch then you can simulate larger batch size by one of these, or by combining them:\n- gradient accumulation: only call optimizer step every n batches is almost the same as using batch size n times larger.  I say almost because batch norm is updated at each mini batch, not every n mini batches.\n- use distributed data parallel, if you have several GPUs.",
      "votes": 5,
      "replies": [
        {
          "id": 1162200,
          "postDate": "2021-01-21T03:31:40.007Z",
          "content": "<p>Very thanks.</p>\n<p>Gradient accumulation is good. It looks like you could theoretically make the batch size infinite. I'd like to try it.</p>",
          "rawMarkdown": "Very thanks.\n\nGradient accumulation is good. It looks like you could theoretically make the batch size infinite. I'd like to try it.",
          "votes": 1
        },
        {
          "id": 1162318,
          "postDate": "2021-01-21T05:20:25.480Z",
          "content": "<p>You may also try <a href=\"https://github.com/MalongTech/research-xbm\" target=\"_blank\">Cross-Batch Memory</a>. Gradient accumulation has its own limitations (e.g. doesn't work on BN layers; doesn't work on contrastive/triplet loss without modifications)</p>",
          "rawMarkdown": "You may also try [Cross-Batch Memory](https://github.com/MalongTech/research-xbm). Gradient accumulation has its own limitations (e.g. doesn't work on BN layers; doesn't work on contrastive/triplet loss without modifications)",
          "votes": 3
        },
        {
          "id": 1162557,
          "postDate": "2021-01-21T08:15:19.923Z",
          "content": "<blockquote>\n  <p>doesn't work on BN layers</p>\n</blockquote>\n<p>That's true! If one batch size is too small, Gradient accumulation might not work. That was helpful. Thank you.</p>",
          "rawMarkdown": "> doesn't work on BN layers\n\nThat's true! If one batch size is too small, Gradient accumulation might not work. That was helpful. Thank you.",
          "votes": 1
        },
        {
          "id": 1178202,
          "postDate": "2021-01-30T18:04:45.427Z",
          "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> </p>\n<p>Were you able to try out COLA with larger batch size ?   Was it helpful ? </p>",
          "rawMarkdown": "@shinmurashinmura \n\nWere you able to try out COLA with larger batch size ?   Was it helpful ? "
        },
        {
          "id": 1178265,
          "postDate": "2021-01-30T19:07:27.277Z",
          "content": "<p>i have tried COLA with gradient accumalation but not successful. The test accuracy of COLA was around 93%.</p>",
          "rawMarkdown": "i have tried COLA with gradient accumalation but not successful. The test accuracy of COLA was around 93%.",
          "votes": 2
        },
        {
          "id": 1178322,
          "postDate": "2021-01-30T19:35:56.363Z",
          "content": "<p>\"test accuracy of COLA was around 93%\". You mean multi-label accuracy not the LWLRAP metric right ?  What LWLRAP score did you get ? </p>",
          "rawMarkdown": "\"test accuracy of COLA was around 93%\". You mean multi-label accuracy not the LWLRAP metric right ?  What LWLRAP score did you get ? "
        },
        {
          "id": 1178339,
          "postDate": "2021-01-30T19:54:43.707Z",
          "content": "<p>If you train COLA, it reports a test accuracy. With bs=256, gradient_accum=8, epochs=100 and max_length=300, the final accuracy of COLA was around 94% (shorter max_length -&gt; higher test accuracy up to 99%). After that I have trained a classifier on 6 second segments: 89% LWLRAP, full clip LWLRAP I don't remember. But the model was really bad on LB.</p>",
          "rawMarkdown": "If you train COLA, it reports a test accuracy. With bs=256, gradient_accum=8, epochs=100 and max_length=300, the final accuracy of COLA was around 94% (shorter max_length -> higher test accuracy up to 99%). After that I have trained a classifier on 6 second segments: 89% LWLRAP, full clip LWLRAP I don't remember. But the model was really bad on LB.",
          "votes": 2
        },
        {
          "id": 1178459,
          "postDate": "2021-01-30T22:17:48.500Z",
          "content": "<p>I don't think COLA works with gradient accumulation because of the nature of contrastive loss. But CrossBatchMemory was proven to be useful in MoCo-style self-supervision as stated in this <a href=\"https://github.com/KevinMusgrave/pytorch-metric-learning/tree/master\" target=\"_blank\">repo</a> </p>\n<p>(cannot post attachment right now, just ctrl+f and search for key words <strong>Using loss functions for unsupervised / self-supervised learning</strong></p>",
          "rawMarkdown": "I don't think COLA works with gradient accumulation because of the nature of contrastive loss. But CrossBatchMemory was proven to be useful in MoCo-style self-supervision as stated in this [repo](https://github.com/KevinMusgrave/pytorch-metric-learning/tree/master) \n\n(cannot post attachment right now, just ctrl+f and search for key words **Using loss functions for unsupervised / self-supervised learning**"
        },
        {
          "id": 1178582,
          "postDate": "2021-01-31T01:53:45.467Z",
          "content": "<p>Thanks for sharing the details.</p>\n<p>I just skim through the COLA but here is some preliminary thoughts on potentially tailoring it for our problem set.</p>\n<p>The original paper assumes same audio clip --&gt; Similarity and different audio clip --&gt; Dissimilarity. In our competition, most part of audio clip this similar/dissimilar relation do not exist ( except for a few tiny window where birds call). </p>\n<p>Maybe the naive approach ( like proposed in the paper)  of drawing positive and negative segment from a batch of audio is not ideal because the signal ( similarity / dissimialrity is just too sparse )</p>\n<p>I am thinking along this line of customize it: ( I have not read the paper or the code in that much detail, so maybe my thought is incorrect  ) </p>\n<p>For bird specie X, we construct audio positive clips by randomly merging 10sec segments where X is calling ( tp ) and construct negative clips by merging 10 sec segments where X is not calling (fp). And then somehow modify the batching and sampling procedure to force the model to learn dissimilarity between tp and fp fro specie X.   And for different batch, we could pass through these positive / negative pair for different specie.</p>\n<p>What do you guys think ? <a href=\"https://www.kaggle.com/jihangz\" target=\"_blank\">@jihangz</a> <a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> </p>",
          "rawMarkdown": "Thanks for sharing the details.\n\nI just skim through the COLA but here is some preliminary thoughts on potentially tailoring it for our problem set.\n\nThe original paper assumes same audio clip --> Similarity and different audio clip --> Dissimilarity. In our competition, most part of audio clip this similar/dissimilar relation do not exist ( except for a few tiny window where birds call). \n\nMaybe the naive approach ( like proposed in the paper)  of drawing positive and negative segment from a batch of audio is not ideal because the signal ( similarity / dissimialrity is just too sparse )\n\nI am thinking along this line of customize it: ( I have not read the paper or the code in that much detail, so maybe my thought is incorrect  ) \n\nFor bird specie X, we construct audio positive clips by randomly merging 10sec segments where X is calling ( tp ) and construct negative clips by merging 10 sec segments where X is not calling (fp). And then somehow modify the batching and sampling procedure to force the model to learn dissimilarity between tp and fp fro specie X.   And for different batch, we could pass through these positive / negative pair for different specie.\n\n\nWhat do you guys think ? @jihangz @tugstugi "
        },
        {
          "id": 1178622,
          "postDate": "2021-01-31T03:02:55.680Z",
          "content": "<p><a href=\"https://www.kaggle.com/jihangz\" target=\"_blank\">@jihangz</a> I have tried also smaller max_length with bigger batch sizes on Quadro 48GB, but still no luck.<br>\n<a href=\"https://www.kaggle.com/garfieldchh\" target=\"_blank\">@garfieldchh</a> I am doubting that the audios with TP and FP are enough.</p>",
          "rawMarkdown": "@jihangz I have tried also smaller max_length with bigger batch sizes on Quadro 48GB, but still no luck.\n@garfieldchh I am doubting that the audios with TP and FP are enough."
        }
      ]
    },
    {
      "id": 1139769,
      "postDate": "2021-01-05T15:51:18.823Z",
      "content": "<p>Thanks for sharing!!! What are your validation strategies and metrics ? Do your CV results correlate with your LB score?</p>",
      "rawMarkdown": "Thanks for sharing!!! What are your validation strategies and metrics ? Do your CV results correlate with your LB score?",
      "votes": 3,
      "replies": [
        {
          "id": 1140471,
          "postDate": "2021-01-06T03:06:36.860Z",
          "content": "<p>CV and LB were roughly correlated, but I didn't trust CV because maybe there is a domain shift.</p>",
          "rawMarkdown": "CV and LB were roughly correlated, but I didn't trust CV because maybe there is a domain shift.",
          "votes": 3
        },
        {
          "id": 1140616,
          "postDate": "2021-01-06T06:06:09.417Z",
          "content": "<p>Thanks for the info. If you dont trust CV results, could you give some hint on how you select models for LB? I'm using LWLRAP in 10-second crop windows for CV but the LB results sometimes seems very random and not well correlated with my CV results. I also try CV using LWLRAP on full clip but it does not improve the correlation.</p>",
          "rawMarkdown": "Thanks for the info. If you dont trust CV results, could you give some hint on how you select models for LB? I'm using LWLRAP in 10-second crop windows for CV but the LB results sometimes seems very random and not well correlated with my CV results. I also try CV using LWLRAP on full clip but it does not improve the correlation.",
          "votes": 1
        },
        {
          "id": 1140723,
          "postDate": "2021-01-06T08:16:21.540Z",
          "content": "<p>Do you use SED with weak label? </p>",
          "rawMarkdown": "Do you use SED with weak label? "
        },
        {
          "id": 1140814,
          "postDate": "2021-01-06T09:38:03.103Z",
          "content": "<p>Yes. I'm using SED (either resnet or efficient net as base model) and weak label.</p>",
          "rawMarkdown": "Yes. I'm using SED (either resnet or efficient net as base model) and weak label."
        },
        {
          "id": 1141063,
          "postDate": "2021-01-06T13:15:10.070Z",
          "content": "<p>OK. Why do CV and LBs different? I think three reasons. However, it is a long text and I want to write a new topic later. Please wait.</p>",
          "rawMarkdown": "OK. Why do CV and LBs different? I think three reasons. However, it is a long text and I want to write a new topic later. Please wait.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1139533,
      "postDate": "2021-01-05T13:15:35.620Z",
      "content": "<p>its interesting to see MIXUP improve the soce. for me, the mixup code provided in the PANN was not improving my score. <br>\nI also found with a larger base model(resnext50) the performance reduced. <br>\npsudolabeling for the missing label did not help me too.<br>\nlonger training did help me after correcting the learning rate.</p>",
      "rawMarkdown": "its interesting to see MIXUP improve the soce. for me, the mixup code provided in the PANN was not improving my score. \nI also found with a larger base model(resnext50) the performance reduced. \npsudolabeling for the missing label did not help me too.\nlonger training did help me after correcting the learning rate.",
      "votes": 4,
      "replies": [
        {
          "id": 1139576,
          "postDate": "2021-01-05T13:55:18.700Z",
          "content": "<p>Thanks for sharing. I found that a larger base model had bad performance too.<br>\nEfficientNetB0 and MobileNetV2 was good performance.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EfficientNetB0</td>\n<td>0.893</td>\n</tr>\n<tr>\n<td>MobileNetV2</td>\n<td>0.893</td>\n</tr>\n<tr>\n<td>ResNest50</td>\n<td>0.868</td>\n</tr>\n</tbody>\n</table>",
          "rawMarkdown": "Thanks for sharing. I found that a larger base model had bad performance too.\nEfficientNetB0 and MobileNetV2 was good performance.\n\n||LB|\n|---|---|\n|EfficientNetB0|0.893|\n|MobileNetV2|0.893|\n|ResNest50|0.868|",
          "votes": 5
        },
        {
          "id": 1141201,
          "postDate": "2021-01-06T14:57:19.280Z",
          "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> Are these results from single models or k-folds?</p>",
          "rawMarkdown": "@shinmurashinmura Are these results from single models or k-folds?",
          "votes": 1
        },
        {
          "id": 1142236,
          "postDate": "2021-01-07T08:45:42.767Z",
          "content": "<p>It's 5folds ensemble result. A single model is down about 0.05 in LB.</p>",
          "rawMarkdown": "It's 5folds ensemble result. A single model is down about 0.05 in LB.",
          "votes": 4
        },
        {
          "id": 1143257,
          "postDate": "2021-01-07T20:11:38.427Z",
          "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> thanks for sharing it!<br>\nI am the beginner and little bit confused on how you ensemble 5 fold models in SED. Is it averaging the <code>framewise outputs</code> or something else?</p>",
          "rawMarkdown": "@shinmurashinmura thanks for sharing it!\nI am the beginner and little bit confused on how you ensemble 5 fold models in SED. Is it averaging the `framewise outputs` or something else?"
        },
        {
          "id": 1143740,
          "postDate": "2021-01-08T03:23:07.303Z",
          "content": "<p>I used <code>framewise_outputs</code> for prediction.</p>\n<ul>\n<li>Getting frame level prediction(y-axis is classes. x-axis is time.)</li>\n<li>prediction = np.max(\"frame level prediction\", axis=\"time axis\")</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2Fe4f92bf08c7e326c579321a67e449a8e%2Ffig.png?generation=1609853933525114&amp;alt=media\" alt=\"\"></p>\n<p>And I made submission per a fold(5times).<br>\nFinally, I averaged these submissions. </p>",
          "rawMarkdown": "I used ```framewise_outputs``` for prediction.\n+ Getting frame level prediction(y-axis is classes. x-axis is time.)\n+ prediction = np.max(\"frame level prediction\", axis=\"time axis\")\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2Fe4f92bf08c7e326c579321a67e449a8e%2Ffig.png?generation=1609853933525114&alt=media)\n\nAnd I made submission per a fold(5times).\nFinally, I averaged these submissions. ",
          "votes": 4
        },
        {
          "id": 1143753,
          "postDate": "2021-01-08T03:36:40.170Z",
          "content": "<p>Thank you so much. One more thing…<br>\nIs it better to keep PERIOD=30 or just use all the data once? which is leading to how can I get a score from several chunks predicted on one <code>record_id</code>?</p>",
          "rawMarkdown": "Thank you so much. One more thing...\nIs it better to keep PERIOD=30 or just use all the data once? which is leading to how can I get a score from several chunks predicted on one `record_id`?"
        },
        {
          "id": 1144081,
          "postDate": "2021-01-08T08:17:03.260Z",
          "content": "<p>I experimented period10 vs period30. The winner was period30. I don't recommend long period. <br>\nBecause the longer period it takes, the more <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209040\" target=\"_blank\">missing labels</a> it contains.</p>",
          "rawMarkdown": "I experimented period10 vs period30. The winner was period30. I don't recommend long period. \nBecause the longer period it takes, the more [missing labels](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209040) it contains.",
          "votes": 1
        },
        {
          "id": 1145276,
          "postDate": "2021-01-09T03:28:12.087Z",
          "content": "<p>Thank you for the answer!<br>\nHow do you aggregate predictions during the inference time.<br>\nFor example, <br>\n<code>PERIOD=30</code>;<br>\nprediction shape of <code>000316da7.flac</code> is <code>(3, 24)</code> How should I reduce it to <code>(1, 24)</code></p>",
          "rawMarkdown": "Thank you for the answer!\nHow do you aggregate predictions during the inference time.\nFor example, \n`PERIOD=30`;\nprediction shape of `000316da7.flac` is `(3, 24)` How should I reduce it to `(1, 24)`"
        },
        {
          "id": 1148839,
          "postDate": "2021-01-11T12:21:40.937Z",
          "content": "<p><a href=\"https://www.kaggle.com/yuvaramsingh\" target=\"_blank\">@yuvaramsingh</a> Is there something wrong with the mixup approach in PANN ? they are mixing two spectrogram with a different coefficient right ? Isn't it what the mixup paper presented ?</p>",
          "rawMarkdown": "@yuvaramsingh Is there something wrong with the mixup approach in PANN ? they are mixing two spectrogram with a different coefficient right ? Isn't it what the mixup paper presented ?"
        },
        {
          "id": 1149630,
          "postDate": "2021-01-12T03:02:49.783Z",
          "content": "<p><a href=\"https://www.kaggle.com/bayartsogtya\" target=\"_blank\">@bayartsogtya</a> <br>\nMy prediction is np.max((3,24), axis=time)</p>\n<p><a href=\"https://www.kaggle.com/ludovick\" target=\"_blank\">@ludovick</a> <br>\nI used MixUp with wave style. I random select 10sec clip(A(48000<em>10) and B(48000</em>10)).<br>\nAnd I mix A with B. For example 0.9A+0.1B. Finally, label is mixed.</p>",
          "rawMarkdown": "@bayartsogtya \nMy prediction is np.max((3,24), axis=time)\n\n@ludovick \nI used MixUp with wave style. I random select 10sec clip(A(48000*10) and B(48000*10)).\nAnd I mix A with B. For example 0.9A+0.1B. Finally, label is mixed.\n",
          "votes": 2
        },
        {
          "id": 1152148,
          "postDate": "2021-01-13T21:20:29.660Z",
          "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a>, thanks for this.<br>\nso for each class, you are getting frame predictions in time..<br>\nand you get max of these frame predictions.<br>\nWhat exactly is 'a frame prediction'. Is it a 10 sec of each 60sec recording in sequence?<br>\nSo 6 frame predictions per record in sequence?<br>\n(sorry for my newbie question)</p>",
          "rawMarkdown": "@shinmurashinmura, thanks for this.\nso for each class, you are getting frame predictions in time..\nand you get max of these frame predictions.\nWhat exactly is 'a frame prediction'. Is it a 10 sec of each 60sec recording in sequence?\nSo 6 frame predictions per record in sequence?\n(sorry for my newbie question)"
        },
        {
          "id": 1152327,
          "postDate": "2021-01-14T03:20:35.663Z",
          "content": "<p>I used 10sec clip weak label training. And I split 60sec test clip per 10 sec.<br>\nAnd I get mini prediction(max(framewise_output, axis=time)) per 10 sec.<br>\nFinal prediction is max of mini prediction(=np.max((6,24), axis=0)).</p>",
          "rawMarkdown": "I used 10sec clip weak label training. And I split 60sec test clip per 10 sec.\nAnd I get mini prediction(max(framewise_output, axis=time)) per 10 sec.\nFinal prediction is max of mini prediction(=np.max((6,24), axis=0)).",
          "votes": 2
        },
        {
          "id": 1153688,
          "postDate": "2021-01-15T05:02:29.100Z",
          "content": "<p><a href=\"https://www.kaggle.com/ludovick\" target=\"_blank\">@ludovick</a> Hi! I was wondering what do you mean by \"they are mixing two spectrogram with a different coefficient\"? Thanks! </p>",
          "rawMarkdown": "@ludovick Hi! I was wondering what do you mean by \"they are mixing two spectrogram with a different coefficient\"? Thanks! "
        },
        {
          "id": 1154694,
          "postDate": "2021-01-15T19:41:02.320Z",
          "content": "<p><a href=\"https://www.kaggle.com/tonychenxyz\" target=\"_blank\">@tonychenxyz</a> Hi, the idea of mixup is to combine two images such as <code>new_image = alpha * image_A + (1-alpha) * image_B</code>. (A and B are two different  images from your training data) The labels are then also mixed : <code>new_label = alpha * label_A + (1-alpha) * label_B</code> . In our competition, we can mix raw signal or spectrogram. For now I saw some drop using mixup on spectrogram, and weird results with mixup on raw signal.</p>",
          "rawMarkdown": "@tonychenxyz Hi, the idea of mixup is to combine two images such as `new_image = alpha * image_A + (1-alpha) * image_B`. (A and B are two different  images from your training data) The labels are then also mixed : `new_label = alpha * label_A + (1-alpha) * label_B` . In our competition, we can mix raw signal or spectrogram. For now I saw some drop using mixup on spectrogram, and weird results with mixup on raw signal."
        },
        {
          "id": 1171720,
          "postDate": "2021-01-27T04:23:59.273Z",
          "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> Thanks for sharing this. It is very useful info. Do we have any special pre-processing for using MobileNetV2? My CV dips by half if I switch from Resnet to Mobilenet. This is my first shot at CNN problems…just wondering if I am making rookie mistake somewhere. I have resized to 224, 224 which is what the model expects and done the preprocessing returned by MobileNetV2. Any guidance or suggestions?</p>",
          "rawMarkdown": "@shinmurashinmura Thanks for sharing this. It is very useful info. Do we have any special pre-processing for using MobileNetV2? My CV dips by half if I switch from Resnet to Mobilenet. This is my first shot at CNN problems...just wondering if I am making rookie mistake somewhere. I have resized to 224, 224 which is what the model expects and done the preprocessing returned by MobileNetV2. Any guidance or suggestions?"
        }
      ]
    },
    {
      "id": 1154880,
      "postDate": "2021-01-16T02:36:46.313Z",
      "content": "<p><a href=\"https://www.kaggle.com/shinmura0\" target=\"_blank\">@shinmura0</a>  <br>\nThanks for sharing the approach. I have 2 questions.</p>\n<h3>Random Cropping:</h3>\n<p>You mentioned that training instance is obtained from \"Random 10sec clip(contain tp sound)\".  I understand that  random crop is obtained by some random shift of the tmin tmax window given in the training data, is this correct ?   If yes,  do you generate multiple training instances (multiple random crops with different offsets ) from a single audio file ?   Also,  do you mind sharing random cropping parameters ( IE: the range of possible offset from the ground truth window).<br>\nI am sorry if these data preprocessing questions sound very basic. I am new to audio competition and SED in general .</p>\n<h3>Testing clip length.</h3>\n<p>I am a bit confused about how long a testing segment is.  Are we using 10 sec or 30 sec testing segments ? <br>\nQuoting discussion:<br>\n\"I experimented period10 vs period30. The winner was period30.\"<br>\n\"And I split 60sec test clip per 10 sec.\"<br>\nIf we are not using 10 seconds long test segment, what is the reason for training on 1 segment length and testing on another segment length ?   The SED notebook you referenced seems to be doing different training / testing segment length too ( 5 period training segment and 30 period testing segment ). I am wondering what is the rationale for doing this ? </p>",
      "rawMarkdown": "\n@shinmura0  \n\nThanks for sharing the approach. I have 2 questions.\n\n### Random Cropping:\n\nYou mentioned that training instance is obtained from \"Random 10sec clip(contain tp sound)\".  I understand that  random crop is obtained by some random shift of the tmin tmax window given in the training data, is this correct ?   If yes,  do you generate multiple training instances (multiple random crops with different offsets ) from a single audio file ?   Also,  do you mind sharing random cropping parameters ( IE: the range of possible offset from the ground truth window).\n\nI am sorry if these data preprocessing questions sound very basic. I am new to audio competition and SED in general .\n\n### Testing clip length.\n\nI am a bit confused about how long a testing segment is.  Are we using 10 sec or 30 sec testing segments ? \n\nQuoting discussion:\n\n\"I experimented period10 vs period30. The winner was period30.\"\n\n\"And I split 60sec test clip per 10 sec.\"\n\nIf we are not using 10 seconds long test segment, what is the reason for training on 1 segment length and testing on another segment length ?   The SED notebook you referenced seems to be doing different training / testing segment length too ( 5 period training segment and 30 period testing segment ). I am wondering what is the rationale for doing this ? \n\n\n\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 1157578,
          "postDate": "2021-01-18T03:03:41.867Z",
          "content": "<blockquote>\n  <p>I understand that random crop is obtained by some random shift of the tmin tmax window given in the training data, is this correct ?</p>\n</blockquote>\n<p>Yes.</p>\n<blockquote>\n  <p>If yes, do you generate multiple training instances (multiple random crops with different offsets ) from a single audio file ? </p>\n</blockquote>\n<p>For example,<br>\ntp segment: 11sec ~ 13sec,<br>\nTraining clip is following.</p>\n<ul>\n<li>case1: 3sec ~ 13sec</li>\n<li>case2: 5sec ~ 15sec</li>\n<li>case3: 11sec ~ 21sec</li>\n</ul>\n<p>Cases 1 to 3 are all possible. And more diverse training clips are also possible. It is random.</p>\n<blockquote>\n  <p>Are we using 10 sec or 30 sec testing segments ?</p>\n</blockquote>\n<p>I used 10 sec testing segments.</p>\n<blockquote>\n  <p>The SED notebook you referenced seems to be doing different training</p>\n</blockquote>\n<p>Clip time can be changed between training and test. However, since it changes the prediction a bit, we experimented with it(test clip:10sec vs 30sec).</p>",
          "rawMarkdown": ">  I understand that random crop is obtained by some random shift of the tmin tmax window given in the training data, is this correct ?\n\nYes.\n\n>  If yes, do you generate multiple training instances (multiple random crops with different offsets ) from a single audio file ? \n\nFor example,\ntp segment: 11sec ~ 13sec,\nTraining clip is following.\n+ case1: 3sec ~ 13sec\n+ case2: 5sec ~ 15sec\n+ case3: 11sec ~ 21sec\n\nCases 1 to 3 are all possible. And more diverse training clips are also possible. It is random.\n\n> Are we using 10 sec or 30 sec testing segments ?\n\nI used 10 sec testing segments.\n\n> The SED notebook you referenced seems to be doing different training\n\nClip time can be changed between training and test. However, since it changes the prediction a bit, we experimented with it(test clip:10sec vs 30sec).",
          "votes": 4
        },
        {
          "id": 1159984,
          "postDate": "2021-01-19T15:39:18.177Z",
          "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a>    Got it.  Thanks for the detailed answers. Really appreciate it </p>",
          "rawMarkdown": "@shinmurashinmura    Got it.  Thanks for the detailed answers. Really appreciate it "
        }
      ]
    },
    {
      "id": 1140375,
      "postDate": "2021-01-06T00:55:07.680Z",
      "content": "<p>Hi<br>\nI use the SED too. But it is not stable, could you tell me what wrong in my notebook?</p>\n<p><a href=\"https://www.kaggle.com/truonghoang/rfcx-pannscnn14att-train-inference\" target=\"_blank\">https://www.kaggle.com/truonghoang/rfcx-pannscnn14att-train-inference</a></p>\n<p>Note: I used the Multi-label + LWLRAP<br>\nThanks.</p>",
      "rawMarkdown": "Hi\nI use the SED too. But it is not stable, could you tell me what wrong in my notebook?\n\n[https://www.kaggle.com/truonghoang/rfcx-pannscnn14att-train-inference](https://www.kaggle.com/truonghoang/rfcx-pannscnn14att-train-inference)\n\nNote: I used the Multi-label + LWLRAP\nThanks.",
      "votes": 1,
      "replies": [
        {
          "id": 1140469,
          "postDate": "2021-01-06T03:05:55.820Z",
          "content": "<p>Do you use pannscnn14? In my experiment, pannscnn14 had LB:0.762(single model).<br>\nThis is not strong model. EfficientNet is recommended.</p>",
          "rawMarkdown": "Do you use pannscnn14? In my experiment, pannscnn14 had LB:0.762(single model).\nThis is not strong model. EfficientNet is recommended.",
          "votes": 4
        },
        {
          "id": 1140788,
          "postDate": "2021-01-06T09:08:55.183Z",
          "content": "<p>EfficientB0 is good for me too.<br>\nbut only in train pipeline. When i submit it, it's so worse. </p>",
          "rawMarkdown": "EfficientB0 is good for me too.\nbut only in train pipeline. When i submit it, it's so worse. ",
          "votes": 1
        },
        {
          "id": 1141061,
          "postDate": "2021-01-06T13:14:07.053Z",
          "content": "<p>Why do CV and LBs different? I think three reasons. However, it is a long text and I want to write a new topic later. Please wait.</p>",
          "rawMarkdown": "Why do CV and LBs different? I think three reasons. However, it is a long text and I want to write a new topic later. Please wait.",
          "votes": 3
        },
        {
          "id": 1147676,
          "postDate": "2021-01-10T16:35:02.750Z",
          "content": "<p>same as <a href=\"https://www.kaggle.com/doanquanvietnamca\" target=\"_blank\">@doanquanvietnamca</a> I did not suceed to have results as high as yours with SED model with Eff architecture</p>",
          "rawMarkdown": "same as @doanquanvietnamca I did not suceed to have results as high as yours with SED model with Eff architecture"
        }
      ]
    },
    {
      "id": 1139401,
      "postDate": "2021-01-05T11:14:31.723Z",
      "content": "<p>So SED is once again showing great results, interesting !</p>\n<p>(I say once again because it won the last audio competition)</p>",
      "rawMarkdown": "So SED is once again showing great results, interesting !\n\n(I say once again because it won the last audio competition)",
      "votes": 1,
      "replies": [
        {
          "id": 1139522,
          "postDate": "2021-01-05T13:00:04.570Z",
          "content": "<p>It's Cornell Birdcall Identification!<br>\nYou are excellent performance in that competition.</p>\n<p>I think SED is a strong method in <strong>multi label task.</strong>Cornell Birdcall Identification and this competition is multi label task. Therefore SED is good performance.</p>",
          "rawMarkdown": "It's Cornell Birdcall Identification!\nYou are excellent performance in that competition.\n\nI think SED is a strong method in **multi label task.**Cornell Birdcall Identification and this competition is multi label task. Therefore SED is good performance.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1143495,
      "postDate": "2021-01-07T23:40:26.367Z",
      "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> Thanks you for sharing . Have you try to use the clipwise output ? if it's the case, how good they are compared to segmentwise output ?</p>",
      "rawMarkdown": "@shinmurashinmura Thanks you for sharing . Have you try to use the clipwise output ? if it's the case, how good they are compared to segmentwise output ?",
      "votes": 2,
      "replies": [
        {
          "id": 1143721,
          "postDate": "2021-01-08T03:03:50.423Z",
          "content": "<p>I don't try <code>clipwise_output</code> prediction. Sorry.</p>",
          "rawMarkdown": "I don't try ```clipwise_output``` prediction. Sorry.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1141258,
      "postDate": "2021-01-06T15:35:49.780Z",
      "content": "<p>Does anyone know of a Tensorflow Notebook that implements SED? All the examples I've seen are in PyTorch.</p>",
      "rawMarkdown": "Does anyone know of a Tensorflow Notebook that implements SED? All the examples I've seen are in PyTorch.",
      "votes": 2
    },
    {
      "id": 1140569,
      "postDate": "2021-01-06T05:38:04.827Z",
      "content": "<p>Tnank you for your great discussion!<br>\nI have a question.<br>\nDo you have any reason that you don't use strong label when training SED.<br>\nIntuitively, I think the score will improve more, but do you have any insights?</p>",
      "rawMarkdown": "Tnank you for your great discussion!\nI have a question.\nDo you have any reason that you don't use strong label when training SED.\nIntuitively, I think the score will improve more, but do you have any insights?",
      "votes": 2,
      "replies": [
        {
          "id": 1140728,
          "postDate": "2021-01-06T08:18:26.257Z",
          "content": "<p>The reason for not using strong label is that missing labels exist.<br>\nAs shown <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209040\" target=\"_blank\">here</a>, there are missing labels in many sound sources.</p>\n<p>If I get random 10sec clip, It might contain missing labels. Robust training(for missing labels) is possible by using weak labels.However, I have not experimented with strong label.it may be better result with using strong label.</p>",
          "rawMarkdown": "The reason for not using strong label is that missing labels exist.\nAs shown [here](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209040), there are missing labels in many sound sources.\n\nIf I get random 10sec clip, It might contain missing labels. Robust training(for missing labels) is possible by using weak labels.However, I have not experimented with strong label.it may be better result with using strong label.",
          "votes": 4
        },
        {
          "id": 1140801,
          "postDate": "2021-01-06T09:24:38.450Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1140800,
          "postDate": "2021-01-06T09:24:38.450Z",
          "content": "<p>Tnank you for your answering!<br>\nIt is very convincing.</p>",
          "rawMarkdown": "Tnank you for your answering!\nIt is very convincing."
        }
      ]
    },
    {
      "id": 1146530,
      "postDate": "2021-01-09T20:32:40.460Z",
      "content": "<p>Thanks for kindly sharing your approach in post and your replies really helped. I am waiting for your topic on \"How to use SED\". </p>",
      "rawMarkdown": "Thanks for kindly sharing your approach in post and your replies really helped. I am waiting for your topic on \"How to use SED\". ",
      "votes": 1
    },
    {
      "id": 1176821,
      "postDate": "2021-01-29T20:10:50.770Z",
      "content": "<p>Thanks! I'm joining this competition now, hope it wont be too late :)</p>",
      "rawMarkdown": "Thanks! I'm joining this competition now, hope it wont be too late :)"
    },
    {
      "id": 1168300,
      "postDate": "2021-01-24T20:59:55.140Z",
      "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> nice write-up. I have a question about the strong labels. How do you define them? From what I understand, lets say you take 0-10 sec segment of the whole clip, and there an event at 1-2 sec and the cnn returns a B,10,C tensor after avg the frequency dimension, do you just say the second segment has an event as a binary label? </p>",
      "rawMarkdown": "@shinmurashinmura nice write-up. I have a question about the strong labels. How do you define them? From what I understand, lets say you take 0-10 sec segment of the whole clip, and there an event at 1-2 sec and the cnn returns a B,10,C tensor after avg the frequency dimension, do you just say the second segment has an event as a binary label? ",
      "replies": [
        {
          "id": 1170158,
          "postDate": "2021-01-26T03:21:23.017Z",
          "content": "<blockquote>\n  <p>do you just say the second segment has an event as a binary label?</p>\n</blockquote>\n<p>Yes. Binary label(second segment) is needed.<br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007#1152329\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007#1152329</a></p>",
          "rawMarkdown": "> do you just say the second segment has an event as a binary label?\n\nYes. Binary label(second segment) is needed.\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007#1152329"
        },
        {
          "id": 1173585,
          "postDate": "2021-01-28T01:35:00.730Z",
          "content": "<p>I see. You mentioned you're using clip labels in the other post; from what I understand clip labels give you the best results? I also tried using frame labels (labeling each column after interpolate), but results are worse</p>",
          "rawMarkdown": "I see. You mentioned you're using clip labels in the other post; from what I understand clip labels give you the best results? I also tried using frame labels (labeling each column after interpolate), but results are worse",
          "votes": 1
        },
        {
          "id": 1173669,
          "postDate": "2021-01-28T03:34:37.917Z",
          "content": "<p>I tried clipwise vs framewise for prediction. Clipwise have improved a bit. You are right.<br>\nI also will use clipwise for prediction. Thank you.</p>",
          "rawMarkdown": "I tried clipwise vs framewise for prediction. Clipwise have improved a bit. You are right.\nI also will use clipwise for prediction. Thank you."
        }
      ]
    },
    {
      "id": 1159132,
      "postDate": "2021-01-19T04:20:08.173Z",
      "content": "<p>Interesting to see SED working so well!</p>",
      "rawMarkdown": "Interesting to see SED working so well!"
    },
    {
      "id": 1140878,
      "postDate": "2021-01-06T10:39:10.847Z",
      "content": "<p>I'm new to this compeition after the NFL has just finished, you share a lot on the disscusion which is great for a starter.<br>\nMaybe you could try to treat FP label as TP label, but in a weak prob way?</p>",
      "rawMarkdown": "I'm new to this compeition after the NFL has just finished, you share a lot on the disscusion which is great for a starter.\nMaybe you could try to treat FP label as TP label, but in a weak prob way?"
    },
    {
      "id": 1140829,
      "postDate": "2021-01-06T09:56:40.437Z",
      "content": "<p>Hi, what are weak labels? What does it mean to work with weak labels?</p>",
      "rawMarkdown": "Hi, what are weak labels? What does it mean to work with weak labels?",
      "replies": [
        {
          "id": 1141029,
          "postDate": "2021-01-06T12:46:54.547Z",
          "content": "<p>What is weak label? Let's check link.<br>\n<a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection</a></p>\n<p>And \"Train SED model with only weak supervision\" help you.</p>",
          "rawMarkdown": "What is weak label? Let's check link.\nhttps://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\n\nAnd \"Train SED model with only weak supervision\" help you.",
          "votes": 1
        },
        {
          "id": 1141074,
          "postDate": "2021-01-06T13:24:20.450Z",
          "content": "<p>Thank you, I read that. It's just that in this problem, in both approaches, we cut a segment out of the audio and claim there is a target class. I don't understand what the difference is? Or is the weak label just the name of the SED approach? </p>",
          "rawMarkdown": "Thank you, I read that. It's just that in this problem, in both approaches, we cut a segment out of the audio and claim there is a target class. I don't understand what the difference is? Or is the weak label just the name of the SED approach? "
        },
        {
          "id": 1142232,
          "postDate": "2021-01-07T08:42:05.230Z",
          "content": "<p>Now, I'm writing new topic \"How to use SED\". Please wait.</p>",
          "rawMarkdown": "Now, I'm writing new topic \"How to use SED\". Please wait.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1139338,
      "postDate": "2021-01-05T10:27:08.297Z",
      "content": "<p>Thanks for sharing! Is your SED using strong supervision and producing clipwise output as final prediction?</p>",
      "rawMarkdown": "Thanks for sharing! Is your SED using strong supervision and producing clipwise output as final prediction?",
      "replies": [
        {
          "id": 1139552,
          "postDate": "2021-01-05T13:31:51.947Z",
          "content": "<p>My SED training is </p>\n<ul>\n<li>Random 10sec clip(contain tp sound)</li>\n<li>The label is given to the whole 10sec clip(Weak label training. No strong label training.).</li>\n</ul>\n<p>Final prediction is </p>\n<ul>\n<li>Getting frame level prediction(y-axis is classes. x-axis is time.)</li>\n<li>prediction = np.max(\"frame level prediction\", axis=\"time axis\")</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2Fe4f92bf08c7e326c579321a67e449a8e%2Ffig.png?generation=1609853933525114&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "My SED training is \n+ Random 10sec clip(contain tp sound)\n+ The label is given to the whole 10sec clip(Weak label training. No strong label training.).\n\nFinal prediction is \n+ Getting frame level prediction(y-axis is classes. x-axis is time.)\n+ prediction = np.max(\"frame level prediction\", axis=\"time axis\")\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2Fe4f92bf08c7e326c579321a67e449a8e%2Ffig.png?generation=1609853933525114&alt=media)",
          "votes": 11
        },
        {
          "id": 1139587,
          "postDate": "2021-01-05T14:01:06.603Z",
          "content": "<p>the visualization looks good . i did not think about viz frame prediction as a 2d image. this will help me understand my prediction even better . thanks </p>",
          "rawMarkdown": "the visualization looks good . i did not think about viz frame prediction as a 2d image. this will help me understand my prediction even better . thanks "
        },
        {
          "id": 1141293,
          "postDate": "2021-01-06T15:54:37.770Z",
          "content": "<p>What image size did you use?</p>",
          "rawMarkdown": "What image size did you use?"
        },
        {
          "id": 1144492,
          "postDate": "2021-01-08T13:41:23.640Z",
          "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> Thanks for the clear graphical explaination. My understanding is that you train with clip-wise label and conduct inference with frame-wise label. Is my understanding correct?</p>",
          "rawMarkdown": "@shinmurashinmura Thanks for the clear graphical explaination. My understanding is that you train with clip-wise label and conduct inference with frame-wise label. Is my understanding correct?"
        },
        {
          "id": 1149631,
          "postDate": "2021-01-12T03:03:15.690Z",
          "content": "<p><a href=\"https://www.kaggle.com/tonychenxyz\" target=\"_blank\">@tonychenxyz</a> </p>\n<blockquote>\n  <p>My understanding is that you train with clip-wise label and conduct inference with frame-wise label.</p>\n</blockquote>\n<p>Yes, I used this style. And for your <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209684\" target=\"_blank\">reference</a> (3. A gap training and prediction)</p>",
          "rawMarkdown": "@tonychenxyz \n\n> My understanding is that you train with clip-wise label and conduct inference with frame-wise label.\n\nYes, I used this style. And for your [reference](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209684) (3. A gap training and prediction)",
          "votes": 1
        },
        {
          "id": 1149973,
          "postDate": "2021-01-12T09:17:07.373Z",
          "content": "<p>Thanks for clarifying!</p>",
          "rawMarkdown": "Thanks for clarifying!"
        }
      ]
    },
    {
      "id": 1139224,
      "postDate": "2021-01-05T08:59:24.270Z",
      "content": "<p>Using SED? Interesting…<br>\nHow high is your single model score?</p>",
      "rawMarkdown": "Using SED? Interesting...\nHow high is your single model score?",
      "replies": [
        {
          "id": 1139527,
          "postDate": "2021-01-05T13:03:02.853Z",
          "content": "<p>I forget single model score. <br>\nBut EfficientNetB3(5fold ensemble) had 0.913 score.</p>",
          "rawMarkdown": "I forget single model score. \nBut EfficientNetB3(5fold ensemble) had 0.913 score.",
          "votes": 8
        },
        {
          "id": 1210793,
          "postDate": "2021-02-19T17:29:03.837Z",
          "content": "<p>I tried replicating this but failed probably missing a processing step. Do you mind <a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> sharing your notebook that got you 0.913? Thanks in advance. :)</p>",
          "rawMarkdown": "I tried replicating this but failed probably missing a processing step. Do you mind @shinmurashinmura sharing your notebook that got you 0.913? Thanks in advance. :)"
        }
      ]
    },
    {
      "id": 1143330,
      "postDate": "2021-01-07T21:04:22.937Z",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!"
    }
  ],
  "comments": [
    {
      "id": 1161714,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-01-20T17:50:54.753000",
      "content": "<blockquote>\n  <p>COLA has an optimal batch size of 1024 in the paper. However, in my environment, the limit of batch size was 64 and therefore it did not work.</p>\n</blockquote>\n<p>If you use pytorch then you can simulate larger batch size by one of these, or by combining them:</p>\n<ul>\n<li>gradient accumulation: only call optimizer step every n batches is almost the same as using batch size n times larger.  I say almost because batch norm is updated at each mini batch, not every n mini batches.</li>\n<li>use distributed data parallel, if you have several GPUs.</li>\n</ul>",
      "votes": 5,
      "replies": [
        {
          "id": 1162200,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-21T03:31:40.007000",
          "content": "<p>Very thanks.</p>\n<p>Gradient accumulation is good. It looks like you could theoretically make the batch size infinite. I'd like to try it.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1162318,
          "author_name": "Darkate",
          "author_url": "",
          "post_date": "2021-01-21T05:20:25.480000",
          "content": "<p>You may also try <a href=\"https://github.com/MalongTech/research-xbm\" target=\"_blank\">Cross-Batch Memory</a>. Gradient accumulation has its own limitations (e.g. doesn't work on BN layers; doesn't work on contrastive/triplet loss without modifications)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1162557,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-21T08:15:19.923000",
          "content": "<blockquote>\n  <p>doesn't work on BN layers</p>\n</blockquote>\n<p>That's true! If one batch size is too small, Gradient accumulation might not work. That was helpful. Thank you.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1178202,
          "author_name": "NakedKoala",
          "author_url": "",
          "post_date": "2021-01-30T18:04:45.427000",
          "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> </p>\n<p>Were you able to try out COLA with larger batch size ?   Was it helpful ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1178265,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "2021-01-30T19:07:27.277000",
          "content": "<p>i have tried COLA with gradient accumalation but not successful. The test accuracy of COLA was around 93%.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1178322,
          "author_name": "NakedKoala",
          "author_url": "",
          "post_date": "2021-01-30T19:35:56.363000",
          "content": "<p>\"test accuracy of COLA was around 93%\". You mean multi-label accuracy not the LWLRAP metric right ?  What LWLRAP score did you get ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1178339,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "2021-01-30T19:54:43.707000",
          "content": "<p>If you train COLA, it reports a test accuracy. With bs=256, gradient_accum=8, epochs=100 and max_length=300, the final accuracy of COLA was around 94% (shorter max_length -&gt; higher test accuracy up to 99%). After that I have trained a classifier on 6 second segments: 89% LWLRAP, full clip LWLRAP I don't remember. But the model was really bad on LB.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1178459,
          "author_name": "Darkate",
          "author_url": "",
          "post_date": "2021-01-30T22:17:48.500000",
          "content": "<p>I don't think COLA works with gradient accumulation because of the nature of contrastive loss. But CrossBatchMemory was proven to be useful in MoCo-style self-supervision as stated in this <a href=\"https://github.com/KevinMusgrave/pytorch-metric-learning/tree/master\" target=\"_blank\">repo</a> </p>\n<p>(cannot post attachment right now, just ctrl+f and search for key words <strong>Using loss functions for unsupervised / self-supervised learning</strong></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1178582,
          "author_name": "NakedKoala",
          "author_url": "",
          "post_date": "2021-01-31T01:53:45.467000",
          "content": "<p>Thanks for sharing the details.</p>\n<p>I just skim through the COLA but here is some preliminary thoughts on potentially tailoring it for our problem set.</p>\n<p>The original paper assumes same audio clip --&gt; Similarity and different audio clip --&gt; Dissimilarity. In our competition, most part of audio clip this similar/dissimilar relation do not exist ( except for a few tiny window where birds call). </p>\n<p>Maybe the naive approach ( like proposed in the paper)  of drawing positive and negative segment from a batch of audio is not ideal because the signal ( similarity / dissimialrity is just too sparse )</p>\n<p>I am thinking along this line of customize it: ( I have not read the paper or the code in that much detail, so maybe my thought is incorrect  ) </p>\n<p>For bird specie X, we construct audio positive clips by randomly merging 10sec segments where X is calling ( tp ) and construct negative clips by merging 10 sec segments where X is not calling (fp). And then somehow modify the batching and sampling procedure to force the model to learn dissimilarity between tp and fp fro specie X.   And for different batch, we could pass through these positive / negative pair for different specie.</p>\n<p>What do you guys think ? <a href=\"https://www.kaggle.com/jihangz\" target=\"_blank\">@jihangz</a> <a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1178622,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "2021-01-31T03:02:55.680000",
          "content": "<p><a href=\"https://www.kaggle.com/jihangz\" target=\"_blank\">@jihangz</a> I have tried also smaller max_length with bigger batch sizes on Quadro 48GB, but still no luck.<br>\n<a href=\"https://www.kaggle.com/garfieldchh\" target=\"_blank\">@garfieldchh</a> I am doubting that the audios with TP and FP are enough.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1139769,
      "author_name": "mangeption",
      "author_url": "",
      "post_date": "2021-01-05T15:51:18.823000",
      "content": "<p>Thanks for sharing!!! What are your validation strategies and metrics ? Do your CV results correlate with your LB score?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1140471,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-06T03:06:36.860000",
          "content": "<p>CV and LB were roughly correlated, but I didn't trust CV because maybe there is a domain shift.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1140616,
          "author_name": "mangeption",
          "author_url": "",
          "post_date": "2021-01-06T06:06:09.417000",
          "content": "<p>Thanks for the info. If you dont trust CV results, could you give some hint on how you select models for LB? I'm using LWLRAP in 10-second crop windows for CV but the LB results sometimes seems very random and not well correlated with my CV results. I also try CV using LWLRAP on full clip but it does not improve the correlation.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1140723,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-06T08:16:21.540000",
          "content": "<p>Do you use SED with weak label? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1140814,
          "author_name": "mangeption",
          "author_url": "",
          "post_date": "2021-01-06T09:38:03.103000",
          "content": "<p>Yes. I'm using SED (either resnet or efficient net as base model) and weak label.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1141063,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-06T13:15:10.070000",
          "content": "<p>OK. Why do CV and LBs different? I think three reasons. However, it is a long text and I want to write a new topic later. Please wait.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1139533,
      "author_name": "yuvaramsingh",
      "author_url": "",
      "post_date": "2021-01-05T13:15:35.620000",
      "content": "<p>its interesting to see MIXUP improve the soce. for me, the mixup code provided in the PANN was not improving my score. <br>\nI also found with a larger base model(resnext50) the performance reduced. <br>\npsudolabeling for the missing label did not help me too.<br>\nlonger training did help me after correcting the learning rate.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1139576,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-05T13:55:18.700000",
          "content": "<p>Thanks for sharing. I found that a larger base model had bad performance too.<br>\nEfficientNetB0 and MobileNetV2 was good performance.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EfficientNetB0</td>\n<td>0.893</td>\n</tr>\n<tr>\n<td>MobileNetV2</td>\n<td>0.893</td>\n</tr>\n<tr>\n<td>ResNest50</td>\n<td>0.868</td>\n</tr>\n</tbody>\n</table>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1141201,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2021-01-06T14:57:19.280000",
          "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> Are these results from single models or k-folds?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1142236,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-07T08:45:42.767000",
          "content": "<p>It's 5folds ensemble result. A single model is down about 0.05 in LB.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1143257,
          "author_name": "Bayartsogt Yadamsuren",
          "author_url": "",
          "post_date": "2021-01-07T20:11:38.427000",
          "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> thanks for sharing it!<br>\nI am the beginner and little bit confused on how you ensemble 5 fold models in SED. Is it averaging the <code>framewise outputs</code> or something else?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1143740,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-08T03:23:07.303000",
          "content": "<p>I used <code>framewise_outputs</code> for prediction.</p>\n<ul>\n<li>Getting frame level prediction(y-axis is classes. x-axis is time.)</li>\n<li>prediction = np.max(\"frame level prediction\", axis=\"time axis\")</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2Fe4f92bf08c7e326c579321a67e449a8e%2Ffig.png?generation=1609853933525114&amp;alt=media\" alt=\"\"></p>\n<p>And I made submission per a fold(5times).<br>\nFinally, I averaged these submissions. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1143753,
          "author_name": "Bayartsogt Yadamsuren",
          "author_url": "",
          "post_date": "2021-01-08T03:36:40.170000",
          "content": "<p>Thank you so much. One more thing…<br>\nIs it better to keep PERIOD=30 or just use all the data once? which is leading to how can I get a score from several chunks predicted on one <code>record_id</code>?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1144081,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-08T08:17:03.260000",
          "content": "<p>I experimented period10 vs period30. The winner was period30. I don't recommend long period. <br>\nBecause the longer period it takes, the more <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209040\" target=\"_blank\">missing labels</a> it contains.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1145276,
          "author_name": "Bayartsogt Yadamsuren",
          "author_url": "",
          "post_date": "2021-01-09T03:28:12.087000",
          "content": "<p>Thank you for the answer!<br>\nHow do you aggregate predictions during the inference time.<br>\nFor example, <br>\n<code>PERIOD=30</code>;<br>\nprediction shape of <code>000316da7.flac</code> is <code>(3, 24)</code> How should I reduce it to <code>(1, 24)</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1148839,
          "author_name": "Shiro",
          "author_url": "",
          "post_date": "2021-01-11T12:21:40.937000",
          "content": "<p><a href=\"https://www.kaggle.com/yuvaramsingh\" target=\"_blank\">@yuvaramsingh</a> Is there something wrong with the mixup approach in PANN ? they are mixing two spectrogram with a different coefficient right ? Isn't it what the mixup paper presented ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1149630,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-12T03:02:49.783000",
          "content": "<p><a href=\"https://www.kaggle.com/bayartsogtya\" target=\"_blank\">@bayartsogtya</a> <br>\nMy prediction is np.max((3,24), axis=time)</p>\n<p><a href=\"https://www.kaggle.com/ludovick\" target=\"_blank\">@ludovick</a> <br>\nI used MixUp with wave style. I random select 10sec clip(A(48000<em>10) and B(48000</em>10)).<br>\nAnd I mix A with B. For example 0.9A+0.1B. Finally, label is mixed.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1152148,
          "author_name": "GUNER",
          "author_url": "",
          "post_date": "2021-01-13T21:20:29.660000",
          "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a>, thanks for this.<br>\nso for each class, you are getting frame predictions in time..<br>\nand you get max of these frame predictions.<br>\nWhat exactly is 'a frame prediction'. Is it a 10 sec of each 60sec recording in sequence?<br>\nSo 6 frame predictions per record in sequence?<br>\n(sorry for my newbie question)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1152327,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-14T03:20:35.663000",
          "content": "<p>I used 10sec clip weak label training. And I split 60sec test clip per 10 sec.<br>\nAnd I get mini prediction(max(framewise_output, axis=time)) per 10 sec.<br>\nFinal prediction is max of mini prediction(=np.max((6,24), axis=0)).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1153688,
          "author_name": "Gold Retriever",
          "author_url": "",
          "post_date": "2021-01-15T05:02:29.100000",
          "content": "<p><a href=\"https://www.kaggle.com/ludovick\" target=\"_blank\">@ludovick</a> Hi! I was wondering what do you mean by \"they are mixing two spectrogram with a different coefficient\"? Thanks! </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1154694,
          "author_name": "Shiro",
          "author_url": "",
          "post_date": "2021-01-15T19:41:02.320000",
          "content": "<p><a href=\"https://www.kaggle.com/tonychenxyz\" target=\"_blank\">@tonychenxyz</a> Hi, the idea of mixup is to combine two images such as <code>new_image = alpha * image_A + (1-alpha) * image_B</code>. (A and B are two different  images from your training data) The labels are then also mixed : <code>new_label = alpha * label_A + (1-alpha) * label_B</code> . In our competition, we can mix raw signal or spectrogram. For now I saw some drop using mixup on spectrogram, and weird results with mixup on raw signal.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1171720,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2021-01-27T04:23:59.273000",
          "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> Thanks for sharing this. It is very useful info. Do we have any special pre-processing for using MobileNetV2? My CV dips by half if I switch from Resnet to Mobilenet. This is my first shot at CNN problems…just wondering if I am making rookie mistake somewhere. I have resized to 224, 224 which is what the model expects and done the preprocessing returned by MobileNetV2. Any guidance or suggestions?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1154880,
      "author_name": "NakedKoala",
      "author_url": "",
      "post_date": "2021-01-16T02:36:46.313000",
      "content": "<p><a href=\"https://www.kaggle.com/shinmura0\" target=\"_blank\">@shinmura0</a>  <br>\nThanks for sharing the approach. I have 2 questions.</p>\n<h3>Random Cropping:</h3>\n<p>You mentioned that training instance is obtained from \"Random 10sec clip(contain tp sound)\".  I understand that  random crop is obtained by some random shift of the tmin tmax window given in the training data, is this correct ?   If yes,  do you generate multiple training instances (multiple random crops with different offsets ) from a single audio file ?   Also,  do you mind sharing random cropping parameters ( IE: the range of possible offset from the ground truth window).<br>\nI am sorry if these data preprocessing questions sound very basic. I am new to audio competition and SED in general .</p>\n<h3>Testing clip length.</h3>\n<p>I am a bit confused about how long a testing segment is.  Are we using 10 sec or 30 sec testing segments ? <br>\nQuoting discussion:<br>\n\"I experimented period10 vs period30. The winner was period30.\"<br>\n\"And I split 60sec test clip per 10 sec.\"<br>\nIf we are not using 10 seconds long test segment, what is the reason for training on 1 segment length and testing on another segment length ?   The SED notebook you referenced seems to be doing different training / testing segment length too ( 5 period training segment and 30 period testing segment ). I am wondering what is the rationale for doing this ? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1157578,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-18T03:03:41.867000",
          "content": "<blockquote>\n  <p>I understand that random crop is obtained by some random shift of the tmin tmax window given in the training data, is this correct ?</p>\n</blockquote>\n<p>Yes.</p>\n<blockquote>\n  <p>If yes, do you generate multiple training instances (multiple random crops with different offsets ) from a single audio file ? </p>\n</blockquote>\n<p>For example,<br>\ntp segment: 11sec ~ 13sec,<br>\nTraining clip is following.</p>\n<ul>\n<li>case1: 3sec ~ 13sec</li>\n<li>case2: 5sec ~ 15sec</li>\n<li>case3: 11sec ~ 21sec</li>\n</ul>\n<p>Cases 1 to 3 are all possible. And more diverse training clips are also possible. It is random.</p>\n<blockquote>\n  <p>Are we using 10 sec or 30 sec testing segments ?</p>\n</blockquote>\n<p>I used 10 sec testing segments.</p>\n<blockquote>\n  <p>The SED notebook you referenced seems to be doing different training</p>\n</blockquote>\n<p>Clip time can be changed between training and test. However, since it changes the prediction a bit, we experimented with it(test clip:10sec vs 30sec).</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1159984,
          "author_name": "NakedKoala",
          "author_url": "",
          "post_date": "2021-01-19T15:39:18.177000",
          "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a>    Got it.  Thanks for the detailed answers. Really appreciate it </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1140375,
      "author_name": "( ͡° ͜ʖ ͡°)",
      "author_url": "",
      "post_date": "2021-01-06T00:55:07.680000",
      "content": "<p>Hi<br>\nI use the SED too. But it is not stable, could you tell me what wrong in my notebook?</p>\n<p><a href=\"https://www.kaggle.com/truonghoang/rfcx-pannscnn14att-train-inference\" target=\"_blank\">https://www.kaggle.com/truonghoang/rfcx-pannscnn14att-train-inference</a></p>\n<p>Note: I used the Multi-label + LWLRAP<br>\nThanks.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1140469,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-06T03:05:55.820000",
          "content": "<p>Do you use pannscnn14? In my experiment, pannscnn14 had LB:0.762(single model).<br>\nThis is not strong model. EfficientNet is recommended.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1140788,
          "author_name": "Manh Lab",
          "author_url": "",
          "post_date": "2021-01-06T09:08:55.183000",
          "content": "<p>EfficientB0 is good for me too.<br>\nbut only in train pipeline. When i submit it, it's so worse. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1141061,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-06T13:14:07.053000",
          "content": "<p>Why do CV and LBs different? I think three reasons. However, it is a long text and I want to write a new topic later. Please wait.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1147676,
          "author_name": "Shiro",
          "author_url": "",
          "post_date": "2021-01-10T16:35:02.750000",
          "content": "<p>same as <a href=\"https://www.kaggle.com/doanquanvietnamca\" target=\"_blank\">@doanquanvietnamca</a> I did not suceed to have results as high as yours with SED model with Eff architecture</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1139401,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2021-01-05T11:14:31.723000",
      "content": "<p>So SED is once again showing great results, interesting !</p>\n<p>(I say once again because it won the last audio competition)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1139522,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-05T13:00:04.570000",
          "content": "<p>It's Cornell Birdcall Identification!<br>\nYou are excellent performance in that competition.</p>\n<p>I think SED is a strong method in <strong>multi label task.</strong>Cornell Birdcall Identification and this competition is multi label task. Therefore SED is good performance.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1143495,
      "author_name": "Shiro",
      "author_url": "",
      "post_date": "2021-01-07T23:40:26.367000",
      "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> Thanks you for sharing . Have you try to use the clipwise output ? if it's the case, how good they are compared to segmentwise output ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1143721,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-08T03:03:50.423000",
          "content": "<p>I don't try <code>clipwise_output</code> prediction. Sorry.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1141258,
      "author_name": "Trevor Dunlap",
      "author_url": "",
      "post_date": "2021-01-06T15:35:49.780000",
      "content": "<p>Does anyone know of a Tensorflow Notebook that implements SED? All the examples I've seen are in PyTorch.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1140569,
      "author_name": "ibuki",
      "author_url": "",
      "post_date": "2021-01-06T05:38:04.827000",
      "content": "<p>Tnank you for your great discussion!<br>\nI have a question.<br>\nDo you have any reason that you don't use strong label when training SED.<br>\nIntuitively, I think the score will improve more, but do you have any insights?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1140728,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-06T08:18:26.257000",
          "content": "<p>The reason for not using strong label is that missing labels exist.<br>\nAs shown <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209040\" target=\"_blank\">here</a>, there are missing labels in many sound sources.</p>\n<p>If I get random 10sec clip, It might contain missing labels. Robust training(for missing labels) is possible by using weak labels.However, I have not experimented with strong label.it may be better result with using strong label.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1140801,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-01-06T09:24:38.450000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1140800,
          "author_name": "ibuki",
          "author_url": "",
          "post_date": "2021-01-06T09:24:38.450000",
          "content": "<p>Tnank you for your answering!<br>\nIt is very convincing.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1146530,
      "author_name": "sin",
      "author_url": "",
      "post_date": "2021-01-09T20:32:40.460000",
      "content": "<p>Thanks for kindly sharing your approach in post and your replies really helped. I am waiting for your topic on \"How to use SED\". </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1176821,
      "author_name": "Enricovoltan",
      "author_url": "",
      "post_date": "2021-01-29T20:10:50.770000",
      "content": "<p>Thanks! I'm joining this competition now, hope it wont be too late :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1168300,
      "author_name": "Shujun",
      "author_url": "",
      "post_date": "2021-01-24T20:59:55.140000",
      "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> nice write-up. I have a question about the strong labels. How do you define them? From what I understand, lets say you take 0-10 sec segment of the whole clip, and there an event at 1-2 sec and the cnn returns a B,10,C tensor after avg the frequency dimension, do you just say the second segment has an event as a binary label? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1170158,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-26T03:21:23.017000",
          "content": "<blockquote>\n  <p>do you just say the second segment has an event as a binary label?</p>\n</blockquote>\n<p>Yes. Binary label(second segment) is needed.<br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007#1152329\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/211007#1152329</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1173585,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2021-01-28T01:35:00.730000",
          "content": "<p>I see. You mentioned you're using clip labels in the other post; from what I understand clip labels give you the best results? I also tried using frame labels (labeling each column after interpolate), but results are worse</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1173669,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-28T03:34:37.917000",
          "content": "<p>I tried clipwise vs framewise for prediction. Clipwise have improved a bit. You are right.<br>\nI also will use clipwise for prediction. Thank you.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1159132,
      "author_name": "nights",
      "author_url": "",
      "post_date": "2021-01-19T04:20:08.173000",
      "content": "<p>Interesting to see SED working so well!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1140878,
      "author_name": "steaphan",
      "author_url": "",
      "post_date": "2021-01-06T10:39:10.847000",
      "content": "<p>I'm new to this compeition after the NFL has just finished, you share a lot on the disscusion which is great for a starter.<br>\nMaybe you could try to treat FP label as TP label, but in a weak prob way?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1140829,
      "author_name": "Kramarenko Vladislav",
      "author_url": "",
      "post_date": "2021-01-06T09:56:40.437000",
      "content": "<p>Hi, what are weak labels? What does it mean to work with weak labels?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1141029,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-06T12:46:54.547000",
          "content": "<p>What is weak label? Let's check link.<br>\n<a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection</a></p>\n<p>And \"Train SED model with only weak supervision\" help you.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1141074,
          "author_name": "Kramarenko Vladislav",
          "author_url": "",
          "post_date": "2021-01-06T13:24:20.450000",
          "content": "<p>Thank you, I read that. It's just that in this problem, in both approaches, we cut a segment out of the audio and claim there is a target class. I don't understand what the difference is? Or is the weak label just the name of the SED approach? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1142232,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-07T08:42:05.230000",
          "content": "<p>Now, I'm writing new topic \"How to use SED\". Please wait.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1139338,
      "author_name": "Darkate",
      "author_url": "",
      "post_date": "2021-01-05T10:27:08.297000",
      "content": "<p>Thanks for sharing! Is your SED using strong supervision and producing clipwise output as final prediction?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1139552,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-05T13:31:51.947000",
          "content": "<p>My SED training is </p>\n<ul>\n<li>Random 10sec clip(contain tp sound)</li>\n<li>The label is given to the whole 10sec clip(Weak label training. No strong label training.).</li>\n</ul>\n<p>Final prediction is </p>\n<ul>\n<li>Getting frame level prediction(y-axis is classes. x-axis is time.)</li>\n<li>prediction = np.max(\"frame level prediction\", axis=\"time axis\")</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2Fe4f92bf08c7e326c579321a67e449a8e%2Ffig.png?generation=1609853933525114&amp;alt=media\" alt=\"\"></p>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 1139587,
          "author_name": "yuvaramsingh",
          "author_url": "",
          "post_date": "2021-01-05T14:01:06.603000",
          "content": "<p>the visualization looks good . i did not think about viz frame prediction as a 2d image. this will help me understand my prediction even better . thanks </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1141293,
          "author_name": "Svitlana Tarasenko ",
          "author_url": "",
          "post_date": "2021-01-06T15:54:37.770000",
          "content": "<p>What image size did you use?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1144492,
          "author_name": "Gold Retriever",
          "author_url": "",
          "post_date": "2021-01-08T13:41:23.640000",
          "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> Thanks for the clear graphical explaination. My understanding is that you train with clip-wise label and conduct inference with frame-wise label. Is my understanding correct?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1149631,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-12T03:03:15.690000",
          "content": "<p><a href=\"https://www.kaggle.com/tonychenxyz\" target=\"_blank\">@tonychenxyz</a> </p>\n<blockquote>\n  <p>My understanding is that you train with clip-wise label and conduct inference with frame-wise label.</p>\n</blockquote>\n<p>Yes, I used this style. And for your <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209684\" target=\"_blank\">reference</a> (3. A gap training and prediction)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1149973,
          "author_name": "Gold Retriever",
          "author_url": "",
          "post_date": "2021-01-12T09:17:07.373000",
          "content": "<p>Thanks for clarifying!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1139224,
      "author_name": "resistance0108",
      "author_url": "",
      "post_date": "2021-01-05T08:59:24.270000",
      "content": "<p>Using SED? Interesting…<br>\nHow high is your single model score?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1139527,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-05T13:03:02.853000",
          "content": "<p>I forget single model score. <br>\nBut EfficientNetB3(5fold ensemble) had 0.913 score.</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 1210793,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2021-02-19T17:29:03.837000",
          "content": "<p>I tried replicating this but failed probably missing a processing step. Do you mind <a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> sharing your notebook that got you 0.913? Thanks in advance. :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1143330,
      "author_name": "Yier",
      "author_url": "",
      "post_date": "2021-01-07T21:04:22.937000",
      "content": "<p>Thank you!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1139171": "My main approach is [SED](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection).\nThe following methods worked well(LB score is 0.9 over).\n+ PANNs SED architecture\n+ BCE Loss\n+ Random 10sec clip\n+ Using only tp file\n+ MixUp\n+ The base model is EfficientNet\n\nBut the following methods didn't work.\n+ Dice Loss\n+ SpecAugment\n+ Using **fp file** as background noise sound(For example, mix 0.9tp + 0.1fp)\n+ Using **fp file** as no label sound (all zero label)\n+ Self Supervised Learning ([COLA](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197805))\n+ Discover missing labels and retrain\n\nCOLA has an optimal batch size of 1024 in the paper. However, in my environment, the limit of batch size was 64 and therefore it did not work.\n\nAlso, as many participants may have noticed, there are many **missing labels** in the tp file sources. I discovered the missing labels and retrained using them, but it did not improve.\n\nNote: I showed missing label example in [this topic](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209040).",
    "1161714": "> COLA has an optimal batch size of 1024 in the paper. However, in my environment, the limit of batch size was 64 and therefore it did not work.\n\nIf you use pytorch then you can simulate larger batch size by one of these, or by combining them:\n- gradient accumulation: only call optimizer step every n batches is almost the same as using batch size n times larger.  I say almost because batch norm is updated at each mini batch, not every n mini batches.\n- use distributed data parallel, if you have several GPUs.",
    "1139769": "Thanks for sharing!!! What are your validation strategies and metrics ? Do your CV results correlate with your LB score?",
    "1139533": "its interesting to see MIXUP improve the soce. for me, the mixup code provided in the PANN was not improving my score. \nI also found with a larger base model(resnext50) the performance reduced. \npsudolabeling for the missing label did not help me too.\nlonger training did help me after correcting the learning rate.",
    "1154880": "\n@shinmura0  \n\nThanks for sharing the approach. I have 2 questions.\n\n### Random Cropping:\n\nYou mentioned that training instance is obtained from \"Random 10sec clip(contain tp sound)\".  I understand that  random crop is obtained by some random shift of the tmin tmax window given in the training data, is this correct ?   If yes,  do you generate multiple training instances (multiple random crops with different offsets ) from a single audio file ?   Also,  do you mind sharing random cropping parameters ( IE: the range of possible offset from the ground truth window).\n\nI am sorry if these data preprocessing questions sound very basic. I am new to audio competition and SED in general .\n\n### Testing clip length.\n\nI am a bit confused about how long a testing segment is.  Are we using 10 sec or 30 sec testing segments ? \n\nQuoting discussion:\n\n\"I experimented period10 vs period30. The winner was period30.\"\n\n\"And I split 60sec test clip per 10 sec.\"\n\nIf we are not using 10 seconds long test segment, what is the reason for training on 1 segment length and testing on another segment length ?   The SED notebook you referenced seems to be doing different training / testing segment length too ( 5 period training segment and 30 period testing segment ). I am wondering what is the rationale for doing this ? \n\n\n\n\n",
    "1140375": "Hi\nI use the SED too. But it is not stable, could you tell me what wrong in my notebook?\n\n[https://www.kaggle.com/truonghoang/rfcx-pannscnn14att-train-inference](https://www.kaggle.com/truonghoang/rfcx-pannscnn14att-train-inference)\n\nNote: I used the Multi-label + LWLRAP\nThanks.",
    "1139401": "So SED is once again showing great results, interesting !\n\n(I say once again because it won the last audio competition)",
    "1143495": "@shinmurashinmura Thanks you for sharing . Have you try to use the clipwise output ? if it's the case, how good they are compared to segmentwise output ?",
    "1141258": "Does anyone know of a Tensorflow Notebook that implements SED? All the examples I've seen are in PyTorch.",
    "1140569": "Tnank you for your great discussion!\nI have a question.\nDo you have any reason that you don't use strong label when training SED.\nIntuitively, I think the score will improve more, but do you have any insights?",
    "1146530": "Thanks for kindly sharing your approach in post and your replies really helped. I am waiting for your topic on \"How to use SED\". ",
    "1176821": "Thanks! I'm joining this competition now, hope it wont be too late :)",
    "1168300": "@shinmurashinmura nice write-up. I have a question about the strong labels. How do you define them? From what I understand, lets say you take 0-10 sec segment of the whole clip, and there an event at 1-2 sec and the cnn returns a B,10,C tensor after avg the frequency dimension, do you just say the second segment has an event as a binary label? ",
    "1159132": "Interesting to see SED working so well!",
    "1140878": "I'm new to this compeition after the NFL has just finished, you share a lot on the disscusion which is great for a starter.\nMaybe you could try to treat FP label as TP label, but in a weak prob way?",
    "1140829": "Hi, what are weak labels? What does it mean to work with weak labels?",
    "1139338": "Thanks for sharing! Is your SED using strong supervision and producing clipwise output as final prediction?",
    "1139224": "Using SED? Interesting...\nHow high is your single model score?",
    "1143330": "Thank you!"
  }
}