{
  "id": 211007,
  "title": "How to use SED",
  "url": "/competitions/rfcx-species-audio-detection/discussion/211007",
  "author_name": "shinmura0",
  "post_date": "2021-01-13T08:16:32.365000",
  "votes": 131,
  "comment_count": 49,
  "views": 0,
  "content": "<p>I use SED(Sound Event Detection) and maybe many participants used it. But SED(PANNs architecture) is difficult to learn for audio competition beginner. I shortly Introduce how to use SED(only PANNs architecture).</p>\n<p>In SED, prediction is classes and time like below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2F315015cca93d0b34eee61e27ec995516%2Ffig1.png?generation=1609987035311073&amp;alt=media\" alt=\"\"></p>\n<h1>1 How to train</h1>\n<p>SED(PANNs architecture)'s output is two type. </p>\n<ul>\n<li>clipwise_output</li>\n<li>framewise_output</li>\n</ul>\n<p>framewise_output is the prediction result of <strong>time and classes</strong>. It can show above figure. Training with framewise_output need time and classes data. And loss function is usually binary cross entropy. Then this label data(contains time and class) is called <strong>\"strong label\"</strong>.</p>\n<p>clipwise_output is the prediction result for <strong>classes of the whole clip</strong>(for example 60sec). It contain no time information(only class information). Training with clipwise_output need only classes data. And loss function is also usually binary cross entropy. Then this label(contains only class) is called <strong>\"weak label\"</strong>.</p>\n<p>In general, strong label training have higher accuracy than weak label training. Because strong labels contain time information.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2Fcfc1dc4d0aacee3b2513a7c44fc3b1d4%2Ffig1.png?generation=1610517949491883&amp;alt=media\" alt=\"\"></p>\n<h1>2 Feature extractor</h1>\n<p>PANNs architecture is above. Feature extractor is the most Important. It is CNN14 in default PANNs architecture. CNN14 is strong model and pretrained model(by audio clip). But in this competition(and <a href=\"https://www.kaggle.com/c/birdsong-recognition\" target=\"_blank\">last competition</a>), CNN14 is not strongest model. In my experiment, EfficientNet is better model. But the input of CNN14(1CH) is different from the one of EfficientNet(3CH). So you need to change the code.</p>\n<p>Here is a sample code(Reference code is <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">here</a>).</p>\n<pre><code>pip install efficientnet_pytorch\n</code></pre>\n<pre><code>from efficientnet_pytorch import EfficientNet\n\nclass feature_extractor(nn.Module):\n    def __init__(self, original):\n        super().__init__()\n        self.model = original\n    def forward(self, x):\n        x= self.model.extract_features(x)\n        return x\n\nclass PANNsCNN14Att(nn.Module):\n    def __init__():\n        self.Enet = feature_extractor(EfficientNet.from_pretrained('efficientnet-b0'))\n    ...\n\n    def cnn_feature_extractor(self, x):\n        return self.Enet(x)\n\n    def forward(self, input, mixup_lambda=None):\n        x, frames_num = self.preprocess(input, mixup_lambda=mixup_lambda)\n        x = torch.cat((x,x,x),1)\n\n        x = self.cnn_feature_extractor(x)\n\n        x = torch.mean(x, dim=3)\n        ...\n</code></pre>\n<p>If errors happen, adjust attblock value or interpolation value.</p>\n<h1>3 How to predict</h1>\n<p>prediction procedure is also two type. </p>\n<ul>\n<li>clipwise_output</li>\n<li>framewise_output</li>\n</ul>\n<p>Prediction with clipwise_output is simple to use. But I think that it is weak for short event. Comparatively, framewise_output prediction is good at short sound event. In this competition, there is many short sound event. Therefore I use framewise_output for prediction. </p>\n<p>My prediction strategy with framewise_outputis is below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2F7207d5745bb1627dcd9733ee7fe4f4a5%2Ffig3.png?generation=1610523433543850&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 1151299,
      "postDate": "2021-01-13T08:16:32.367Z",
      "content": "<p>I use SED(Sound Event Detection) and maybe many participants used it. But SED(PANNs architecture) is difficult to learn for audio competition beginner. I shortly Introduce how to use SED(only PANNs architecture).</p>\n<p>In SED, prediction is classes and time like below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2F315015cca93d0b34eee61e27ec995516%2Ffig1.png?generation=1609987035311073&amp;alt=media\" alt=\"\"></p>\n<h1>1 How to train</h1>\n<p>SED(PANNs architecture)'s output is two type. </p>\n<ul>\n<li>clipwise_output</li>\n<li>framewise_output</li>\n</ul>\n<p>framewise_output is the prediction result of <strong>time and classes</strong>. It can show above figure. Training with framewise_output need time and classes data. And loss function is usually binary cross entropy. Then this label data(contains time and class) is called <strong>\"strong label\"</strong>.</p>\n<p>clipwise_output is the prediction result for <strong>classes of the whole clip</strong>(for example 60sec). It contain no time information(only class information). Training with clipwise_output need only classes data. And loss function is also usually binary cross entropy. Then this label(contains only class) is called <strong>\"weak label\"</strong>.</p>\n<p>In general, strong label training have higher accuracy than weak label training. Because strong labels contain time information.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2Fcfc1dc4d0aacee3b2513a7c44fc3b1d4%2Ffig1.png?generation=1610517949491883&amp;alt=media\" alt=\"\"></p>\n<h1>2 Feature extractor</h1>\n<p>PANNs architecture is above. Feature extractor is the most Important. It is CNN14 in default PANNs architecture. CNN14 is strong model and pretrained model(by audio clip). But in this competition(and <a href=\"https://www.kaggle.com/c/birdsong-recognition\" target=\"_blank\">last competition</a>), CNN14 is not strongest model. In my experiment, EfficientNet is better model. But the input of CNN14(1CH) is different from the one of EfficientNet(3CH). So you need to change the code.</p>\n<p>Here is a sample code(Reference code is <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">here</a>).</p>\n<pre><code>pip install efficientnet_pytorch\n</code></pre>\n<pre><code>from efficientnet_pytorch import EfficientNet\n\nclass feature_extractor(nn.Module):\n    def __init__(self, original):\n        super().__init__()\n        self.model = original\n    def forward(self, x):\n        x= self.model.extract_features(x)\n        return x\n\nclass PANNsCNN14Att(nn.Module):\n    def __init__():\n        self.Enet = feature_extractor(EfficientNet.from_pretrained('efficientnet-b0'))\n    ...\n\n    def cnn_feature_extractor(self, x):\n        return self.Enet(x)\n\n    def forward(self, input, mixup_lambda=None):\n        x, frames_num = self.preprocess(input, mixup_lambda=mixup_lambda)\n        x = torch.cat((x,x,x),1)\n\n        x = self.cnn_feature_extractor(x)\n\n        x = torch.mean(x, dim=3)\n        ...\n</code></pre>\n<p>If errors happen, adjust attblock value or interpolation value.</p>\n<h1>3 How to predict</h1>\n<p>prediction procedure is also two type. </p>\n<ul>\n<li>clipwise_output</li>\n<li>framewise_output</li>\n</ul>\n<p>Prediction with clipwise_output is simple to use. But I think that it is weak for short event. Comparatively, framewise_output prediction is good at short sound event. In this competition, there is many short sound event. Therefore I use framewise_output for prediction. </p>\n<p>My prediction strategy with framewise_outputis is below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2F7207d5745bb1627dcd9733ee7fe4f4a5%2Ffig3.png?generation=1610523433543850&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I use SED(Sound Event Detection) and maybe many participants used it. But SED(PANNs architecture) is difficult to learn for audio competition beginner. I shortly Introduce how to use SED(only PANNs architecture).\n\nIn SED, prediction is classes and time like below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2F315015cca93d0b34eee61e27ec995516%2Ffig1.png?generation=1609987035311073&alt=media)\n\n# 1 How to train\nSED(PANNs architecture)'s output is two type. \n+ clipwise_output\n+ framewise_output\n\nframewise_output is the prediction result of **time and classes**. It can show above figure. Training with framewise_output need time and classes data. And loss function is usually binary cross entropy. Then this label data(contains time and class) is called **\"strong label\"**.\n\nclipwise_output is the prediction result for **classes of the whole clip**(for example 60sec). It contain no time information(only class information). Training with clipwise_output need only classes data. And loss function is also usually binary cross entropy. Then this label(contains only class) is called **\"weak label\"**.\n\nIn general, strong label training have higher accuracy than weak label training. Because strong labels contain time information.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2Fcfc1dc4d0aacee3b2513a7c44fc3b1d4%2Ffig1.png?generation=1610517949491883&alt=media)\n\n# 2 Feature extractor\n\nPANNs architecture is above. Feature extractor is the most Important. It is CNN14 in default PANNs architecture. CNN14 is strong model and pretrained model(by audio clip). But in this competition(and [last competition](https://www.kaggle.com/c/birdsong-recognition)), CNN14 is not strongest model. In my experiment, EfficientNet is better model. But the input of CNN14(1CH) is different from the one of EfficientNet(3CH). So you need to change the code.\n\nHere is a sample code(Reference code is [here](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection)).\n\n```\npip install efficientnet_pytorch\n```\n\n```\nfrom efficientnet_pytorch import EfficientNet\n\nclass feature_extractor(nn.Module):\n    def __init__(self, original):\n        super().__init__()\n        self.model = original\n    def forward(self, x):\n        x= self.model.extract_features(x)\n        return x\n    \nclass PANNsCNN14Att(nn.Module):\n    def __init__():\n        self.Enet = feature_extractor(EfficientNet.from_pretrained('efficientnet-b0'))\n    ...\n\n    def cnn_feature_extractor(self, x):\n        return self.Enet(x)\n\n    def forward(self, input, mixup_lambda=None):\n        x, frames_num = self.preprocess(input, mixup_lambda=mixup_lambda)\n        x = torch.cat((x,x,x),1)\n\n        x = self.cnn_feature_extractor(x)\n        \n        x = torch.mean(x, dim=3)\n        ...\n```\n\nIf errors happen, adjust attblock value or interpolation value.\n\n# 3 How to predict\nprediction procedure is also two type. \n+ clipwise_output\n+ framewise_output\n\nPrediction with clipwise_output is simple to use. But I think that it is weak for short event. Comparatively, framewise_output prediction is good at short sound event. In this competition, there is many short sound event. Therefore I use framewise_output for prediction. \n\nMy prediction strategy with framewise_outputis is below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2F7207d5745bb1627dcd9733ee7fe4f4a5%2Ffig3.png?generation=1610523433543850&alt=media)\n",
      "votes": 130
    },
    {
      "id": 1151606,
      "postDate": "2021-01-13T12:51:52.590Z",
      "content": "<p>If this helps anybody, this is the forward function of an SED model with what happens to the tensor sizes. The architecture is almost the same as the one <a href=\"https://www.kaggle.com/gopidurgaprasad/rfcx-sed-model-stater\" target=\"_blank\">here</a>.</p>\n<pre><code>def forward(self, x):  # BS x C x F x T\n    x = self.encoder.extract_features(x)  # BS x nb_ft x f x t\n    x = torch.mean(x, dim=2)  # BS x nb_ft x t\n\n    x1 = F.max_pool1d(x, kernel_size=3, stride=1, padding=1)\n    x2 = F.avg_pool1d(x, kernel_size=3, stride=1, padding=1)\n    x = x1 + x2  # BS x nb_ft x t\n\n    x = self.dropout(x).transpose(1, 2)  # BS x t x nb_ft\n    x = self.fc(x)  # BS x t x h  (h = 1024)\n\n    x = x.transpose(1, 2)  # BS x h x t\n    clipwise_output, norm_att, segmentwise_output = self.att_block(\n        x\n    )  # BS x num_classes, BS x num_classes x t, BS x num_classes x t\n\n    segmentwise_output = segmentwise_output.transpose(\n        1, 2\n    )  # BS x t x num_classes\n\n    # Upscale back to original size\n    framewise_output = interpolate(segmentwise_output, ratio=32)  # BS x T x num_classes\n\n    return framewise_output, clipwise_output\n</code></pre>",
      "rawMarkdown": "If this helps anybody, this is the forward function of an SED model with what happens to the tensor sizes. The architecture is almost the same as the one [here](https://www.kaggle.com/gopidurgaprasad/rfcx-sed-model-stater).\n\n```\ndef forward(self, x):  # BS x C x F x T\n    x = self.encoder.extract_features(x)  # BS x nb_ft x f x t\n    x = torch.mean(x, dim=2)  # BS x nb_ft x t\n\n    x1 = F.max_pool1d(x, kernel_size=3, stride=1, padding=1)\n    x2 = F.avg_pool1d(x, kernel_size=3, stride=1, padding=1)\n    x = x1 + x2  # BS x nb_ft x t\n\n    x = self.dropout(x).transpose(1, 2)  # BS x t x nb_ft\n    x = self.fc(x)  # BS x t x h  (h = 1024)\n\n    x = x.transpose(1, 2)  # BS x h x t\n    clipwise_output, norm_att, segmentwise_output = self.att_block(\n        x\n    )  # BS x num_classes, BS x num_classes x t, BS x num_classes x t\n\n    segmentwise_output = segmentwise_output.transpose(\n        1, 2\n    )  # BS x t x num_classes\n\n    # Upscale back to original size\n    framewise_output = interpolate(segmentwise_output, ratio=32)  # BS x T x num_classes\n\n    return framewise_output, clipwise_output\n```",
      "votes": 12,
      "replies": [
        {
          "id": 1151613,
          "postDate": "2021-01-13T13:01:21.683Z",
          "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> do you happen to know what's the reason everyone sums max and avg pooling? And why not to go with only one of them or concatenation?</p>\n<p>I did some experiments on simple image datasets and summing avg and max was never the best option. </p>",
          "rawMarkdown": "@theoviel do you happen to know what's the reason everyone sums max and avg pooling? And why not to go with only one of them or concatenation?\n\nI did some experiments on simple image datasets and summing avg and max was never the best option. ",
          "votes": 1
        },
        {
          "id": 1151632,
          "postDate": "2021-01-13T13:14:17.087Z",
          "content": "<p>That's how it's done in the original repo <a href=\"https://github.com/qiuqiangkong/audioset_tagging_cnn\" target=\"_blank\">https://github.com/qiuqiangkong/audioset_tagging_cnn</a><br>\nChanging that is probably a good idea indeed</p>",
          "rawMarkdown": "That's how it's done in the original repo https://github.com/qiuqiangkong/audioset_tagging_cnn\nChanging that is probably a good idea indeed",
          "votes": 1
        },
        {
          "id": 1152517,
          "postDate": "2021-01-14T08:40:46.457Z",
          "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> I have actually done a simple test, but only on a single fold.<br>\nI tried solo average and concatenate. Solo average was not good. </p>\n<p>Although you should take this with a grain of salt, it makes sense to me because max tends to work very good with small objects which is the case here. </p>\n<p>Authors tuned the params on AudioSet which perhaps has many long and short interval sounds in which case combination of max and avg would make sense. </p>",
          "rawMarkdown": "@theoviel I have actually done a simple test, but only on a single fold.\nI tried solo average and concatenate. Solo average was not good. \n\nAlthough you should take this with a grain of salt, it makes sense to me because max tends to work very good with small objects which is the case here. \n\nAuthors tuned the params on AudioSet which perhaps has many long and short interval sounds in which case combination of max and avg would make sense. ",
          "votes": 1
        },
        {
          "id": 1152583,
          "postDate": "2021-01-14T09:37:48.947Z",
          "content": "<p>Thanks for sharing,<br>\nI've noticed that results vary a lot which makes it hard to draw conclusions. I'm guessing focusing on small changes such as this one won't make a big difference in the end</p>",
          "rawMarkdown": "Thanks for sharing,\nI've noticed that results vary a lot which makes it hard to draw conclusions. I'm guessing focusing on small changes such as this one won't make a big difference in the end",
          "votes": 1
        },
        {
          "id": 1152604,
          "postDate": "2021-01-14T10:06:57.480Z",
          "content": "<p>Could not agree more.</p>",
          "rawMarkdown": "Could not agree more.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1162726,
      "postDate": "2021-01-21T09:52:40.543Z",
      "content": "<p>Really nice write-up. Also I have found code of SED models from the winner of the previous Sound competition<br>\n <a href=\"https://github.com/ryanwongsa/kaggle-birdsong-recognition/blob/master/src/models/sed_models.py\" target=\"_blank\">https://github.com/ryanwongsa/kaggle-birdsong-recognition/blob/master/src/models/sed_models.py</a></p>",
      "rawMarkdown": "Really nice write-up. Also I have found code of SED models from the winner of the previous Sound competition\n https://github.com/ryanwongsa/kaggle-birdsong-recognition/blob/master/src/models/sed_models.py",
      "votes": 6,
      "replies": [
        {
          "id": 1162958,
          "postDate": "2021-01-21T12:28:31.060Z",
          "content": "<p>Thank you.</p>",
          "rawMarkdown": "Thank you."
        }
      ]
    },
    {
      "id": 1192149,
      "postDate": "2021-02-09T02:25:06.437Z",
      "content": "<p>Thanks for sharing the kernel. In your case, how about using focal loss instead of BCE?</p>",
      "rawMarkdown": "Thanks for sharing the kernel. In your case, how about using focal loss instead of BCE?",
      "votes": 1,
      "replies": [
        {
          "id": 1192225,
          "postDate": "2021-02-09T03:48:01.747Z",
          "content": "<p>I don't know much about focal loss. But I think focal loss is strong in the imbalance data.<br>\nThis competition data is <strong>not</strong> big imbalance. Therefore If you use focal loss instead of BCE, it won't work. Maybe…</p>",
          "rawMarkdown": "I don't know much about focal loss. But I think focal loss is strong in the imbalance data.\nThis competition data is **not** big imbalance. Therefore If you use focal loss instead of BCE, it won't work. Maybe..."
        }
      ]
    },
    {
      "id": 1173831,
      "postDate": "2021-01-28T06:05:49.163Z",
      "content": "<p>Thanks for sharing this model explanation. For the curious, the SED paperswithcode page gives some details: <a href=\"https://paperswithcode.com/task/sound-event-detection\" target=\"_blank\">https://paperswithcode.com/task/sound-event-detection</a>. </p>",
      "rawMarkdown": "Thanks for sharing this model explanation. For the curious, the SED paperswithcode page gives some details: https://paperswithcode.com/task/sound-event-detection. ",
      "votes": 1
    },
    {
      "id": 1151311,
      "postDate": "2021-01-13T08:29:28.070Z",
      "content": "<p>for how much epoch did you train the model to get above 0.9x and the time period ?<br>\ni am training for  time periods such as 30 sec as it is improving my cv score though </p>",
      "rawMarkdown": "for how much epoch did you train the model to get above 0.9x and the time period ?\ni am training for  time periods such as 30 sec as it is improving my cv score though ",
      "votes": 1,
      "replies": [
        {
          "id": 1151621,
          "postDate": "2021-01-13T13:08:06.573Z",
          "content": "<p>LB 0.9X is below condition.<br>\nepoch:30<br>\ntime period:10sec</p>\n<p>I tried 60sec ,10sec and 5sec period. The best one was 10sec period.</p>",
          "rawMarkdown": "LB 0.9X is below condition.\nepoch:30\ntime period:10sec\n\nI tried 60sec ,10sec and 5sec period. The best one was 10sec period.\n",
          "votes": 2
        },
        {
          "id": 1151726,
          "postDate": "2021-01-13T14:14:14.243Z",
          "content": "<p>thats great thanx for the reply  !</p>",
          "rawMarkdown": "thats great thanx for the reply  !"
        },
        {
          "id": 1151742,
          "postDate": "2021-01-13T14:26:41.623Z",
          "content": "<p>I am a bit confused, can you please elaborate on 'time period'. Is it a time interval on which we make frame_wise prediction?</p>",
          "rawMarkdown": "I am a bit confused, can you please elaborate on 'time period'. Is it a time interval on which we make frame_wise prediction?"
        },
        {
          "id": 1152328,
          "postDate": "2021-01-14T03:20:38.033Z",
          "content": "<p>'time period' is the length of clip to make training data. I used 10sec 'time period'.<br>\nAnd I get random 10 sec clip from 60sec training data.</p>\n<p>In inference time, I split 60sec test clip per 10 sec.<br>\nAnd I get mini prediction(max(framewise_output, axis=time)) per 10 sec.<br>\nFinal prediction is max of mini prediction(=np.max((6,24), axis=0)).</p>",
          "rawMarkdown": "'time period' is the length of clip to make training data. I used 10sec 'time period'.\nAnd I get random 10 sec clip from 60sec training data.\n\nIn inference time, I split 60sec test clip per 10 sec.\nAnd I get mini prediction(max(framewise_output, axis=time)) per 10 sec.\nFinal prediction is max of mini prediction(=np.max((6,24), axis=0)).",
          "votes": 5
        },
        {
          "id": 1152422,
          "postDate": "2021-01-14T06:18:25.430Z",
          "content": "<p>But i think that getting max of framewise_output for 60-sec clip and then applying sigmoid is the equal to \"Final prediction is max of mini prediction\". The difference is only getting max after/before activation (sigmoid), but sigmoid is the monotonic function.</p>",
          "rawMarkdown": "But i think that getting max of framewise_output for 60-sec clip and then applying sigmoid is the equal to \"Final prediction is max of mini prediction\". The difference is only getting max after/before activation (sigmoid), but sigmoid is the monotonic function.",
          "votes": 1
        },
        {
          "id": 1152503,
          "postDate": "2021-01-14T08:19:45.097Z",
          "content": "<p><a href=\"https://www.kaggle.com/sergeyverbitskiy\" target=\"_blank\">@sergeyverbitskiy</a> </p>\n<p>I experimented below condition(train: weak label, inference: framewise_output).</p>\n<ul>\n<li>train:10sec clip / inference:60 sec clip</li>\n<li>train:10sec clip / inference:(10 sec clip) * 6</li>\n</ul>\n<p>The latter is good score. So I used max of 6 mini predictions.</p>",
          "rawMarkdown": "@sergeyverbitskiy \n\nI experimented below condition(train: weak label, inference: framewise_output).\n+ train:10sec clip / inference:60 sec clip\n+ train:10sec clip / inference:(10 sec clip) * 6\n\nThe latter is good score. So I used max of 6 mini predictions.",
          "votes": 1
        },
        {
          "id": 1152596,
          "postDate": "2021-01-14T09:56:23.287Z",
          "content": "<p>I am sure your 0.9+ LB score uses some other tricks :)</p>",
          "rawMarkdown": "I am sure your 0.9+ LB score uses some other tricks :)",
          "votes": 7
        },
        {
          "id": 1152598,
          "postDate": "2021-01-14T10:01:23.343Z",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> I share your opinion but I wonder how can you be sure :) </p>",
          "rawMarkdown": "@philippsinger I share your opinion but I wonder how can you be sure :) "
        },
        {
          "id": 1152616,
          "postDate": "2021-01-14T10:19:20.747Z",
          "content": "<p>trick might be heavy augmentations , i did some heavy aug and it improves my cv  ridiculosly , havent tried lb yet :) </p>",
          "rawMarkdown": "trick might be heavy augmentations , i did some heavy aug and it improves my cv  ridiculosly , havent tried lb yet :) ",
          "votes": 1
        },
        {
          "id": 1152642,
          "postDate": "2021-01-14T10:39:46.147Z",
          "content": "<p>Because apparently people have not been able to replicate it, and there are no public kernels with 0.9+</p>\n<p><a href=\"https://www.kaggle.com/trooperog\" target=\"_blank\">@trooperog</a> interesting, I think OP mentioned more augs didnt help him. Where are you applying augs?</p>",
          "rawMarkdown": "Because apparently people have not been able to replicate it, and there are no public kernels with 0.9+\n\n@trooperog interesting, I think OP mentioned more augs didnt help him. Where are you applying augs?",
          "votes": 3
        },
        {
          "id": 1152671,
          "postDate": "2021-01-14T10:58:52.837Z",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> i am applying augs to the raw waveforms , results are pretty intresting earlier without the augs i was overfitting as i my train metric was above 0.9+ and validation was stuck at 0.82x and after applying augs validation is reaching to 0.88x in 20 epochs of training  .</p>",
          "rawMarkdown": "@philippsinger i am applying augs to the raw waveforms , results are pretty intresting earlier without the augs i was overfitting as i my train metric was above 0.9+ and validation was stuck at 0.82x and after applying augs validation is reaching to 0.88x in 20 epochs of training  .",
          "votes": 10
        },
        {
          "id": 1159828,
          "postDate": "2021-01-19T13:55:58.920Z",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Looks like you've figured out the \"tricks\" :)</p>",
          "rawMarkdown": "@philippsinger Looks like you've figured out the \"tricks\" :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1163553,
      "postDate": "2021-01-21T18:39:59.607Z",
      "content": "<p><a href=\"https://www.kaggle.com/shinmura0\" target=\"_blank\">@shinmura0</a></p>\n<p>Just wanna confirm if I understand all details correctly.</p>\n<p>So when we compute train and val loss. We are doing <code>loss(clipwise_output, target)</code>.</p>\n<p>But when we compute LWRAP metric and generating prediction we are doing <code>metric(framewise_output.max(dim=time), target)</code>.</p>\n<p>Is my understanding correct ? </p>",
      "rawMarkdown": "@shinmura0\n\nJust wanna confirm if I understand all details correctly.\n\nSo when we compute train and val loss. We are doing `loss(clipwise_output, target)`.\n\nBut when we compute LWRAP metric and generating prediction we are doing `metric(framewise_output.max(dim=time), target)`.\n\nIs my understanding correct ? ",
      "votes": 2,
      "replies": [
        {
          "id": 1163776,
          "postDate": "2021-01-21T23:24:52.357Z",
          "content": "<p>Yes. You are correct.</p>",
          "rawMarkdown": "Yes. You are correct."
        },
        {
          "id": 1163784,
          "postDate": "2021-01-21T23:48:14.777Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1368955%2F5de886aab541ee5e47736e75cfbf7fc1%2FScreen%20Shot%202021-01-21%20at%206.42.10%20PM.png?generation=1611272580923419&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/shinmura0\" target=\"_blank\">@shinmura0</a> </p>\n<p>Did you empirically validate that max(framewise_output) is superior to clipwise_output or did you just make the decision based on in tuition ? </p>\n<p>I am also doing the same SED architecture and trying to replicate your 0.90 + performance. Above chart is what I found. </p>\n<p>I am doing validation on entire image (separate validation image to 6 10secs clip  and do same thing as inference ), not just a close crop of annotated event. It looks like lwlrap_clipwise is consistently better than lwlrap_framewise, for local validation.</p>\n<p>What do you think ?  Is local validation completely garbage so that I should not trust this finding. or is it likely because that I am doing something wrong with my implementation</p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1368955%2F5de886aab541ee5e47736e75cfbf7fc1%2FScreen%20Shot%202021-01-21%20at%206.42.10%20PM.png?generation=1611272580923419&alt=media)\n\n@shinmura0 \n\nDid you empirically validate that max(framewise_output) is superior to clipwise_output or did you just make the decision based on in tuition ? \n\nI am also doing the same SED architecture and trying to replicate your 0.90 + performance. Above chart is what I found. \n\nI am doing validation on entire image (separate validation image to 6 10secs clip  and do same thing as inference ), not just a close crop of annotated event. It looks like lwlrap_clipwise is consistently better than lwlrap_framewise, for local validation.\n\nWhat do you think ?  Is local validation completely garbage so that I should not trust this finding. or is it likely because that I am doing something wrong with my implementation",
          "votes": 3
        }
      ]
    },
    {
      "id": 1163131,
      "postDate": "2021-01-21T14:58:33.117Z",
      "content": "<p>Also I would like to ask about some points:</p>\n<ul>\n<li>You have mentioned that you take random crop with <code>tp</code> label.  Do you apply the same strategy for validation set. If so, I am a bit confused, because your validation score will be different for different runs on the same model and data</li>\n<li>You have stated that you have models with 0.89 - 0.91 LB score. What is your CV score for this model. And do you compute CV only on crops with <code>tp</code> labels or you make prediction as on test (6 pieces, each 10 seconds and max(axis=time)) and then compute LWLRAP on the whole recording?<br>\nThanks in advance!</li>\n</ul>",
      "rawMarkdown": "Also I would like to ask about some points:\n- You have mentioned that you take random crop with `tp` label.  Do you apply the same strategy for validation set. If so, I am a bit confused, because your validation score will be different for different runs on the same model and data\n- You have stated that you have models with 0.89 - 0.91 LB score. What is your CV score for this model. And do you compute CV only on crops with `tp` labels or you make prediction as on test (6 pieces, each 10 seconds and max(axis=time)) and then compute LWLRAP on the whole recording?\nThanks in advance!",
      "votes": 2,
      "replies": [
        {
          "id": 1164224,
          "postDate": "2021-01-22T08:15:57.387Z",
          "content": "<blockquote>\n  <p>Do you apply the same strategy for validation set.</p>\n</blockquote>\n<p>Yes. Your suggestion is correct. However, since seed is fixed, the validation clip is the same in every experiment, and I think the missing labels are the reason why the validation is not stable.<br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209684\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209684</a></p>\n<blockquote>\n  <p>And do you compute CV only on crops with tp labels</p>\n</blockquote>\n<p>Yes. I used 10sec random clip for validation. I don't use the whole recording(60sec) for validation. Because there is missing labels in the whole recording.</p>",
          "rawMarkdown": "> Do you apply the same strategy for validation set.\n\nYes. Your suggestion is correct. However, since seed is fixed, the validation clip is the same in every experiment, and I think the missing labels are the reason why the validation is not stable.\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209684\n\n> And do you compute CV only on crops with tp labels\n\nYes. I used 10sec random clip for validation. I don't use the whole recording(60sec) for validation. Because there is missing labels in the whole recording.\n\n\n",
          "votes": 1
        },
        {
          "id": 1164369,
          "postDate": "2021-01-22T10:28:10.710Z",
          "content": "<p>Thanks for your answer!<br>\nAlso I am bit confused with logic of taking a big enough clip:<br>\nIf we have lots of missed values, we can assume that the probability of having missed value in our training clip increasing with the length of the clip. So it is logically to use smaller clips</p>",
          "rawMarkdown": "Thanks for your answer!\nAlso I am bit confused with logic of taking a big enough clip:\nIf we have lots of missed values, we can assume that the probability of having missed value in our training clip increasing with the length of the clip. So it is logically to use smaller clips",
          "votes": 1
        },
        {
          "id": 1164929,
          "postDate": "2021-01-22T16:16:41.483Z",
          "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> </p>\n<p>I was using the crop validation as well but I found it dramatically overestimating the LB score.</p>\n<p>In my setup, I was getting 0.87 in crop based validation but LB is only 0.74 - 0.78. </p>\n<p>Then, I switch to clip based validation and the CV - LB gap becomes much smaller. </p>\n<p>I understand your logic on using crop instead of full clip for validation ( missing labels --&gt; Shorter clip has smaller chance of missing -&gt; more trust-worthy number ) </p>\n<p>I am curious on this: Do you find crop based validation to be more correlated with LB score than clip based validation score , even though it is overly optimistic ? </p>",
          "rawMarkdown": "@shinmurashinmura \n\nI was using the crop validation as well but I found it dramatically overestimating the LB score.\n\nIn my setup, I was getting 0.87 in crop based validation but LB is only 0.74 - 0.78. \n\nThen, I switch to clip based validation and the CV - LB gap becomes much smaller. \n\nI understand your logic on using crop instead of full clip for validation ( missing labels --> Shorter clip has smaller chance of missing -> more trust-worthy number ) \n\nI am curious on this: Do you find crop based validation to be more correlated with LB score than clip based validation score , even though it is overly optimistic ? ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1162693,
      "postDate": "2021-01-21T09:26:02.760Z",
      "content": "<p>Thank you for the explanation! I recently just started trying out the SED architecture and i am confused over some of the details. My understanding is that :</p>\n<ol>\n<li>During training, randomly crop the clip (e.g 10s) and use all the labels associated with the entire clip (but i seen implementation that used just the labels associated within the cropped timeframe)</li>\n<li>Compute loss over the <code>clipwise_output</code> </li>\n<li>Validate/Inference with the <code>framewise_output</code>.<br>\nIs this flow the correct logic? </li>\n</ol>",
      "rawMarkdown": "Thank you for the explanation! I recently just started trying out the SED architecture and i am confused over some of the details. My understanding is that :\n1. During training, randomly crop the clip (e.g 10s) and use all the labels associated with the entire clip (but i seen implementation that used just the labels associated within the cropped timeframe)\n2. Compute loss over the `clipwise_output` \n3. Validate/Inference with the `framewise_output`.\nIs this flow the correct logic? ",
      "votes": 2,
      "replies": [
        {
          "id": 1162986,
          "postDate": "2021-01-21T12:43:10.570Z",
          "content": "<p>2,3 : I think you are correct.</p>\n<blockquote>\n  <ol>\n  <li>the labels associated with the entire clip</li>\n  </ol>\n</blockquote>\n<p>Does \"entire clip\" mean 10sec crop?  labels is not given a whole clip(60sec). So training crop is made by containing tp segment.<br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830#1157578\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830#1157578</a></p>",
          "rawMarkdown": "2,3 : I think you are correct.\n\n> 1. the labels associated with the entire clip\n\nDoes \"entire clip\" mean 10sec crop?  labels is not given a whole clip(60sec). So training crop is made by containing tp segment.\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830#1157578",
          "votes": 1
        },
        {
          "id": 1163027,
          "postDate": "2021-01-21T13:13:51.467Z",
          "content": "<p>Thank you for the above clarification, i really appreciate it! I have followed closely your discussion and what i understand is that </p>\n<ol>\n<li><p>e.g a 60s clip with labels [0,1,2,3,4] , and assuming the first 10s crop has labels [0,1,2]. From your discussion, the ground truth for the first 10s cropped segment should follow [0,1,2], instead of the whole [0,1,2,3,4].</p></li>\n<li><p>I came across this discussion <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/200922#1102470\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/200922#1102470</a> which i believe encourage the use of [0,1,2,3,4] instead. So i was pretty confused by this.    </p></li>\n</ol>",
          "rawMarkdown": "Thank you for the above clarification, i really appreciate it! I have followed closely your discussion and what i understand is that \n1. e.g a 60s clip with labels [0,1,2,3,4] , and assuming the first 10s crop has labels [0,1,2]. From your discussion, the ground truth for the first 10s cropped segment should follow [0,1,2], instead of the whole [0,1,2,3,4].\n\n2.  I came across this discussion https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/200922#1102470 which i believe encourage the use of [0,1,2,3,4] instead. So i was pretty confused by this.    \n "
        }
      ]
    },
    {
      "id": 1158361,
      "postDate": "2021-01-18T14:25:11.217Z",
      "content": "<p>Thanks for nice sharing! How did you crop the audio?</p>",
      "rawMarkdown": "Thanks for nice sharing! How did you crop the audio?",
      "replies": [
        {
          "id": 1159087,
          "postDate": "2021-01-19T03:23:28.753Z",
          "content": "<p>I used time shift and make a clip.<br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830#1154880\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830#1154880</a></p>",
          "rawMarkdown": "I used time shift and make a clip.\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830#1154880"
        }
      ]
    },
    {
      "id": 1157005,
      "postDate": "2021-01-17T15:24:55.837Z",
      "content": "<p>Thanks so much for taking the time to put this together for us. I have one quick question about your inference method. You are taking a max for framewise_output during inference. Why not just leave it to the attention mechanism that you trained up? That's its job after all.</p>",
      "rawMarkdown": "Thanks so much for taking the time to put this together for us. I have one quick question about your inference method. You are taking a max for framewise_output during inference. Why not just leave it to the attention mechanism that you trained up? That's its job after all."
    },
    {
      "id": 1154760,
      "postDate": "2021-01-15T21:42:30.090Z",
      "content": "<p>Hi! Nice idea ! Did you pre trained the model on audio set ? Or maybe have done semi supervised like Mixmatch with data from test set or rainforest audio on YouTube ? </p>",
      "rawMarkdown": "Hi! Nice idea ! Did you pre trained the model on audio set ? Or maybe have done semi supervised like Mixmatch with data from test set or rainforest audio on YouTube ? ",
      "replies": [
        {
          "id": 1155593,
          "postDate": "2021-01-16T13:42:58.900Z",
          "content": "<blockquote>\n  <p>Did you pre trained the model on audio set ?<br>\n  No. But I think it is effective.</p>\n  <p>have done semi supervised like Mixmatch<br>\n  No. I don't touch test data. I train with tp data and use trained model's prediction for test data.</p>\n</blockquote>",
          "rawMarkdown": "> Did you pre trained the model on audio set ?\nNo. But I think it is effective.\n\n> have done semi supervised like Mixmatch\nNo. I don't touch test data. I train with tp data and use trained model's prediction for test data."
        }
      ]
    },
    {
      "id": 1154599,
      "postDate": "2021-01-15T18:11:58.413Z",
      "content": "<p>Hi, </p>\n<p>Thanks for your explanation. I was waiting for it 😄</p>",
      "rawMarkdown": "Hi, \n\nThanks for your explanation. I was waiting for it 😄"
    },
    {
      "id": 1152125,
      "postDate": "2021-01-13T20:42:31.493Z",
      "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> Thank you for this post! I have a question - since you train on Framewise Output (i.e. time and class) how are you inputting both time and class information into BCE? Are you simply taking bce(y_true_one_hot, np.max(framewise_output, axis=time))? When looking here (<a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection</a>) the Loss is based on clipwise output. </p>",
      "rawMarkdown": "@shinmurashinmura Thank you for this post! I have a question - since you train on Framewise Output (i.e. time and class) how are you inputting both time and class information into BCE? Are you simply taking bce(y_true_one_hot, np.max(framewise_output, axis=time))? When looking here (https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection) the Loss is based on clipwise output. ",
      "replies": [
        {
          "id": 1152329,
          "postDate": "2021-01-14T03:20:41.997Z",
          "content": "<p>In reference notebook, training method is \"weak label training\". So clipwise output is fitted whole clip label. On the other hand, In \"strong label training\", you need to prepare time and class data like below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2Fee6a92ad7e3490e9a1611e824b49b18d%2Ffig4.png?generation=1610584861906771&amp;alt=media\" alt=\"\"></p>\n<p>In this style, you can train model using framewise output with BCE.</p>",
          "rawMarkdown": "In reference notebook, training method is \"weak label training\". So clipwise output is fitted whole clip label. On the other hand, In \"strong label training\", you need to prepare time and class data like below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2Fee6a92ad7e3490e9a1611e824b49b18d%2Ffig4.png?generation=1610584861906771&alt=media)\n\nIn this style, you can train model using framewise output with BCE.",
          "votes": 3
        },
        {
          "id": 1158248,
          "postDate": "2021-01-18T13:34:40.577Z",
          "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> Thanks for sharing nice post ! I tried to build strong label training. loss was going down, but CV was stuck in the 0.2x while training. maybe I made a mistake.<br>\nIf there is 10 period, 1501 frames(320 hop_size), <code>framewise output</code> should be `(bs, 1501, 24). so if I feed (bs, 1501, 24) target in BCE with framewise output, It's gonna be strong label training. Am I understanding as a right way ?  Thanks :)</p>",
          "rawMarkdown": "@shinmurashinmura Thanks for sharing nice post ! I tried to build strong label training. loss was going down, but CV was stuck in the 0.2x while training. maybe I made a mistake.\nIf there is 10 period, 1501 frames(320 hop_size), `framewise output` should be `(bs, 1501, 24). so if I feed (bs, 1501, 24) target in BCE with framewise output, It's gonna be strong label training. Am I understanding as a right way ?  Thanks :)"
        },
        {
          "id": 1159086,
          "postDate": "2021-01-19T03:22:53.373Z",
          "content": "<p>Yes. Your are correct. <br>\nI also try to train the model with strong label. Loss and CV improved. But LB did not imporve.</p>",
          "rawMarkdown": "Yes. Your are correct. \nI also try to train the model with strong label. Loss and CV improved. But LB did not imporve.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1151445,
      "postDate": "2021-01-13T10:29:57.847Z",
      "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a>  have you tried any audio augmentations (like Guassian noise etc.)? If any could you enlighten us ?</p>",
      "rawMarkdown": "@shinmurashinmura  have you tried any audio augmentations (like Guassian noise etc.)? If any could you enlighten us ?",
      "replies": [
        {
          "id": 1151617,
          "postDate": "2021-01-13T13:04:42.873Z",
          "content": "<p>I tried to add white noise and pink noise. But they didn't work.</p>",
          "rawMarkdown": "I tried to add white noise and pink noise. But they didn't work.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1163289,
      "postDate": "2021-01-21T16:15:13.230Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1163449,
          "postDate": "2021-01-21T17:20:36.990Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1163484,
          "postDate": "2021-01-21T17:37:55.167Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1151606,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2021-01-13T12:51:52.590000",
      "content": "<p>If this helps anybody, this is the forward function of an SED model with what happens to the tensor sizes. The architecture is almost the same as the one <a href=\"https://www.kaggle.com/gopidurgaprasad/rfcx-sed-model-stater\" target=\"_blank\">here</a>.</p>\n<pre><code>def forward(self, x):  # BS x C x F x T\n    x = self.encoder.extract_features(x)  # BS x nb_ft x f x t\n    x = torch.mean(x, dim=2)  # BS x nb_ft x t\n\n    x1 = F.max_pool1d(x, kernel_size=3, stride=1, padding=1)\n    x2 = F.avg_pool1d(x, kernel_size=3, stride=1, padding=1)\n    x = x1 + x2  # BS x nb_ft x t\n\n    x = self.dropout(x).transpose(1, 2)  # BS x t x nb_ft\n    x = self.fc(x)  # BS x t x h  (h = 1024)\n\n    x = x.transpose(1, 2)  # BS x h x t\n    clipwise_output, norm_att, segmentwise_output = self.att_block(\n        x\n    )  # BS x num_classes, BS x num_classes x t, BS x num_classes x t\n\n    segmentwise_output = segmentwise_output.transpose(\n        1, 2\n    )  # BS x t x num_classes\n\n    # Upscale back to original size\n    framewise_output = interpolate(segmentwise_output, ratio=32)  # BS x T x num_classes\n\n    return framewise_output, clipwise_output\n</code></pre>",
      "votes": 12,
      "replies": [
        {
          "id": 1151613,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2021-01-13T13:01:21.683000",
          "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> do you happen to know what's the reason everyone sums max and avg pooling? And why not to go with only one of them or concatenation?</p>\n<p>I did some experiments on simple image datasets and summing avg and max was never the best option. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1151632,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2021-01-13T13:14:17.087000",
          "content": "<p>That's how it's done in the original repo <a href=\"https://github.com/qiuqiangkong/audioset_tagging_cnn\" target=\"_blank\">https://github.com/qiuqiangkong/audioset_tagging_cnn</a><br>\nChanging that is probably a good idea indeed</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1152517,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2021-01-14T08:40:46.457000",
          "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> I have actually done a simple test, but only on a single fold.<br>\nI tried solo average and concatenate. Solo average was not good. </p>\n<p>Although you should take this with a grain of salt, it makes sense to me because max tends to work very good with small objects which is the case here. </p>\n<p>Authors tuned the params on AudioSet which perhaps has many long and short interval sounds in which case combination of max and avg would make sense. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1152583,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2021-01-14T09:37:48.947000",
          "content": "<p>Thanks for sharing,<br>\nI've noticed that results vary a lot which makes it hard to draw conclusions. I'm guessing focusing on small changes such as this one won't make a big difference in the end</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1152604,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2021-01-14T10:06:57.480000",
          "content": "<p>Could not agree more.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1162726,
      "author_name": "Volodymyr",
      "author_url": "",
      "post_date": "2021-01-21T09:52:40.543000",
      "content": "<p>Really nice write-up. Also I have found code of SED models from the winner of the previous Sound competition<br>\n <a href=\"https://github.com/ryanwongsa/kaggle-birdsong-recognition/blob/master/src/models/sed_models.py\" target=\"_blank\">https://github.com/ryanwongsa/kaggle-birdsong-recognition/blob/master/src/models/sed_models.py</a></p>",
      "votes": 6,
      "replies": [
        {
          "id": 1162958,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-21T12:28:31.060000",
          "content": "<p>Thank you.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1192149,
      "author_name": "Wonkwang",
      "author_url": "",
      "post_date": "2021-02-09T02:25:06.437000",
      "content": "<p>Thanks for sharing the kernel. In your case, how about using focal loss instead of BCE?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1192225,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-02-09T03:48:01.747000",
          "content": "<p>I don't know much about focal loss. But I think focal loss is strong in the imbalance data.<br>\nThis competition data is <strong>not</strong> big imbalance. Therefore If you use focal loss instead of BCE, it won't work. Maybe…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1173831,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2021-01-28T06:05:49.163000",
      "content": "<p>Thanks for sharing this model explanation. For the curious, the SED paperswithcode page gives some details: <a href=\"https://paperswithcode.com/task/sound-event-detection\" target=\"_blank\">https://paperswithcode.com/task/sound-event-detection</a>. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1151311,
      "author_name": "Shubham Thapa",
      "author_url": "",
      "post_date": "2021-01-13T08:29:28.070000",
      "content": "<p>for how much epoch did you train the model to get above 0.9x and the time period ?<br>\ni am training for  time periods such as 30 sec as it is improving my cv score though </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1151621,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-13T13:08:06.573000",
          "content": "<p>LB 0.9X is below condition.<br>\nepoch:30<br>\ntime period:10sec</p>\n<p>I tried 60sec ,10sec and 5sec period. The best one was 10sec period.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1151726,
          "author_name": "Shubham Thapa",
          "author_url": "",
          "post_date": "2021-01-13T14:14:14.243000",
          "content": "<p>thats great thanx for the reply  !</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1151742,
          "author_name": "Robert Kim",
          "author_url": "",
          "post_date": "2021-01-13T14:26:41.623000",
          "content": "<p>I am a bit confused, can you please elaborate on 'time period'. Is it a time interval on which we make frame_wise prediction?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1152328,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-14T03:20:38.033000",
          "content": "<p>'time period' is the length of clip to make training data. I used 10sec 'time period'.<br>\nAnd I get random 10 sec clip from 60sec training data.</p>\n<p>In inference time, I split 60sec test clip per 10 sec.<br>\nAnd I get mini prediction(max(framewise_output, axis=time)) per 10 sec.<br>\nFinal prediction is max of mini prediction(=np.max((6,24), axis=0)).</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1152422,
          "author_name": "Sergey Verbitskiy",
          "author_url": "",
          "post_date": "2021-01-14T06:18:25.430000",
          "content": "<p>But i think that getting max of framewise_output for 60-sec clip and then applying sigmoid is the equal to \"Final prediction is max of mini prediction\". The difference is only getting max after/before activation (sigmoid), but sigmoid is the monotonic function.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1152503,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-14T08:19:45.097000",
          "content": "<p><a href=\"https://www.kaggle.com/sergeyverbitskiy\" target=\"_blank\">@sergeyverbitskiy</a> </p>\n<p>I experimented below condition(train: weak label, inference: framewise_output).</p>\n<ul>\n<li>train:10sec clip / inference:60 sec clip</li>\n<li>train:10sec clip / inference:(10 sec clip) * 6</li>\n</ul>\n<p>The latter is good score. So I used max of 6 mini predictions.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1152596,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-01-14T09:56:23.287000",
          "content": "<p>I am sure your 0.9+ LB score uses some other tricks :)</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1152598,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2021-01-14T10:01:23.343000",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> I share your opinion but I wonder how can you be sure :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1152616,
          "author_name": "Shubham Thapa",
          "author_url": "",
          "post_date": "2021-01-14T10:19:20.747000",
          "content": "<p>trick might be heavy augmentations , i did some heavy aug and it improves my cv  ridiculosly , havent tried lb yet :) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1152642,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-01-14T10:39:46.147000",
          "content": "<p>Because apparently people have not been able to replicate it, and there are no public kernels with 0.9+</p>\n<p><a href=\"https://www.kaggle.com/trooperog\" target=\"_blank\">@trooperog</a> interesting, I think OP mentioned more augs didnt help him. Where are you applying augs?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1152671,
          "author_name": "Shubham Thapa",
          "author_url": "",
          "post_date": "2021-01-14T10:58:52.837000",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> i am applying augs to the raw waveforms , results are pretty intresting earlier without the augs i was overfitting as i my train metric was above 0.9+ and validation was stuck at 0.82x and after applying augs validation is reaching to 0.88x in 20 epochs of training  .</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 1159828,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2021-01-19T13:55:58.920000",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Looks like you've figured out the \"tricks\" :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1163553,
      "author_name": "NakedKoala",
      "author_url": "",
      "post_date": "2021-01-21T18:39:59.607000",
      "content": "<p><a href=\"https://www.kaggle.com/shinmura0\" target=\"_blank\">@shinmura0</a></p>\n<p>Just wanna confirm if I understand all details correctly.</p>\n<p>So when we compute train and val loss. We are doing <code>loss(clipwise_output, target)</code>.</p>\n<p>But when we compute LWRAP metric and generating prediction we are doing <code>metric(framewise_output.max(dim=time), target)</code>.</p>\n<p>Is my understanding correct ? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1163776,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-21T23:24:52.357000",
          "content": "<p>Yes. You are correct.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1163784,
          "author_name": "NakedKoala",
          "author_url": "",
          "post_date": "2021-01-21T23:48:14.777000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1368955%2F5de886aab541ee5e47736e75cfbf7fc1%2FScreen%20Shot%202021-01-21%20at%206.42.10%20PM.png?generation=1611272580923419&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/shinmura0\" target=\"_blank\">@shinmura0</a> </p>\n<p>Did you empirically validate that max(framewise_output) is superior to clipwise_output or did you just make the decision based on in tuition ? </p>\n<p>I am also doing the same SED architecture and trying to replicate your 0.90 + performance. Above chart is what I found. </p>\n<p>I am doing validation on entire image (separate validation image to 6 10secs clip  and do same thing as inference ), not just a close crop of annotated event. It looks like lwlrap_clipwise is consistently better than lwlrap_framewise, for local validation.</p>\n<p>What do you think ?  Is local validation completely garbage so that I should not trust this finding. or is it likely because that I am doing something wrong with my implementation</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1163131,
      "author_name": "Volodymyr",
      "author_url": "",
      "post_date": "2021-01-21T14:58:33.117000",
      "content": "<p>Also I would like to ask about some points:</p>\n<ul>\n<li>You have mentioned that you take random crop with <code>tp</code> label.  Do you apply the same strategy for validation set. If so, I am a bit confused, because your validation score will be different for different runs on the same model and data</li>\n<li>You have stated that you have models with 0.89 - 0.91 LB score. What is your CV score for this model. And do you compute CV only on crops with <code>tp</code> labels or you make prediction as on test (6 pieces, each 10 seconds and max(axis=time)) and then compute LWLRAP on the whole recording?<br>\nThanks in advance!</li>\n</ul>",
      "votes": 2,
      "replies": [
        {
          "id": 1164224,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-22T08:15:57.387000",
          "content": "<blockquote>\n  <p>Do you apply the same strategy for validation set.</p>\n</blockquote>\n<p>Yes. Your suggestion is correct. However, since seed is fixed, the validation clip is the same in every experiment, and I think the missing labels are the reason why the validation is not stable.<br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209684\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209684</a></p>\n<blockquote>\n  <p>And do you compute CV only on crops with tp labels</p>\n</blockquote>\n<p>Yes. I used 10sec random clip for validation. I don't use the whole recording(60sec) for validation. Because there is missing labels in the whole recording.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1164369,
          "author_name": "Volodymyr",
          "author_url": "",
          "post_date": "2021-01-22T10:28:10.710000",
          "content": "<p>Thanks for your answer!<br>\nAlso I am bit confused with logic of taking a big enough clip:<br>\nIf we have lots of missed values, we can assume that the probability of having missed value in our training clip increasing with the length of the clip. So it is logically to use smaller clips</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1164929,
          "author_name": "NakedKoala",
          "author_url": "",
          "post_date": "2021-01-22T16:16:41.483000",
          "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> </p>\n<p>I was using the crop validation as well but I found it dramatically overestimating the LB score.</p>\n<p>In my setup, I was getting 0.87 in crop based validation but LB is only 0.74 - 0.78. </p>\n<p>Then, I switch to clip based validation and the CV - LB gap becomes much smaller. </p>\n<p>I understand your logic on using crop instead of full clip for validation ( missing labels --&gt; Shorter clip has smaller chance of missing -&gt; more trust-worthy number ) </p>\n<p>I am curious on this: Do you find crop based validation to be more correlated with LB score than clip based validation score , even though it is overly optimistic ? </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1162693,
      "author_name": "aqx",
      "author_url": "",
      "post_date": "2021-01-21T09:26:02.760000",
      "content": "<p>Thank you for the explanation! I recently just started trying out the SED architecture and i am confused over some of the details. My understanding is that :</p>\n<ol>\n<li>During training, randomly crop the clip (e.g 10s) and use all the labels associated with the entire clip (but i seen implementation that used just the labels associated within the cropped timeframe)</li>\n<li>Compute loss over the <code>clipwise_output</code> </li>\n<li>Validate/Inference with the <code>framewise_output</code>.<br>\nIs this flow the correct logic? </li>\n</ol>",
      "votes": 2,
      "replies": [
        {
          "id": 1162986,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-21T12:43:10.570000",
          "content": "<p>2,3 : I think you are correct.</p>\n<blockquote>\n  <ol>\n  <li>the labels associated with the entire clip</li>\n  </ol>\n</blockquote>\n<p>Does \"entire clip\" mean 10sec crop?  labels is not given a whole clip(60sec). So training crop is made by containing tp segment.<br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830#1157578\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830#1157578</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1163027,
          "author_name": "aqx",
          "author_url": "",
          "post_date": "2021-01-21T13:13:51.467000",
          "content": "<p>Thank you for the above clarification, i really appreciate it! I have followed closely your discussion and what i understand is that </p>\n<ol>\n<li><p>e.g a 60s clip with labels [0,1,2,3,4] , and assuming the first 10s crop has labels [0,1,2]. From your discussion, the ground truth for the first 10s cropped segment should follow [0,1,2], instead of the whole [0,1,2,3,4].</p></li>\n<li><p>I came across this discussion <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/200922#1102470\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/200922#1102470</a> which i believe encourage the use of [0,1,2,3,4] instead. So i was pretty confused by this.    </p></li>\n</ol>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1158361,
      "author_name": "Manh Lab",
      "author_url": "",
      "post_date": "2021-01-18T14:25:11.217000",
      "content": "<p>Thanks for nice sharing! How did you crop the audio?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1159087,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-19T03:23:28.753000",
          "content": "<p>I used time shift and make a clip.<br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830#1154880\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830#1154880</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1157005,
      "author_name": "Alexander Soare",
      "author_url": "",
      "post_date": "2021-01-17T15:24:55.837000",
      "content": "<p>Thanks so much for taking the time to put this together for us. I have one quick question about your inference method. You are taking a max for framewise_output during inference. Why not just leave it to the attention mechanism that you trained up? That's its job after all.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1154760,
      "author_name": "Benjamin35",
      "author_url": "",
      "post_date": "2021-01-15T21:42:30.090000",
      "content": "<p>Hi! Nice idea ! Did you pre trained the model on audio set ? Or maybe have done semi supervised like Mixmatch with data from test set or rainforest audio on YouTube ? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1155593,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-16T13:42:58.900000",
          "content": "<blockquote>\n  <p>Did you pre trained the model on audio set ?<br>\n  No. But I think it is effective.</p>\n  <p>have done semi supervised like Mixmatch<br>\n  No. I don't touch test data. I train with tp data and use trained model's prediction for test data.</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1154599,
      "author_name": "Sinan Calisir",
      "author_url": "",
      "post_date": "2021-01-15T18:11:58.413000",
      "content": "<p>Hi, </p>\n<p>Thanks for your explanation. I was waiting for it 😄</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1152125,
      "author_name": "Trevor Dunlap",
      "author_url": "",
      "post_date": "2021-01-13T20:42:31.493000",
      "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> Thank you for this post! I have a question - since you train on Framewise Output (i.e. time and class) how are you inputting both time and class information into BCE? Are you simply taking bce(y_true_one_hot, np.max(framewise_output, axis=time))? When looking here (<a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection</a>) the Loss is based on clipwise output. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1152329,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-14T03:20:41.997000",
          "content": "<p>In reference notebook, training method is \"weak label training\". So clipwise output is fitted whole clip label. On the other hand, In \"strong label training\", you need to prepare time and class data like below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2Fee6a92ad7e3490e9a1611e824b49b18d%2Ffig4.png?generation=1610584861906771&amp;alt=media\" alt=\"\"></p>\n<p>In this style, you can train model using framewise output with BCE.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1158248,
          "author_name": "JIN",
          "author_url": "",
          "post_date": "2021-01-18T13:34:40.577000",
          "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> Thanks for sharing nice post ! I tried to build strong label training. loss was going down, but CV was stuck in the 0.2x while training. maybe I made a mistake.<br>\nIf there is 10 period, 1501 frames(320 hop_size), <code>framewise output</code> should be `(bs, 1501, 24). so if I feed (bs, 1501, 24) target in BCE with framewise output, It's gonna be strong label training. Am I understanding as a right way ?  Thanks :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1159086,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-19T03:22:53.373000",
          "content": "<p>Yes. Your are correct. <br>\nI also try to train the model with strong label. Loss and CV improved. But LB did not imporve.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1151445,
      "author_name": "RAHUL SINGH INDA",
      "author_url": "",
      "post_date": "2021-01-13T10:29:57.847000",
      "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a>  have you tried any audio augmentations (like Guassian noise etc.)? If any could you enlighten us ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1151617,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-13T13:04:42.873000",
          "content": "<p>I tried to add white noise and pink noise. But they didn't work.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1163289,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-21T16:15:13.230000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1163449,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-01-21T17:20:36.990000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1163484,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-01-21T17:37:55.167000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1151299": "I use SED(Sound Event Detection) and maybe many participants used it. But SED(PANNs architecture) is difficult to learn for audio competition beginner. I shortly Introduce how to use SED(only PANNs architecture).\n\nIn SED, prediction is classes and time like below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2F315015cca93d0b34eee61e27ec995516%2Ffig1.png?generation=1609987035311073&alt=media)\n\n# 1 How to train\nSED(PANNs architecture)'s output is two type. \n+ clipwise_output\n+ framewise_output\n\nframewise_output is the prediction result of **time and classes**. It can show above figure. Training with framewise_output need time and classes data. And loss function is usually binary cross entropy. Then this label data(contains time and class) is called **\"strong label\"**.\n\nclipwise_output is the prediction result for **classes of the whole clip**(for example 60sec). It contain no time information(only class information). Training with clipwise_output need only classes data. And loss function is also usually binary cross entropy. Then this label(contains only class) is called **\"weak label\"**.\n\nIn general, strong label training have higher accuracy than weak label training. Because strong labels contain time information.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2Fcfc1dc4d0aacee3b2513a7c44fc3b1d4%2Ffig1.png?generation=1610517949491883&alt=media)\n\n# 2 Feature extractor\n\nPANNs architecture is above. Feature extractor is the most Important. It is CNN14 in default PANNs architecture. CNN14 is strong model and pretrained model(by audio clip). But in this competition(and [last competition](https://www.kaggle.com/c/birdsong-recognition)), CNN14 is not strongest model. In my experiment, EfficientNet is better model. But the input of CNN14(1CH) is different from the one of EfficientNet(3CH). So you need to change the code.\n\nHere is a sample code(Reference code is [here](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection)).\n\n```\npip install efficientnet_pytorch\n```\n\n```\nfrom efficientnet_pytorch import EfficientNet\n\nclass feature_extractor(nn.Module):\n    def __init__(self, original):\n        super().__init__()\n        self.model = original\n    def forward(self, x):\n        x= self.model.extract_features(x)\n        return x\n    \nclass PANNsCNN14Att(nn.Module):\n    def __init__():\n        self.Enet = feature_extractor(EfficientNet.from_pretrained('efficientnet-b0'))\n    ...\n\n    def cnn_feature_extractor(self, x):\n        return self.Enet(x)\n\n    def forward(self, input, mixup_lambda=None):\n        x, frames_num = self.preprocess(input, mixup_lambda=mixup_lambda)\n        x = torch.cat((x,x,x),1)\n\n        x = self.cnn_feature_extractor(x)\n        \n        x = torch.mean(x, dim=3)\n        ...\n```\n\nIf errors happen, adjust attblock value or interpolation value.\n\n# 3 How to predict\nprediction procedure is also two type. \n+ clipwise_output\n+ framewise_output\n\nPrediction with clipwise_output is simple to use. But I think that it is weak for short event. Comparatively, framewise_output prediction is good at short sound event. In this competition, there is many short sound event. Therefore I use framewise_output for prediction. \n\nMy prediction strategy with framewise_outputis is below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2F7207d5745bb1627dcd9733ee7fe4f4a5%2Ffig3.png?generation=1610523433543850&alt=media)\n",
    "1151606": "If this helps anybody, this is the forward function of an SED model with what happens to the tensor sizes. The architecture is almost the same as the one [here](https://www.kaggle.com/gopidurgaprasad/rfcx-sed-model-stater).\n\n```\ndef forward(self, x):  # BS x C x F x T\n    x = self.encoder.extract_features(x)  # BS x nb_ft x f x t\n    x = torch.mean(x, dim=2)  # BS x nb_ft x t\n\n    x1 = F.max_pool1d(x, kernel_size=3, stride=1, padding=1)\n    x2 = F.avg_pool1d(x, kernel_size=3, stride=1, padding=1)\n    x = x1 + x2  # BS x nb_ft x t\n\n    x = self.dropout(x).transpose(1, 2)  # BS x t x nb_ft\n    x = self.fc(x)  # BS x t x h  (h = 1024)\n\n    x = x.transpose(1, 2)  # BS x h x t\n    clipwise_output, norm_att, segmentwise_output = self.att_block(\n        x\n    )  # BS x num_classes, BS x num_classes x t, BS x num_classes x t\n\n    segmentwise_output = segmentwise_output.transpose(\n        1, 2\n    )  # BS x t x num_classes\n\n    # Upscale back to original size\n    framewise_output = interpolate(segmentwise_output, ratio=32)  # BS x T x num_classes\n\n    return framewise_output, clipwise_output\n```",
    "1162726": "Really nice write-up. Also I have found code of SED models from the winner of the previous Sound competition\n https://github.com/ryanwongsa/kaggle-birdsong-recognition/blob/master/src/models/sed_models.py",
    "1192149": "Thanks for sharing the kernel. In your case, how about using focal loss instead of BCE?",
    "1173831": "Thanks for sharing this model explanation. For the curious, the SED paperswithcode page gives some details: https://paperswithcode.com/task/sound-event-detection. ",
    "1151311": "for how much epoch did you train the model to get above 0.9x and the time period ?\ni am training for  time periods such as 30 sec as it is improving my cv score though ",
    "1163553": "@shinmura0\n\nJust wanna confirm if I understand all details correctly.\n\nSo when we compute train and val loss. We are doing `loss(clipwise_output, target)`.\n\nBut when we compute LWRAP metric and generating prediction we are doing `metric(framewise_output.max(dim=time), target)`.\n\nIs my understanding correct ? ",
    "1163131": "Also I would like to ask about some points:\n- You have mentioned that you take random crop with `tp` label.  Do you apply the same strategy for validation set. If so, I am a bit confused, because your validation score will be different for different runs on the same model and data\n- You have stated that you have models with 0.89 - 0.91 LB score. What is your CV score for this model. And do you compute CV only on crops with `tp` labels or you make prediction as on test (6 pieces, each 10 seconds and max(axis=time)) and then compute LWLRAP on the whole recording?\nThanks in advance!",
    "1162693": "Thank you for the explanation! I recently just started trying out the SED architecture and i am confused over some of the details. My understanding is that :\n1. During training, randomly crop the clip (e.g 10s) and use all the labels associated with the entire clip (but i seen implementation that used just the labels associated within the cropped timeframe)\n2. Compute loss over the `clipwise_output` \n3. Validate/Inference with the `framewise_output`.\nIs this flow the correct logic? ",
    "1158361": "Thanks for nice sharing! How did you crop the audio?",
    "1157005": "Thanks so much for taking the time to put this together for us. I have one quick question about your inference method. You are taking a max for framewise_output during inference. Why not just leave it to the attention mechanism that you trained up? That's its job after all.",
    "1154760": "Hi! Nice idea ! Did you pre trained the model on audio set ? Or maybe have done semi supervised like Mixmatch with data from test set or rainforest audio on YouTube ? ",
    "1154599": "Hi, \n\nThanks for your explanation. I was waiting for it 😄",
    "1152125": "@shinmurashinmura Thank you for this post! I have a question - since you train on Framewise Output (i.e. time and class) how are you inputting both time and class information into BCE? Are you simply taking bce(y_true_one_hot, np.max(framewise_output, axis=time))? When looking here (https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection) the Loss is based on clipwise output. ",
    "1151445": "@shinmurashinmura  have you tried any audio augmentations (like Guassian noise etc.)? If any could you enlighten us ?",
    "1163289": ""
  }
}