{
  "id": 213418,
  "title": "Help needed in understanding the PANNs SED arch",
  "url": "/competitions/rfcx-species-audio-detection/discussion/213418",
  "author_name": "",
  "post_date": "2021-01-22T18:14:58.382998900Z",
  "votes": 11,
  "comment_count": 4,
  "views": 0,
  "content": "<p>In this competition, I have been using the PANNs SED architecture.</p>\n<p>I reference these 2 notebooks for model implementation:</p>\n<p><a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">Introduction to Sound Event Detection</a></p>\n<p><a href=\"https://www.kaggle.com/gopidurgaprasad/rfcx-sed-model-stater\" target=\"_blank\">SDE Starter</a></p>\n<p>And I have read these papers to try understanding the intuition and rationale behind some of the architecture details:</p>\n<p><a href=\"https://arxiv.org/abs/1912.10211\" target=\"_blank\">PANNs paper</a><br>\n<a href=\"http://www.cs.cmu.edu/~yunwang/papers/cmu-thesis.pdf\" target=\"_blank\">Polyphonic Sound Event Detection with Weak Labeling Paper</a></p>\n<p>Strangely, <a href=\"https://arxiv.org/abs/1912.10211\" target=\"_blank\">PANNs paper</a> talks mostly about the backbone CNN feature extractor and contains no detail on the SED related details of the architecture.</p>\n<p>As a result, I still have a lot of questions about why things are done in certain way in the architecture. I am new to SED and Audio related competition… maybe I just lack basic knowledge in this field.. But I think there are many people in the same boat as myself. So I just would like to post my doubts here and hope that discussion / resource pointer could help clarify things, so every SED arch user benefits. ( Understand little details better empowers meaningful arch modification) </p>\n<p>All Codes below copied from <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">Introduction to Sound Event Detection</a></p>\n<h3>Detail 1: Applying 2d batch norm on frequency axis</h3>\n<pre><code>def preprocess(self, input, mixup_lambda=None):\n        # t1 = time.time()\n        x = self.spectrogram_extractor(input)  # (batch_size, 1, time_steps, freq_bins)\n        x = self.logmel_extractor(x)  # (batch_size, 1, time_steps, mel_bins)\n\n        frames_num = x.shape[2]\n\n        x = x.transpose(1, 3)\n        x = self.bn0(x)\n        x = x.transpose(1, 3)\n</code></pre>\n<p>Can someone explain the intuition behind this step ( last 3 lines)  ? If this thing has a name in the literature, can you please reference it so that I can do further reading on it.</p>\n<h3>Detail 2: Mean aggregation over frequency Axis ?</h3>\n<pre><code>def forward(self, input, mixup_lambda=None):\n        \"\"\"\n        Input: (batch_size, data_length)\"\"\"\n        x, frames_num = self.preprocess(input, mixup_lambda=mixup_lambda)\n\n        # Output shape (batch size, channels, time, frequency)\n        x = self.cnn_feature_extractor(x)\n\n        # Aggregate in frequency axis\n        x = torch.mean(x, dim=3)\n</code></pre>\n<p>Not understanding purpose of mean aggregation here as well.   Is this a common practice in SED ? <br>\nIf so, any pointer to resource explaining its effect ?   What about other aggregations different from mean ?  Would it be useful ? </p>\n<h3>Detail3:  Apply dense layer over channel axis after pooling</h3>\n<pre><code>def forward(self, input, mixup_lambda=None):\n        \"\"\"\n        Input: (batch_size, data_length)\"\"\"\n        x, frames_num = self.preprocess(input, mixup_lambda=mixup_lambda)\n\n        # Output shape (batch size, channels, time, frequency)\n        x = self.cnn_feature_extractor(x)\n\n        # Aggregate in frequency axis\n        x = torch.mean(x, dim=3)\n\n        x1 = F.max_pool1d(x, kernel_size=3, stride=1, padding=1)\n        x2 = F.avg_pool1d(x, kernel_size=3, stride=1, padding=1)\n        x = x1 + x2\n\n        x = F.dropout(x, p=0.5, training=self.training)\n        x = x.transpose(1, 2)\n        x = F.relu_(self.fc1(x))\n        x = x.transpose(1, 2)\n        x = F.dropout(x, p=0.5, training=self.training)\n</code></pre>\n<p>What is the purpose and intuition for last 5 lines ?   Does this thing have a name ? </p>\n<h3>Details4: Interpolate</h3>\n<pre><code>        (clipwise_output, norm_att, segmentwise_output) = self.att_block(x)\n        segmentwise_output = segmentwise_output.transpose(1, 2)\n\n        # Get framewise output\n        framewise_output = interpolate(segmentwise_output,\n                                       self.interpolate_ratio)\n        framewise_output = pad_framewise_output(framewise_output, frames_num)\n\n        output_dict = {\n            'framewise_output': framewise_output,\n            'clipwise_output': clipwise_output\n        }\n\n        return output_dict\n</code></pre>\n<p>Although the code comment has explaination  - \"Interpolate data in time domain. This is used to compensate the resolution reduction in downsampling of a CNN \". I still don't feel I understand it deeply enough.  A pointer to the original paper where this technique was proposed would be helpful </p>\n<p>SED gurus  <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a>   <a href=\"https://www.kaggle.com/shinmura0\" target=\"_blank\">@shinmura0</a> , would you guys be able to comment and educate me and other on these details ?  Thanks </p>",
  "messages": [
    {
      "id": "1165139",
      "postDate": "01/22/2021 18:14:58",
      "content": "<p>In this competition, I have been using the PANNs SED architecture.</p>\n<p>I reference these 2 notebooks for model implementation:</p>\n<p><a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">Introduction to Sound Event Detection</a></p>\n<p><a href=\"https://www.kaggle.com/gopidurgaprasad/rfcx-sed-model-stater\" target=\"_blank\">SDE Starter</a></p>\n<p>And I have read these papers to try understanding the intuition and rationale behind some of the architecture details:</p>\n<p><a href=\"https://arxiv.org/abs/1912.10211\" target=\"_blank\">PANNs paper</a><br>\n<a href=\"http://www.cs.cmu.edu/~yunwang/papers/cmu-thesis.pdf\" target=\"_blank\">Polyphonic Sound Event Detection with Weak Labeling Paper</a></p>\n<p>Strangely, <a href=\"https://arxiv.org/abs/1912.10211\" target=\"_blank\">PANNs paper</a> talks mostly about the backbone CNN feature extractor and contains no detail on the SED related details of the architecture.</p>\n<p>As a result, I still have a lot of questions about why things are done in certain way in the architecture. I am new to SED and Audio related competition… maybe I just lack basic knowledge in this field.. But I think there are many people in the same boat as myself. So I just would like to post my doubts here and hope that discussion / resource pointer could help clarify things, so every SED arch user benefits. ( Understand little details better empowers meaningful arch modification) </p>\n<p>All Codes below copied from <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">Introduction to Sound Event Detection</a></p>\n<h3>Detail 1: Applying 2d batch norm on frequency axis</h3>\n<pre><code>def preprocess(self, input, mixup_lambda=None):\n        # t1 = time.time()\n        x = self.spectrogram_extractor(input)  # (batch_size, 1, time_steps, freq_bins)\n        x = self.logmel_extractor(x)  # (batch_size, 1, time_steps, mel_bins)\n\n        frames_num = x.shape[2]\n\n        x = x.transpose(1, 3)\n        x = self.bn0(x)\n        x = x.transpose(1, 3)\n</code></pre>\n<p>Can someone explain the intuition behind this step ( last 3 lines)  ? If this thing has a name in the literature, can you please reference it so that I can do further reading on it.</p>\n<h3>Detail 2: Mean aggregation over frequency Axis ?</h3>\n<pre><code>def forward(self, input, mixup_lambda=None):\n        \"\"\"\n        Input: (batch_size, data_length)\"\"\"\n        x, frames_num = self.preprocess(input, mixup_lambda=mixup_lambda)\n\n        # Output shape (batch size, channels, time, frequency)\n        x = self.cnn_feature_extractor(x)\n\n        # Aggregate in frequency axis\n        x = torch.mean(x, dim=3)\n</code></pre>\n<p>Not understanding purpose of mean aggregation here as well.   Is this a common practice in SED ? <br>\nIf so, any pointer to resource explaining its effect ?   What about other aggregations different from mean ?  Would it be useful ? </p>\n<h3>Detail3:  Apply dense layer over channel axis after pooling</h3>\n<pre><code>def forward(self, input, mixup_lambda=None):\n        \"\"\"\n        Input: (batch_size, data_length)\"\"\"\n        x, frames_num = self.preprocess(input, mixup_lambda=mixup_lambda)\n\n        # Output shape (batch size, channels, time, frequency)\n        x = self.cnn_feature_extractor(x)\n\n        # Aggregate in frequency axis\n        x = torch.mean(x, dim=3)\n\n        x1 = F.max_pool1d(x, kernel_size=3, stride=1, padding=1)\n        x2 = F.avg_pool1d(x, kernel_size=3, stride=1, padding=1)\n        x = x1 + x2\n\n        x = F.dropout(x, p=0.5, training=self.training)\n        x = x.transpose(1, 2)\n        x = F.relu_(self.fc1(x))\n        x = x.transpose(1, 2)\n        x = F.dropout(x, p=0.5, training=self.training)\n</code></pre>\n<p>What is the purpose and intuition for last 5 lines ?   Does this thing have a name ? </p>\n<h3>Details4: Interpolate</h3>\n<pre><code>        (clipwise_output, norm_att, segmentwise_output) = self.att_block(x)\n        segmentwise_output = segmentwise_output.transpose(1, 2)\n\n        # Get framewise output\n        framewise_output = interpolate(segmentwise_output,\n                                       self.interpolate_ratio)\n        framewise_output = pad_framewise_output(framewise_output, frames_num)\n\n        output_dict = {\n            'framewise_output': framewise_output,\n            'clipwise_output': clipwise_output\n        }\n\n        return output_dict\n</code></pre>\n<p>Although the code comment has explaination  - \"Interpolate data in time domain. This is used to compensate the resolution reduction in downsampling of a CNN \". I still don't feel I understand it deeply enough.  A pointer to the original paper where this technique was proposed would be helpful </p>\n<p>SED gurus  <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a>   <a href=\"https://www.kaggle.com/shinmura0\" target=\"_blank\">@shinmura0</a> , would you guys be able to comment and educate me and other on these details ?  Thanks </p>",
      "rawMarkdown": "In this competition, I have been using the PANNs SED architecture.\n\nI reference these 2 notebooks for model implementation:\n\n[Introduction to Sound Event Detection](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection)\n\n[SDE Starter](https://www.kaggle.com/gopidurgaprasad/rfcx-sed-model-stater)\n\nAnd I have read these papers to try understanding the intuition and rationale behind some of the architecture details:\n\n[PANNs paper](https://arxiv.org/abs/1912.10211)\n[Polyphonic Sound Event Detection with Weak Labeling Paper](http://www.cs.cmu.edu/~yunwang/papers/cmu-thesis.pdf)\n\nStrangely, [PANNs paper](https://arxiv.org/abs/1912.10211) talks mostly about the backbone CNN feature extractor and contains no detail on the SED related details of the architecture.\n\nAs a result, I still have a lot of questions about why things are done in certain way in the architecture. I am new to SED and Audio related competition... maybe I just lack basic knowledge in this field.. But I think there are many people in the same boat as myself. So I just would like to post my doubts here and hope that discussion / resource pointer could help clarify things, so every SED arch user benefits. ( Understand little details better empowers meaningful arch modification) \n\n\nAll Codes below copied from [Introduction to Sound Event Detection](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection)\n\n\n### Detail 1: Applying 2d batch norm on frequency axis\n\n\n```\ndef preprocess(self, input, mixup_lambda=None):\n        # t1 = time.time()\n        x = self.spectrogram_extractor(input)  # (batch_size, 1, time_steps, freq_bins)\n        x = self.logmel_extractor(x)  # (batch_size, 1, time_steps, mel_bins)\n\n        frames_num = x.shape[2]\n\n        x = x.transpose(1, 3)\n        x = self.bn0(x)\n        x = x.transpose(1, 3)\n```\nCan someone explain the intuition behind this step ( last 3 lines)  ? If this thing has a name in the literature, can you please reference it so that I can do further reading on it.\n\n### Detail 2: Mean aggregation over frequency Axis ? \n\n```\ndef forward(self, input, mixup_lambda=None):\n        \"\"\"\n        Input: (batch_size, data_length)\"\"\"\n        x, frames_num = self.preprocess(input, mixup_lambda=mixup_lambda)\n\n        # Output shape (batch size, channels, time, frequency)\n        x = self.cnn_feature_extractor(x)\n        \n        # Aggregate in frequency axis\n        x = torch.mean(x, dim=3)\n```\n\nNot understanding purpose of mean aggregation here as well.   Is this a common practice in SED ? \nIf so, any pointer to resource explaining its effect ?   What about other aggregations different from mean ?  Would it be useful ? \n\n### Detail3:  Apply dense layer over channel axis after pooling\n\n```\ndef forward(self, input, mixup_lambda=None):\n        \"\"\"\n        Input: (batch_size, data_length)\"\"\"\n        x, frames_num = self.preprocess(input, mixup_lambda=mixup_lambda)\n\n        # Output shape (batch size, channels, time, frequency)\n        x = self.cnn_feature_extractor(x)\n        \n        # Aggregate in frequency axis\n        x = torch.mean(x, dim=3)\n\n        x1 = F.max_pool1d(x, kernel_size=3, stride=1, padding=1)\n        x2 = F.avg_pool1d(x, kernel_size=3, stride=1, padding=1)\n        x = x1 + x2\n\n        x = F.dropout(x, p=0.5, training=self.training)\n        x = x.transpose(1, 2)\n        x = F.relu_(self.fc1(x))\n        x = x.transpose(1, 2)\n        x = F.dropout(x, p=0.5, training=self.training)\n\n```\n\nWhat is the purpose and intuition for last 5 lines ?   Does this thing have a name ? \n\n\n### Details4: Interpolate\n\n```\n        (clipwise_output, norm_att, segmentwise_output) = self.att_block(x)\n        segmentwise_output = segmentwise_output.transpose(1, 2)\n\n        # Get framewise output\n        framewise_output = interpolate(segmentwise_output,\n                                       self.interpolate_ratio)\n        framewise_output = pad_framewise_output(framewise_output, frames_num)\n\n        output_dict = {\n            'framewise_output': framewise_output,\n            'clipwise_output': clipwise_output\n        }\n\n        return output_dict\n```\n\nAlthough the code comment has explaination  - \"Interpolate data in time domain. This is used to compensate the resolution reduction in downsampling of a CNN \". I still don't feel I understand it deeply enough.  A pointer to the original paper where this technique was proposed would be helpful \n\nSED gurus  @hidehisaarai1213   @shinmura0 , would you guys be able to comment and educate me and other on these details ?  Thanks",
      "votes": null
    },
    {
      "id": "1199857",
      "postDate": "02/14/2021 06:38:38",
      "content": "<p>a pity this remains unanswered…some of them are pertinent questions..</p>",
      "rawMarkdown": "a pity this remains unanswered...some of them are pertinent questions..",
      "votes": null
    },
    {
      "id": "1199895",
      "postDate": "02/14/2021 08:06:28",
      "content": "<p>Sorry for not answering for a while…I was off kaggle a bit January and unaware of the mention.<br>\nSince, I'm also not an expert on this domain (I just knew the PANNs paper and worked a bit with it - still not an expert), I'm not sure whether I can answer correctly</p>\n<ol>\n<li>Applying 2d batch norm on frequency axis</li>\n</ol>\n<p>Not sure…I just used the original implementation. Maybe to deal with noise on log-melspectrogram?</p>\n<ol>\n<li>Detail 2: Mean aggregation over frequency Axis ?</li>\n</ol>\n<p>I think this is only to make the result tensor shape (batch_size, n_channels, n_frames). You can use other aggregation methods like taking max over this axis or combination of mean and max, whatever. Note that at this point, the output of Convolution layers is downsized a lot, so the shape along with frequency axis is 2 or so in original PANNs architecture(Cnn14).</p>\n<ol>\n<li>Detail3: Apply dense layer over channel axis after pooling</li>\n</ol>\n<p>Also not sure…maybe to add some nonlinearity?I just followed the original implementation. I think this kind of implementation details often does not have deep meaning.</p>\n<ol>\n<li>Details4: Interpolate</li>\n</ol>\n<p>The reason for this is quite simple - the output of CNN is quite downscaled, and this part just make the shape of the tensor the same as that of the input along with time axis. It does not add useful information (actually it just repeat the original tensor) but this helps the training implementation easier. With this interpolation, we only need to prepare label whose shape is the same as the input if we are to use strong (time-annotated) label. Otherwise we need to calculate the output tensor shape to make the label. Also, this is similar to the interpolation used in FCN for Semantic Segmentation. </p>",
      "rawMarkdown": "Sorry for not answering for a while...I was off kaggle a bit January and unaware of the mention.\nSince, I'm also not an expert on this domain (I just knew the PANNs paper and worked a bit with it - still not an expert), I'm not sure whether I can answer correctly\n\n1. Applying 2d batch norm on frequency axis\n\nNot sure...I just used the original implementation. Maybe to deal with noise on log-melspectrogram?\n\n2. Detail 2: Mean aggregation over frequency Axis ?\n\nI think this is only to make the result tensor shape (batch\\_size, n\\_channels, n\\_frames). You can use other aggregation methods like taking max over this axis or combination of mean and max, whatever. Note that at this point, the output of Convolution layers is downsized a lot, so the shape along with frequency axis is 2 or so in original PANNs architecture(Cnn14).\n\n3. Detail3: Apply dense layer over channel axis after pooling\n\nAlso not sure...maybe to add some nonlinearity?I just followed the original implementation. I think this kind of implementation details often does not have deep meaning.\n\n4. Details4: Interpolate\n\nThe reason for this is quite simple - the output of CNN is quite downscaled, and this part just make the shape of the tensor the same as that of the input along with time axis. It does not add useful information (actually it just repeat the original tensor) but this helps the training implementation easier. With this interpolation, we only need to prepare label whose shape is the same as the input if we are to use strong (time-annotated) label. Otherwise we need to calculate the output tensor shape to make the label. Also, this is similar to the interpolation used in FCN for Semantic Segmentation.",
      "votes": null
    },
    {
      "id": "1199947",
      "postDate": "02/14/2021 09:12:33",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a>! wish you all the best…u r close to top 10 :) still 3 more days and 15 attempts left</p>",
      "rawMarkdown": "Thanks @hidehisaarai1213! wish you all the best...u r close to top 10 :) still 3 more days and 15 attempts left",
      "votes": null
    },
    {
      "id": "1201176",
      "postDate": "02/15/2021 08:01:54",
      "content": "<p>Hi, </p>\n<p>I share my interpretations here (might not be correct).</p>\n<ol>\n<li><p>Applying 2d batch norm on frequency axis<br>\nAudio data are often mean-std normalization along the frequency dimension. If global mean and standard deviation vectors can not be computed before hand, 2d batch norm on the frequency axis is often used. </p></li>\n<li><p>Mean aggregation over frequency Axis?<br>\nThis is to convert the time-frequency output of the CNN in to a sequence in time. Other pooling methods are also applicable. </p></li>\n<li><p>Apply dense layer over channel axis after pooling<br>\nAs mentioned by <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a>, this is to introduce some non-linearity.</p></li>\n</ol>",
      "rawMarkdown": "Hi, \n\nI share my interpretations here (might not be correct).\n\n1. Applying 2d batch norm on frequency axis\nAudio data are often mean-std normalization along the frequency dimension. If global mean and standard deviation vectors can not be computed before hand, 2d batch norm on the frequency axis is often used. \n\n2. Mean aggregation over frequency Axis?\nThis is to convert the time-frequency output of the CNN in to a sequence in time. Other pooling methods are also applicable. \n\n3. Apply dense layer over channel axis after pooling\nAs mentioned by @hidehisaarai1213, this is to introduce some non-linearity.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1199857,
      "author_name": "allohvk",
      "author_url": "",
      "post_date": "02/14/2021 06:38:38",
      "content": "<p>a pity this remains unanswered…some of them are pertinent questions..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1199895,
      "author_name": "hidehisaarai1213",
      "author_url": "",
      "post_date": "02/14/2021 08:06:28",
      "content": "<p>Sorry for not answering for a while…I was off kaggle a bit January and unaware of the mention.<br>\nSince, I'm also not an expert on this domain (I just knew the PANNs paper and worked a bit with it - still not an expert), I'm not sure whether I can answer correctly</p>\n<ol>\n<li>Applying 2d batch norm on frequency axis</li>\n</ol>\n<p>Not sure…I just used the original implementation. Maybe to deal with noise on log-melspectrogram?</p>\n<ol>\n<li>Detail 2: Mean aggregation over frequency Axis ?</li>\n</ol>\n<p>I think this is only to make the result tensor shape (batch_size, n_channels, n_frames). You can use other aggregation methods like taking max over this axis or combination of mean and max, whatever. Note that at this point, the output of Convolution layers is downsized a lot, so the shape along with frequency axis is 2 or so in original PANNs architecture(Cnn14).</p>\n<ol>\n<li>Detail3: Apply dense layer over channel axis after pooling</li>\n</ol>\n<p>Also not sure…maybe to add some nonlinearity?I just followed the original implementation. I think this kind of implementation details often does not have deep meaning.</p>\n<ol>\n<li>Details4: Interpolate</li>\n</ol>\n<p>The reason for this is quite simple - the output of CNN is quite downscaled, and this part just make the shape of the tensor the same as that of the input along with time axis. It does not add useful information (actually it just repeat the original tensor) but this helps the training implementation easier. With this interpolation, we only need to prepare label whose shape is the same as the input if we are to use strong (time-annotated) label. Otherwise we need to calculate the output tensor shape to make the label. Also, this is similar to the interpolation used in FCN for Semantic Segmentation. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1199947,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "02/14/2021 09:12:33",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a>! wish you all the best…u r close to top 10 :) still 3 more days and 15 attempts left</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1201176,
      "author_name": "thomeou",
      "author_url": "",
      "post_date": "02/15/2021 08:01:54",
      "content": "<p>Hi, </p>\n<p>I share my interpretations here (might not be correct).</p>\n<ol>\n<li><p>Applying 2d batch norm on frequency axis<br>\nAudio data are often mean-std normalization along the frequency dimension. If global mean and standard deviation vectors can not be computed before hand, 2d batch norm on the frequency axis is often used. </p></li>\n<li><p>Mean aggregation over frequency Axis?<br>\nThis is to convert the time-frequency output of the CNN in to a sequence in time. Other pooling methods are also applicable. </p></li>\n<li><p>Apply dense layer over channel axis after pooling<br>\nAs mentioned by <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a>, this is to introduce some non-linearity.</p></li>\n</ol>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1165139": "In this competition, I have been using the PANNs SED architecture.\n\nI reference these 2 notebooks for model implementation:\n\n[Introduction to Sound Event Detection](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection)\n\n[SDE Starter](https://www.kaggle.com/gopidurgaprasad/rfcx-sed-model-stater)\n\nAnd I have read these papers to try understanding the intuition and rationale behind some of the architecture details:\n\n[PANNs paper](https://arxiv.org/abs/1912.10211)\n[Polyphonic Sound Event Detection with Weak Labeling Paper](http://www.cs.cmu.edu/~yunwang/papers/cmu-thesis.pdf)\n\nStrangely, [PANNs paper](https://arxiv.org/abs/1912.10211) talks mostly about the backbone CNN feature extractor and contains no detail on the SED related details of the architecture.\n\nAs a result, I still have a lot of questions about why things are done in certain way in the architecture. I am new to SED and Audio related competition... maybe I just lack basic knowledge in this field.. But I think there are many people in the same boat as myself. So I just would like to post my doubts here and hope that discussion / resource pointer could help clarify things, so every SED arch user benefits. ( Understand little details better empowers meaningful arch modification) \n\n\nAll Codes below copied from [Introduction to Sound Event Detection](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection)\n\n\n### Detail 1: Applying 2d batch norm on frequency axis\n\n\n```\ndef preprocess(self, input, mixup_lambda=None):\n        # t1 = time.time()\n        x = self.spectrogram_extractor(input)  # (batch_size, 1, time_steps, freq_bins)\n        x = self.logmel_extractor(x)  # (batch_size, 1, time_steps, mel_bins)\n\n        frames_num = x.shape[2]\n\n        x = x.transpose(1, 3)\n        x = self.bn0(x)\n        x = x.transpose(1, 3)\n```\nCan someone explain the intuition behind this step ( last 3 lines)  ? If this thing has a name in the literature, can you please reference it so that I can do further reading on it.\n\n### Detail 2: Mean aggregation over frequency Axis ? \n\n```\ndef forward(self, input, mixup_lambda=None):\n        \"\"\"\n        Input: (batch_size, data_length)\"\"\"\n        x, frames_num = self.preprocess(input, mixup_lambda=mixup_lambda)\n\n        # Output shape (batch size, channels, time, frequency)\n        x = self.cnn_feature_extractor(x)\n        \n        # Aggregate in frequency axis\n        x = torch.mean(x, dim=3)\n```\n\nNot understanding purpose of mean aggregation here as well.   Is this a common practice in SED ? \nIf so, any pointer to resource explaining its effect ?   What about other aggregations different from mean ?  Would it be useful ? \n\n### Detail3:  Apply dense layer over channel axis after pooling\n\n```\ndef forward(self, input, mixup_lambda=None):\n        \"\"\"\n        Input: (batch_size, data_length)\"\"\"\n        x, frames_num = self.preprocess(input, mixup_lambda=mixup_lambda)\n\n        # Output shape (batch size, channels, time, frequency)\n        x = self.cnn_feature_extractor(x)\n        \n        # Aggregate in frequency axis\n        x = torch.mean(x, dim=3)\n\n        x1 = F.max_pool1d(x, kernel_size=3, stride=1, padding=1)\n        x2 = F.avg_pool1d(x, kernel_size=3, stride=1, padding=1)\n        x = x1 + x2\n\n        x = F.dropout(x, p=0.5, training=self.training)\n        x = x.transpose(1, 2)\n        x = F.relu_(self.fc1(x))\n        x = x.transpose(1, 2)\n        x = F.dropout(x, p=0.5, training=self.training)\n\n```\n\nWhat is the purpose and intuition for last 5 lines ?   Does this thing have a name ? \n\n\n### Details4: Interpolate\n\n```\n        (clipwise_output, norm_att, segmentwise_output) = self.att_block(x)\n        segmentwise_output = segmentwise_output.transpose(1, 2)\n\n        # Get framewise output\n        framewise_output = interpolate(segmentwise_output,\n                                       self.interpolate_ratio)\n        framewise_output = pad_framewise_output(framewise_output, frames_num)\n\n        output_dict = {\n            'framewise_output': framewise_output,\n            'clipwise_output': clipwise_output\n        }\n\n        return output_dict\n```\n\nAlthough the code comment has explaination  - \"Interpolate data in time domain. This is used to compensate the resolution reduction in downsampling of a CNN \". I still don't feel I understand it deeply enough.  A pointer to the original paper where this technique was proposed would be helpful \n\nSED gurus  @hidehisaarai1213   @shinmura0 , would you guys be able to comment and educate me and other on these details ?  Thanks",
    "1199857": "a pity this remains unanswered...some of them are pertinent questions..",
    "1199895": "Sorry for not answering for a while...I was off kaggle a bit January and unaware of the mention.\nSince, I'm also not an expert on this domain (I just knew the PANNs paper and worked a bit with it - still not an expert), I'm not sure whether I can answer correctly\n\n1. Applying 2d batch norm on frequency axis\n\nNot sure...I just used the original implementation. Maybe to deal with noise on log-melspectrogram?\n\n2. Detail 2: Mean aggregation over frequency Axis ?\n\nI think this is only to make the result tensor shape (batch\\_size, n\\_channels, n\\_frames). You can use other aggregation methods like taking max over this axis or combination of mean and max, whatever. Note that at this point, the output of Convolution layers is downsized a lot, so the shape along with frequency axis is 2 or so in original PANNs architecture(Cnn14).\n\n3. Detail3: Apply dense layer over channel axis after pooling\n\nAlso not sure...maybe to add some nonlinearity?I just followed the original implementation. I think this kind of implementation details often does not have deep meaning.\n\n4. Details4: Interpolate\n\nThe reason for this is quite simple - the output of CNN is quite downscaled, and this part just make the shape of the tensor the same as that of the input along with time axis. It does not add useful information (actually it just repeat the original tensor) but this helps the training implementation easier. With this interpolation, we only need to prepare label whose shape is the same as the input if we are to use strong (time-annotated) label. Otherwise we need to calculate the output tensor shape to make the label. Also, this is similar to the interpolation used in FCN for Semantic Segmentation.",
    "1199947": "Thanks @hidehisaarai1213! wish you all the best...u r close to top 10 :) still 3 more days and 15 attempts left",
    "1201176": "Hi, \n\nI share my interpretations here (might not be correct).\n\n1. Applying 2d batch norm on frequency axis\nAudio data are often mean-std normalization along the frequency dimension. If global mean and standard deviation vectors can not be computed before hand, 2d batch norm on the frequency axis is often used. \n\n2. Mean aggregation over frequency Axis?\nThis is to convert the time-frequency output of the CNN in to a sequence in time. Other pooling methods are also applicable. \n\n3. Apply dense layer over channel axis after pooling\nAs mentioned by @hidehisaarai1213, this is to introduce some non-linearity."
  },
  "source": "meta"
}