{
  "id": 315860,
  "title": "PCEN: a better pre-processing technique of STFT spectrograms",
  "url": "/competitions/birdclef-2022/discussion/315860",
  "author_name": "",
  "post_date": "2022-03-30T04:13:00.407695700Z",
  "votes": 33,
  "comment_count": 11,
  "views": 0,
  "content": "<h1>Description</h1>\n<p>I will share a technique which is called Per-channel energy normalization (PCEN)[1].<br>\nIn short, it processes time-frequency domain signals more sophisticated way than simple log-compression. </p>\n<p>The main feature is:</p>\n<p>1) \"Gaussianization\": make intensity of time-frequency bins close to Gaussian distribution<br>\n2) \"Whitening\": disentangle inter-channel correlations</p>\n<p>This technique was also used by several teams in the DCASE 2021 task5[2] and reported it contributed to improved performance.</p>\n<p>I also released a notebook on PCEN[3], so check it out if you are interested.</p>\n<h1>Reference</h1>\n<p>[1] <a href=\"https://arxiv.org/abs/1607.05666\" target=\"_blank\">Wang, Y., Getreuer, P., Hughes, T., Lyon, R. F., &amp; Saurous, R. A. (2017, March). Trainable frontend for robust and far-field keyword spotting. In Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on (pp. 5670-5674). IEEE.</a><br>\n[2] <a href=\"https://dcase.community/challenge2021/task-few-shot-bioacoustic-event-detection\" target=\"_blank\">DCASE 2021 task 5, Few-shot Bioacoustic Event Detection</a><br>\n[3] <a href=\"https://www.kaggle.com/code/tatamikenn/birdclef-22-per-channel-energy-normalization\" target=\"_blank\">BirdCLEF'22: Per-channel energy normalization\n</a></p>\n<p><a href=\"https://ibb.co/zZX5mYw\"><img src=\"https://i.ibb.co/9WbvpLj/pcen.png\" alt=\"pcen\"></a></p>",
  "messages": [
    {
      "id": "1739461",
      "postDate": "03/30/2022 04:13:00",
      "content": "<h1>Description</h1>\n<p>I will share a technique which is called Per-channel energy normalization (PCEN)[1].<br>\nIn short, it processes time-frequency domain signals more sophisticated way than simple log-compression. </p>\n<p>The main feature is:</p>\n<p>1) \"Gaussianization\": make intensity of time-frequency bins close to Gaussian distribution<br>\n2) \"Whitening\": disentangle inter-channel correlations</p>\n<p>This technique was also used by several teams in the DCASE 2021 task5[2] and reported it contributed to improved performance.</p>\n<p>I also released a notebook on PCEN[3], so check it out if you are interested.</p>\n<h1>Reference</h1>\n<p>[1] <a href=\"https://arxiv.org/abs/1607.05666\" target=\"_blank\">Wang, Y., Getreuer, P., Hughes, T., Lyon, R. F., &amp; Saurous, R. A. (2017, March). Trainable frontend for robust and far-field keyword spotting. In Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on (pp. 5670-5674). IEEE.</a><br>\n[2] <a href=\"https://dcase.community/challenge2021/task-few-shot-bioacoustic-event-detection\" target=\"_blank\">DCASE 2021 task 5, Few-shot Bioacoustic Event Detection</a><br>\n[3] <a href=\"https://www.kaggle.com/code/tatamikenn/birdclef-22-per-channel-energy-normalization\" target=\"_blank\">BirdCLEF'22: Per-channel energy normalization\n</a></p>\n<p><a href=\"https://ibb.co/zZX5mYw\"><img src=\"https://i.ibb.co/9WbvpLj/pcen.png\" alt=\"pcen\"></a></p>",
      "rawMarkdown": "# Description\n\nI will share a technique which is called Per-channel energy normalization (PCEN)[1].\nIn short, it processes time-frequency domain signals more sophisticated way than simple log-compression. \n\nThe main feature is:\n\n1) \"Gaussianization\": make intensity of time-frequency bins close to Gaussian distribution\n2) \"Whitening\": disentangle inter-channel correlations\n\nThis technique was also used by several teams in the DCASE 2021 task5[2] and reported it contributed to improved performance.\n\nI also released a notebook on PCEN[3], so check it out if you are interested.\n\n# Reference\n\n[1] [Wang, Y., Getreuer, P., Hughes, T., Lyon, R. F., & Saurous, R. A. (2017, March). Trainable frontend for robust and far-field keyword spotting. In Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on (pp. 5670-5674). IEEE.](https://arxiv.org/abs/1607.05666)\n[2] [DCASE 2021 task 5, Few-shot Bioacoustic Event Detection](https://dcase.community/challenge2021/task-few-shot-bioacoustic-event-detection)\n[3] [BirdCLEF'22: Per-channel energy normalization\n](https://www.kaggle.com/code/tatamikenn/birdclef-22-per-channel-energy-normalization)\n\n<a href=\"https://ibb.co/zZX5mYw\"><img src=\"https://i.ibb.co/9WbvpLj/pcen.png\" alt=\"pcen\" border=\"0\"></a>",
      "votes": null
    },
    {
      "id": "1741748",
      "postDate": "04/01/2022 05:45:44",
      "content": "<p>I tried this method and it improved my score. Thank you for sharing!</p>",
      "rawMarkdown": "I tried this method and it improved my score. Thank you for sharing!",
      "votes": null
    },
    {
      "id": "1741757",
      "postDate": "04/01/2022 05:55:55",
      "content": "<p>That’s good. I think this technique is becoming de facto standard for the audio task. Thanks for commenting.</p>",
      "rawMarkdown": "That’s good. I think this technique is becoming de facto standard for the audio task. Thanks for commenting.",
      "votes": null
    },
    {
      "id": "1742312",
      "postDate": "04/01/2022 17:10:57",
      "content": "<p>This is really interesting information. Thanks for sharing :)</p>",
      "rawMarkdown": "This is really interesting information. Thanks for sharing :)",
      "votes": null
    },
    {
      "id": "1747142",
      "postDate": "04/06/2022 12:14:06",
      "content": "<p>Good work, thank you </p>",
      "rawMarkdown": "Good work, thank you",
      "votes": null
    },
    {
      "id": "1794200",
      "postDate": "05/18/2022 15:28:56",
      "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> PCEN seems to be giving a slightly lower score for me. Is it working for your model? The images are so much more crystal clear, but somehow looks like model prefers the hazier spect. </p>",
      "rawMarkdown": "tatamikenn PCEN seems to be giving a slightly lower score for me. Is it working for your model? The images are so much more crystal clear, but somehow looks like model prefers the hazier spect.",
      "votes": null
    },
    {
      "id": "1794241",
      "postDate": "05/18/2022 15:59:23",
      "content": "<p><a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a> As a matter of fact, I have not tried PCEN on my model yet, although I have implemented it. Personally, I think it would be better to apply PCEN stochastically, since some noise may be necessary during training. It should be better always applying it during inference.</p>",
      "rawMarkdown": "allohvk As a matter of fact, I have not tried PCEN on my model yet, although I have implemented it. Personally, I think it would be better to apply PCEN stochastically, since some noise may be necessary during training. It should be better always applying it during inference.",
      "votes": null
    },
    {
      "id": "1794243",
      "postDate": "05/18/2022 16:05:28",
      "content": "<p>In the 2020 Cornell bird call identification competition, the 6th place solution introduces the method inputting PCEN and log compression on different channels, making 3-channel input model. If you can afford it, it would be good to try this method as well.</p>",
      "rawMarkdown": "In the 2020 Cornell bird call identification competition, the 6th place solution introduces the method inputting PCEN and log compression on different channels, making 3-channel input model. If you can afford it, it would be good to try this method as well.",
      "votes": null
    },
    {
      "id": "1794248",
      "postDate": "05/18/2022 16:12:30",
      "content": "<p>However, I think it is difficult to judge the effectiveness of PCEN alone from the LB score, since the choice of threshold is much more critical on this competition.</p>",
      "rawMarkdown": "However, I think it is difficult to judge the effectiveness of PCEN alone from the LB score, since the choice of threshold is much more critical on this competition.",
      "votes": null
    },
    {
      "id": "1794285",
      "postDate": "05/18/2022 16:38:18",
      "content": "<p>Thanks. Your comments make sense! I am a bit under the weather these days but if I get better I will definitely try it out. Until then it is a simple \"tweak the threshold and submit\" for me. Dont have the  energy to try out anything radically different.<br>\nI assumed PCEN normalization may not be same as conventional normalization that we do, but the jury is open on that one. I think there have been experiments trying out normalizing channelwise as well as across channels with both giving mixed results.</p>\n<p>I think the solution is to somehow use it stochastically as u suggested</p>",
      "rawMarkdown": "Thanks. Your comments make sense! I am a bit under the weather these days but if I get better I will definitely try it out. Until then it is a simple \"tweak the threshold and submit\" for me. Dont have the  energy to try out anything radically different.\nI assumed PCEN normalization may not be same as conventional normalization that we do, but the jury is open on that one. I think there have been experiments trying out normalizing channelwise as well as across channels with both giving mixed results.\n\nI think the solution is to somehow use it stochastically as u suggested",
      "votes": null
    },
    {
      "id": "1794431",
      "postDate": "05/18/2022 20:34:06",
      "content": "<p><a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a> 👍 Hope you will get well.</p>",
      "rawMarkdown": "allohvk 👍 Hope you will get well.",
      "votes": null
    },
    {
      "id": "1797560",
      "postDate": "05/22/2022 04:54:35",
      "content": "<p>One more thing to note: the hyper parameters chosen in the notebook might not necessarily optimal. If you implement PCEN as a neural network layer, you can learn hyper parameter to adapt to the dataset. There are also some open-source implementations available.</p>",
      "rawMarkdown": "One more thing to note: the hyper parameters chosen in the notebook might not necessarily optimal. If you implement PCEN as a neural network layer, you can learn hyper parameter to adapt to the dataset. There are also some open-source implementations available.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1741748,
      "author_name": "kucsikz",
      "author_url": "",
      "post_date": "04/01/2022 05:45:44",
      "content": "<p>I tried this method and it improved my score. Thank you for sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1741757,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "04/01/2022 05:55:55",
          "content": "<p>That’s good. I think this technique is becoming de facto standard for the audio task. Thanks for commenting.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1794200,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "05/18/2022 15:28:56",
          "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> PCEN seems to be giving a slightly lower score for me. Is it working for your model? The images are so much more crystal clear, but somehow looks like model prefers the hazier spect. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1794241,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "05/18/2022 15:59:23",
          "content": "<p><a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a> As a matter of fact, I have not tried PCEN on my model yet, although I have implemented it. Personally, I think it would be better to apply PCEN stochastically, since some noise may be necessary during training. It should be better always applying it during inference.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1794243,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "05/18/2022 16:05:28",
          "content": "<p>In the 2020 Cornell bird call identification competition, the 6th place solution introduces the method inputting PCEN and log compression on different channels, making 3-channel input model. If you can afford it, it would be good to try this method as well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1794248,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "05/18/2022 16:12:30",
          "content": "<p>However, I think it is difficult to judge the effectiveness of PCEN alone from the LB score, since the choice of threshold is much more critical on this competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1794285,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "05/18/2022 16:38:18",
          "content": "<p>Thanks. Your comments make sense! I am a bit under the weather these days but if I get better I will definitely try it out. Until then it is a simple \"tweak the threshold and submit\" for me. Dont have the  energy to try out anything radically different.<br>\nI assumed PCEN normalization may not be same as conventional normalization that we do, but the jury is open on that one. I think there have been experiments trying out normalizing channelwise as well as across channels with both giving mixed results.</p>\n<p>I think the solution is to somehow use it stochastically as u suggested</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1794431,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "05/18/2022 20:34:06",
          "content": "<p><a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a> 👍 Hope you will get well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1797560,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "05/22/2022 04:54:35",
          "content": "<p>One more thing to note: the hyper parameters chosen in the notebook might not necessarily optimal. If you implement PCEN as a neural network layer, you can learn hyper parameter to adapt to the dataset. There are also some open-source implementations available.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1742312,
      "author_name": "gwanghan",
      "author_url": "",
      "post_date": "04/01/2022 17:10:57",
      "content": "<p>This is really interesting information. Thanks for sharing :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1747142,
      "author_name": "hichambellafkir",
      "author_url": "",
      "post_date": "04/06/2022 12:14:06",
      "content": "<p>Good work, thank you </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1739461": "# Description\n\nI will share a technique which is called Per-channel energy normalization (PCEN)[1].\nIn short, it processes time-frequency domain signals more sophisticated way than simple log-compression. \n\nThe main feature is:\n\n1) \"Gaussianization\": make intensity of time-frequency bins close to Gaussian distribution\n2) \"Whitening\": disentangle inter-channel correlations\n\nThis technique was also used by several teams in the DCASE 2021 task5[2] and reported it contributed to improved performance.\n\nI also released a notebook on PCEN[3], so check it out if you are interested.\n\n# Reference\n\n[1] [Wang, Y., Getreuer, P., Hughes, T., Lyon, R. F., & Saurous, R. A. (2017, March). Trainable frontend for robust and far-field keyword spotting. In Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on (pp. 5670-5674). IEEE.](https://arxiv.org/abs/1607.05666)\n[2] [DCASE 2021 task 5, Few-shot Bioacoustic Event Detection](https://dcase.community/challenge2021/task-few-shot-bioacoustic-event-detection)\n[3] [BirdCLEF'22: Per-channel energy normalization\n](https://www.kaggle.com/code/tatamikenn/birdclef-22-per-channel-energy-normalization)\n\n<a href=\"https://ibb.co/zZX5mYw\"><img src=\"https://i.ibb.co/9WbvpLj/pcen.png\" alt=\"pcen\" border=\"0\"></a>",
    "1741748": "I tried this method and it improved my score. Thank you for sharing!",
    "1741757": "That’s good. I think this technique is becoming de facto standard for the audio task. Thanks for commenting.",
    "1742312": "This is really interesting information. Thanks for sharing :)",
    "1747142": "Good work, thank you",
    "1794200": "tatamikenn PCEN seems to be giving a slightly lower score for me. Is it working for your model? The images are so much more crystal clear, but somehow looks like model prefers the hazier spect.",
    "1794241": "allohvk As a matter of fact, I have not tried PCEN on my model yet, although I have implemented it. Personally, I think it would be better to apply PCEN stochastically, since some noise may be necessary during training. It should be better always applying it during inference.",
    "1794243": "In the 2020 Cornell bird call identification competition, the 6th place solution introduces the method inputting PCEN and log compression on different channels, making 3-channel input model. If you can afford it, it would be good to try this method as well.",
    "1794248": "However, I think it is difficult to judge the effectiveness of PCEN alone from the LB score, since the choice of threshold is much more critical on this competition.",
    "1794285": "Thanks. Your comments make sense! I am a bit under the weather these days but if I get better I will definitely try it out. Until then it is a simple \"tweak the threshold and submit\" for me. Dont have the  energy to try out anything radically different.\nI assumed PCEN normalization may not be same as conventional normalization that we do, but the jury is open on that one. I think there have been experiments trying out normalizing channelwise as well as across channels with both giving mixed results.\n\nI think the solution is to somehow use it stochastically as u suggested",
    "1794431": "allohvk 👍 Hope you will get well.",
    "1797560": "One more thing to note: the hyper parameters chosen in the notebook might not necessarily optimal. If you implement PCEN as a neural network layer, you can learn hyper parameter to adapt to the dataset. There are also some open-source implementations available."
  },
  "source": "meta"
}