{
  "id": 91859,
  "title": "Anyone tried PCEN in front end?",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/91859",
  "author_name": "Dick Lyon",
  "post_date": "2019-05-09T19:05:34.277000",
  "votes": 7,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Is anyone using PCEN in their Freesound Audio Tagging front end?  </p>\n\n<p>I'm been collecting success stories (and otherwise) about the \"per-channel energy normalization\" approach as alternative to the logarithm in mel filterbank front ends.  We've seen successes in keyword spotting, speech recognition, bird and whale call detection, speech dysarthria detection, etc., so it seems it might help more generally for sound classification.  If anyone tried it and didn't get a win, let me know and maybe I can suggest better parameters.  Or if you don't find an implementation you can use, let me know and I'll try to help.</p>\n\n<p>Some pointers:\nLot of papers, mostly successes from 2018/2019:  <a href=\"https://scholar.google.com/scholar?hl=en&amp;as_sdt=1%2C5&amp;q=pcen+%22per+channel+energy+normalization%22\">https://scholar.google.com/scholar?hl=en&amp;as_sdt=1%2C5&amp;q=pcen+%22per+channel+energy+normalization%22</a>\nA librosa implementation: <a href=\"https://librosa.github.io/librosa/generated/librosa.core.pcen.html\">https://librosa.github.io/librosa/generated/librosa.core.pcen.html</a>\nA Lasagne layer: <a href=\"https://gist.github.com/f0k/c837bcf0bfde189ca16eab63637839cb\">https://gist.github.com/f0k/c837bcf0bfde189ca16eab63637839cb</a></p>\n\n<p>Dick Lyon</p>",
  "messages": [
    {
      "id": 529386,
      "postDate": "2019-05-09T19:05:34.277Z",
      "content": "<p>Is anyone using PCEN in their Freesound Audio Tagging front end?  </p>\n\n<p>I'm been collecting success stories (and otherwise) about the \"per-channel energy normalization\" approach as alternative to the logarithm in mel filterbank front ends.  We've seen successes in keyword spotting, speech recognition, bird and whale call detection, speech dysarthria detection, etc., so it seems it might help more generally for sound classification.  If anyone tried it and didn't get a win, let me know and maybe I can suggest better parameters.  Or if you don't find an implementation you can use, let me know and I'll try to help.</p>\n\n<p>Some pointers:\nLot of papers, mostly successes from 2018/2019:  <a href=\"https://scholar.google.com/scholar?hl=en&amp;as_sdt=1%2C5&amp;q=pcen+%22per+channel+energy+normalization%22\">https://scholar.google.com/scholar?hl=en&amp;as_sdt=1%2C5&amp;q=pcen+%22per+channel+energy+normalization%22</a>\nA librosa implementation: <a href=\"https://librosa.github.io/librosa/generated/librosa.core.pcen.html\">https://librosa.github.io/librosa/generated/librosa.core.pcen.html</a>\nA Lasagne layer: <a href=\"https://gist.github.com/f0k/c837bcf0bfde189ca16eab63637839cb\">https://gist.github.com/f0k/c837bcf0bfde189ca16eab63637839cb</a></p>\n\n<p>Dick Lyon</p>",
      "rawMarkdown": "Is anyone using PCEN in their Freesound Audio Tagging front end?  \n\nI'm been collecting success stories (and otherwise) about the \"per-channel energy normalization\" approach as alternative to the logarithm in mel filterbank front ends.  We've seen successes in keyword spotting, speech recognition, bird and whale call detection, speech dysarthria detection, etc., so it seems it might help more generally for sound classification.  If anyone tried it and didn't get a win, let me know and maybe I can suggest better parameters.  Or if you don't find an implementation you can use, let me know and I'll try to help.\n\nSome pointers:\nLot of papers, mostly successes from 2018/2019:  https://scholar.google.com/scholar?hl=en&amp;as_sdt=1%2C5&amp;q=pcen+%22per+channel+energy+normalization%22\nA librosa implementation: https://librosa.github.io/librosa/generated/librosa.core.pcen.html\nA Lasagne layer: https://gist.github.com/f0k/c837bcf0bfde189ca16eab63637839cb\n\nDick Lyon",
      "votes": 7
    },
    {
      "id": 529662,
      "postDate": "2019-05-10T13:07:00.803Z",
      "content": "<p>I'm using it. Time will tell how useful it is.\nI've been reading your book. Humbled by your level of knowledge. Thank you for sharing it.</p>",
      "rawMarkdown": "I'm using it. Time will tell how useful it is.\nI've been reading your book. Humbled by your level of knowledge. Thank you for sharing it.",
      "votes": 1,
      "replies": [
        {
          "id": 529744,
          "postDate": "2019-05-10T16:16:39.367Z",
          "content": "<p>Thanks for the comment.  For others who might not know, a free online version can be found via my book blog at <a href=\"http://machinehearing.org\">http://machinehearing.org</a></p>",
          "rawMarkdown": "Thanks for the comment.  For others who might not know, a free online version can be found via my book blog at http://machinehearing.org",
          "votes": 3
        }
      ]
    },
    {
      "id": 563316,
      "postDate": "2019-06-28T06:14:06.457Z",
      "content": "<p>We used PCEN too. Parameters were the same as you recommended.\nSingle model with PCEN did not improve CV,\nbut boosted our LB about 0.01 when we ensemble with the model which trained with log mels.</p>",
      "rawMarkdown": "We used PCEN too. Parameters were the same as you recommended.\nSingle model with PCEN did not improve CV,\nbut boosted our LB about 0.01 when we ensemble with the model which trained with log mels."
    },
    {
      "id": 562664,
      "postDate": "2019-06-27T12:53:04.733Z",
      "content": "<p>Hello Dick,</p>\n\n<p>Is it meaningful to use pcen-filter on raw spectrogram instead of mel?</p>",
      "rawMarkdown": "Hello Dick,\n\nIs it meaningful to use pcen-filter on raw spectrogram instead of mel?"
    },
    {
      "id": 529736,
      "postDate": "2019-05-10T15:59:39.783Z",
      "content": "<p>I've replaced db with pcen, there is no improvement till now. </p>\n\n<p>Fig. 1 in the paper shows that pcen will enhance transient events (bird calls) while discarding\nstationary noise (insects) as well as slow changes in loudness (vehicle). </p>\n\n<p><a href=\"https://bmcfee.github.io/papers/spl2019_pcen.pdf\">https://bmcfee.github.io/papers/spl2019_pcen.pdf</a></p>",
      "rawMarkdown": "I've replaced db with pcen, there is no improvement till now. \n\nFig. 1 in the paper shows that pcen will enhance transient events (bird calls) while discarding\nstationary noise (insects) as well as slow changes in loudness (vehicle). \n\nhttps://bmcfee.github.io/papers/spl2019_pcen.pdf",
      "replies": [
        {
          "id": 529743,
          "postDate": "2019-05-10T16:15:32.833Z",
          "content": "<p>You might try reducing the parameters to closer to 0.  You can get as close to log as you like this way, and most likely find an optimum that's not all the way at the corner of the parameter space.  Show me what parameters you tried?</p>",
          "rawMarkdown": "You might try reducing the parameters to closer to 0.  You can get as close to log as you like this way, and most likely find an optimum that's not all the way at the corner of the parameter space.  Show me what parameters you tried?"
        },
        {
          "id": 529780,
          "postDate": "2019-05-10T18:23:07.580Z",
          "content": "<p>I just use librosa.pcen(S), with the default parameters. Maybe I should try others.</p>",
          "rawMarkdown": "I just use librosa.pcen(S), with the default parameters. Maybe I should try others."
        },
        {
          "id": 529792,
          "postDate": "2019-05-10T19:00:44.343Z",
          "content": "<p>That default set from our paper is a rather extreme setting for keyword detection of near and far talkers.  For recordings that have been recorded or somewhat adjusted to be in a reasonable dynamic range, as I expect we find at Freesound, less aggressive normalization (\"gain\" they call it here) is appropriate.</p>\n\n<p>instead of:  gain=0.98, bias=2, power=0.5, time_constant=0.4, eps=1e-06,</p>\n\n<p>try:  gain=0.5, bias=0.001, power=0.2, time_constant=0.4, eps=1e-9,</p>\n\n<p>The low values of all the parameters here will make a result closer to the log baseline.  When that works (hopefully not worse than log), try increasing them.  The BirdVox guys suggest a shorter time constant in their domain, but for Freesound, who knows?</p>\n\n<p>The trickiest parameter is the bias, because it depends on how your signals are scaled.  I'd suggest starting low (0.001) and also trying 0.01, 0.1, and 1.  Or do some stats on the intermediate partially normalized signals and pick a bias near the mean of those.  It may not be terribly sensitive, but you need to find the right ballpark.</p>\n\n<p>I'm guessing the signals are scaled such that the energies out of the mel filterbank are less than 1 (from waveforms in -1 to +1).  If waveforms are scaled -32K to +32K, like the integers in 16-bit .wav files, then the bias and eps numbers should be a lot higher.</p>\n\n<p>Dick</p>",
          "rawMarkdown": "That default set from our paper is a rather extreme setting for keyword detection of near and far talkers.  For recordings that have been recorded or somewhat adjusted to be in a reasonable dynamic range, as I expect we find at Freesound, less aggressive normalization (\"gain\" they call it here) is appropriate.\n\ninstead of:  gain=0.98, bias=2, power=0.5, time_constant=0.4, eps=1e-06,\n\ntry:  gain=0.5, bias=0.001, power=0.2, time_constant=0.4, eps=1e-9,\n\nThe low values of all the parameters here will make a result closer to the log baseline.  When that works (hopefully not worse than log), try increasing them.  The BirdVox guys suggest a shorter time constant in their domain, but for Freesound, who knows?\n\nThe trickiest parameter is the bias, because it depends on how your signals are scaled.  I'd suggest starting low (0.001) and also trying 0.01, 0.1, and 1.  Or do some stats on the intermediate partially normalized signals and pick a bias near the mean of those.  It may not be terribly sensitive, but you need to find the right ballpark.\n\nI'm guessing the signals are scaled such that the energies out of the mel filterbank are less than 1 (from waveforms in -1 to +1).  If waveforms are scaled -32K to +32K, like the integers in 16-bit .wav files, then the bias and eps numbers should be a lot higher.\n\nDick\n\n\n",
          "votes": 4
        },
        {
          "id": 529803,
          "postDate": "2019-05-10T19:44:32.003Z",
          "content": "<p>I will learn and try. Thanks a lot for your instruction.</p>",
          "rawMarkdown": "I will learn and try. Thanks a lot for your instruction."
        },
        {
          "id": 531978,
          "postDate": "2019-05-16T00:27:42.233Z",
          "content": "<p>Hi Dick, </p>\n\n<p>I haven’t used PCEN but I can add my small contribution about Freesound audio. Recordings are uploaded by a community of users and, although most of them are well-recorded, there are exceptions. Loudness differences are not uncommon; in fact, in the Freesound Annotator that we use to manually label data, we had to apply a loudness corrector to mitigate this effect, which was somewhat annoying while annotating.</p>\n\n<p>I’ve seen a few cases where the recordings have little energy, which can be similar to a far-field condition. I’ve also witnessed recordings that almost saturate. And I’ve also seen some clips that feature a (not so) small snippet of silence before/after the target event (for which PCEN could be more appropriate than log-mel). Again, it’s not that these cases happen a lot, but certainly they are not negligible. Just wanted to share this in case it helps to PCEN users.</p>\n\n<p>Considering the not-small amount of training data and the number of PCEN params, perhaps it's more reasonable to include PCEN in the first layer and learn its params through backprop?</p>\n\n<p>Eduardo</p>",
          "rawMarkdown": "Hi Dick, \n\nI haven’t used PCEN but I can add my small contribution about Freesound audio. Recordings are uploaded by a community of users and, although most of them are well-recorded, there are exceptions. Loudness differences are not uncommon; in fact, in the Freesound Annotator that we use to manually label data, we had to apply a loudness corrector to mitigate this effect, which was somewhat annoying while annotating.\n\nI’ve seen a few cases where the recordings have little energy, which can be similar to a far-field condition. I’ve also witnessed recordings that almost saturate. And I’ve also seen some clips that feature a (not so) small snippet of silence before/after the target event (for which PCEN could be more appropriate than log-mel). Again, it’s not that these cases happen a lot, but certainly they are not negligible. Just wanted to share this in case it helps to PCEN users.\n\nConsidering the not-small amount of training data and the number of PCEN params, perhaps it's more reasonable to include PCEN in the first layer and learn its params through backprop?\n\nEduardo\n ",
          "votes": 5
        },
        {
          "id": 558420,
          "postDate": "2019-06-22T09:28:45Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 564715,
          "postDate": "2019-06-29T21:14:46.813Z",
          "content": "<p>I used PCEN and I also tried tuning the parameters via gradient descent - see this kernel <a href=\"https://www.kaggle.com/simongrest/trainable-pcen-frontend-in-pytorch\">https://www.kaggle.com/simongrest/trainable-pcen-frontend-in-pytorch</a></p>\n\n<p>I had mixed success - using PCEN gave me a small bump over plain <code>power_to_db</code>, but I wasn't able get better results by making the PCEN parameters trainable.</p>\n\n<p>Thanks <a href=\"/dicklyon\">@dicklyon</a> for drawing our attention to this in the first place - I found your paper really interesting.</p>",
          "rawMarkdown": "I used PCEN and I also tried tuning the parameters via gradient descent - see this kernel https://www.kaggle.com/simongrest/trainable-pcen-frontend-in-pytorch\n\nI had mixed success - using PCEN gave me a small bump over plain `power_to_db`, but I wasn't able get better results by making the PCEN parameters trainable.\n\nThanks @dicklyon for drawing our attention to this in the first place - I found your paper really interesting."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 529662,
      "author_name": "robga",
      "author_url": "",
      "post_date": "2019-05-10T13:07:00.803000",
      "content": "<p>I'm using it. Time will tell how useful it is.\nI've been reading your book. Humbled by your level of knowledge. Thank you for sharing it.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 529744,
          "author_name": "Dick Lyon",
          "author_url": "",
          "post_date": "2019-05-10T16:16:39.367000",
          "content": "<p>Thanks for the comment.  For others who might not know, a free online version can be found via my book blog at <a href=\"http://machinehearing.org\">http://machinehearing.org</a></p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 563316,
      "author_name": "gege",
      "author_url": "",
      "post_date": "2019-06-28T06:14:06.457000",
      "content": "<p>We used PCEN too. Parameters were the same as you recommended.\nSingle model with PCEN did not improve CV,\nbut boosted our LB about 0.01 when we ensemble with the model which trained with log mels.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 562664,
      "author_name": "Oleh Shliazhko",
      "author_url": "",
      "post_date": "2019-06-27T12:53:04.733000",
      "content": "<p>Hello Dick,</p>\n\n<p>Is it meaningful to use pcen-filter on raw spectrogram instead of mel?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 529736,
      "author_name": "wqk",
      "author_url": "",
      "post_date": "2019-05-10T15:59:39.783000",
      "content": "<p>I've replaced db with pcen, there is no improvement till now. </p>\n\n<p>Fig. 1 in the paper shows that pcen will enhance transient events (bird calls) while discarding\nstationary noise (insects) as well as slow changes in loudness (vehicle). </p>\n\n<p><a href=\"https://bmcfee.github.io/papers/spl2019_pcen.pdf\">https://bmcfee.github.io/papers/spl2019_pcen.pdf</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 529743,
          "author_name": "Dick Lyon",
          "author_url": "",
          "post_date": "2019-05-10T16:15:32.833000",
          "content": "<p>You might try reducing the parameters to closer to 0.  You can get as close to log as you like this way, and most likely find an optimum that's not all the way at the corner of the parameter space.  Show me what parameters you tried?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 529780,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2019-05-10T18:23:07.580000",
          "content": "<p>I just use librosa.pcen(S), with the default parameters. Maybe I should try others.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 529792,
          "author_name": "Dick Lyon",
          "author_url": "",
          "post_date": "2019-05-10T19:00:44.343000",
          "content": "<p>That default set from our paper is a rather extreme setting for keyword detection of near and far talkers.  For recordings that have been recorded or somewhat adjusted to be in a reasonable dynamic range, as I expect we find at Freesound, less aggressive normalization (\"gain\" they call it here) is appropriate.</p>\n\n<p>instead of:  gain=0.98, bias=2, power=0.5, time_constant=0.4, eps=1e-06,</p>\n\n<p>try:  gain=0.5, bias=0.001, power=0.2, time_constant=0.4, eps=1e-9,</p>\n\n<p>The low values of all the parameters here will make a result closer to the log baseline.  When that works (hopefully not worse than log), try increasing them.  The BirdVox guys suggest a shorter time constant in their domain, but for Freesound, who knows?</p>\n\n<p>The trickiest parameter is the bias, because it depends on how your signals are scaled.  I'd suggest starting low (0.001) and also trying 0.01, 0.1, and 1.  Or do some stats on the intermediate partially normalized signals and pick a bias near the mean of those.  It may not be terribly sensitive, but you need to find the right ballpark.</p>\n\n<p>I'm guessing the signals are scaled such that the energies out of the mel filterbank are less than 1 (from waveforms in -1 to +1).  If waveforms are scaled -32K to +32K, like the integers in 16-bit .wav files, then the bias and eps numbers should be a lot higher.</p>\n\n<p>Dick</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 529803,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2019-05-10T19:44:32.003000",
          "content": "<p>I will learn and try. Thanks a lot for your instruction.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 531978,
          "author_name": "Eduardo Fonseca",
          "author_url": "",
          "post_date": "2019-05-16T00:27:42.233000",
          "content": "<p>Hi Dick, </p>\n\n<p>I haven’t used PCEN but I can add my small contribution about Freesound audio. Recordings are uploaded by a community of users and, although most of them are well-recorded, there are exceptions. Loudness differences are not uncommon; in fact, in the Freesound Annotator that we use to manually label data, we had to apply a loudness corrector to mitigate this effect, which was somewhat annoying while annotating.</p>\n\n<p>I’ve seen a few cases where the recordings have little energy, which can be similar to a far-field condition. I’ve also witnessed recordings that almost saturate. And I’ve also seen some clips that feature a (not so) small snippet of silence before/after the target event (for which PCEN could be more appropriate than log-mel). Again, it’s not that these cases happen a lot, but certainly they are not negligible. Just wanted to share this in case it helps to PCEN users.</p>\n\n<p>Considering the not-small amount of training data and the number of PCEN params, perhaps it's more reasonable to include PCEN in the first layer and learn its params through backprop?</p>\n\n<p>Eduardo</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 558420,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-22T09:28:45",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 564715,
          "author_name": "Simon Grest",
          "author_url": "",
          "post_date": "2019-06-29T21:14:46.813000",
          "content": "<p>I used PCEN and I also tried tuning the parameters via gradient descent - see this kernel <a href=\"https://www.kaggle.com/simongrest/trainable-pcen-frontend-in-pytorch\">https://www.kaggle.com/simongrest/trainable-pcen-frontend-in-pytorch</a></p>\n\n<p>I had mixed success - using PCEN gave me a small bump over plain <code>power_to_db</code>, but I wasn't able get better results by making the PCEN parameters trainable.</p>\n\n<p>Thanks <a href=\"/dicklyon\">@dicklyon</a> for drawing our attention to this in the first place - I found your paper really interesting.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "529386": "Is anyone using PCEN in their Freesound Audio Tagging front end?  \n\nI'm been collecting success stories (and otherwise) about the \"per-channel energy normalization\" approach as alternative to the logarithm in mel filterbank front ends.  We've seen successes in keyword spotting, speech recognition, bird and whale call detection, speech dysarthria detection, etc., so it seems it might help more generally for sound classification.  If anyone tried it and didn't get a win, let me know and maybe I can suggest better parameters.  Or if you don't find an implementation you can use, let me know and I'll try to help.\n\nSome pointers:\nLot of papers, mostly successes from 2018/2019:  https://scholar.google.com/scholar?hl=en&amp;as_sdt=1%2C5&amp;q=pcen+%22per+channel+energy+normalization%22\nA librosa implementation: https://librosa.github.io/librosa/generated/librosa.core.pcen.html\nA Lasagne layer: https://gist.github.com/f0k/c837bcf0bfde189ca16eab63637839cb\n\nDick Lyon",
    "529662": "I'm using it. Time will tell how useful it is.\nI've been reading your book. Humbled by your level of knowledge. Thank you for sharing it.",
    "563316": "We used PCEN too. Parameters were the same as you recommended.\nSingle model with PCEN did not improve CV,\nbut boosted our LB about 0.01 when we ensemble with the model which trained with log mels.",
    "562664": "Hello Dick,\n\nIs it meaningful to use pcen-filter on raw spectrogram instead of mel?",
    "529736": "I've replaced db with pcen, there is no improvement till now. \n\nFig. 1 in the paper shows that pcen will enhance transient events (bird calls) while discarding\nstationary noise (insects) as well as slow changes in loudness (vehicle). \n\nhttps://bmcfee.github.io/papers/spl2019_pcen.pdf"
  }
}