{
  "id": 171133,
  "title": "Why normalization matter when mixing bird sounds ?!",
  "url": "/competitions/birdsong-recognition/discussion/171133",
  "author_name": "kkiller",
  "post_date": "2020-07-30T14:57:03.390000",
  "votes": 30,
  "comment_count": 25,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1498852%2F0559ecb0ee51031face0724230c51090%2FXC178525_PER54_20190116_un_normalized.PNG?generation=1596121738571626&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1498852%2F3f1112de680c8ad46839bec0d1c52145%2FXC178525_PER54_20190116_normalized.PNG?generation=1596122490291744&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 952012,
      "postDate": "2020-07-30T14:57:03.390Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1498852%2F0559ecb0ee51031face0724230c51090%2FXC178525_PER54_20190116_un_normalized.PNG?generation=1596121738571626&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1498852%2F3f1112de680c8ad46839bec0d1c52145%2FXC178525_PER54_20190116_normalized.PNG?generation=1596122490291744&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1498852%2F0559ecb0ee51031face0724230c51090%2FXC178525_PER54_20190116_un_normalized.PNG?generation=1596121738571626&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1498852%2F3f1112de680c8ad46839bec0d1c52145%2FXC178525_PER54_20190116_normalized.PNG?generation=1596122490291744&amp;alt=media)\n",
      "votes": 30
    },
    {
      "id": 972051,
      "postDate": "2020-08-16T07:38:36.700Z",
      "content": "<p>as an illustration:</p>\n<p>top to bottom:<br>\nmelspec(noise)<br>\nmelspec(bird)<br>\nmelspec(noise+bird)<br>\nmelspec(noise)+melspec(bird)</p>\n<p>left: orginal spectrum, right: after my denoise algorithm</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F924b03b72af0252f68c00859d640c100%2FSelection_035.png?generation=1597563514500825&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "as an illustration:\n\ntop to bottom:\nmelspec(noise)\nmelspec(bird)\nmelspec(noise+bird)\nmelspec(noise)+melspec(bird)\n\nleft: orginal spectrum, right: after my denoise algorithm\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F924b03b72af0252f68c00859d640c100%2FSelection_035.png?generation=1597563514500825&alt=media)",
      "votes": 6,
      "replies": [
        {
          "id": 972511,
          "postDate": "2020-08-16T15:49:46.410Z",
          "content": "<p>finally some  validation results of original (left) vs 50% noise(right)</p>\n<p>noise creation = melspec( 0.5 x normalise(bird) +  0.5 x normalise(noise)) <br>\ngraylevel = framewise probability (trained with sigmoid loss). each frame width is about 150 ms</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F19a53ebbf14867ffc02328f5ef0d5e93%2FSelection_025.png?generation=1597592929981383&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fea8a5ea25c3fb38c4a12d411d347a181%2FSelection_028.png?generation=1597592958181521&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F579efb5326e485c4fe7280c2b4135c27%2FSelection_029.png?generation=1597592984235530&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "finally some  validation results of original (left) vs 50% noise(right)\n\nnoise creation = melspec( 0.5 x normalise(bird) +  0.5 x normalise(noise)) \ngraylevel = framewise probability (trained with sigmoid loss). each frame width is about 150 ms\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F19a53ebbf14867ffc02328f5ef0d5e93%2FSelection_025.png?generation=1597592929981383&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fea8a5ea25c3fb38c4a12d411d347a181%2FSelection_028.png?generation=1597592958181521&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F579efb5326e485c4fe7280c2b4135c27%2FSelection_029.png?generation=1597592984235530&alt=media)",
          "votes": 1,
          "replies": [
            {
              "id": 972557,
              "postDate": "2020-08-16T16:26:44.327Z",
              "content": "<p>Heng, from your pictures, it looks like adding noise makes the model worse… I can agree with those results because I tried adding Gaussian Noise, Pink Noise, and also background noise from other soundscape recordings and my LB score worsened, even after changing thresholds around.</p>",
              "rawMarkdown": "Heng, from your pictures, it looks like adding noise makes the model worse... I can agree with those results because I tried adding Gaussian Noise, Pink Noise, and also background noise from other soundscape recordings and my LB score worsened, even after changing thresholds around.",
              "votes": 1
            },
            {
              "id": 972560,
              "postDate": "2020-08-16T16:30:32.387Z",
              "content": "<p>the model was trained without noise.<br>\ni believed training with noise would improve results.</p>\n<p>i am trying self-supervised contrastive loss for noise training. results should be out in a few days.</p>\n<p>loss = binary cross entropy (wave, label) + contrastive loss  (wave+noise1, wave+noise1)</p>",
              "rawMarkdown": "the model was trained without noise.\ni believed training with noise would improve results.\n\ni am trying self-supervised contrastive loss for noise training. results should be out in a few days.\n\nloss = binary cross entropy (wave, label) + contrastive loss  (wave+noise1, wave+noise1)",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 952562,
      "postDate": "2020-07-31T03:45:32.993Z",
      "content": "<p><a href=\"https://arxiv.org/abs/1711.10282\">BC learning</a> may help you.\nOn page 4, It proposes equation (2), which is mixed to account for sound pressure level(G1 and G2).\nAnd in this case, r=0.5.</p>",
      "rawMarkdown": "[BC learning](https://arxiv.org/abs/1711.10282) may help you.\nOn page 4, It proposes equation (2), which is mixed to account for sound pressure level(G1 and G2).\nAnd in this case, r=0.5.",
      "votes": 5
    },
    {
      "id": 952551,
      "postDate": "2020-07-31T03:31:27.417Z",
      "content": "<p>Quite intuitive explanation!</p>\n\n<p>I haven't tried this yet but maybe <a href=\"https://scaper.readthedocs.io/en/latest/installation.html\">Scaper</a> is useful for mixing audio clips (this is also shared by <a href=\"/hengck23\">@hengck23</a> in this <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/169538\">post</a>)</p>\n\n<p>By the way, have you already tried mixing audio clips? If you don't mind can you share us how it works? About myself, I haven't tried yet but I think it worth trying.</p>",
      "rawMarkdown": "Quite intuitive explanation!\n\nI haven't tried this yet but maybe [Scaper](https://scaper.readthedocs.io/en/latest/installation.html) is useful for mixing audio clips (this is also shared by @hengck23 in this [post](https://www.kaggle.com/c/birdsong-recognition/discussion/169538))\n\nBy the way, have you already tried mixing audio clips? If you don't mind can you share us how it works? About myself, I haven't tried yet but I think it worth trying.",
      "votes": 6,
      "replies": [
        {
          "id": 952860,
          "postDate": "2020-07-31T09:19:38.820Z",
          "content": "<p>Thanks for your words <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> .</p>\n\n<p>Yes I've already tried mixing some bird songs. This stackoverflow <a href=\"https://stackoverflow.com/questions/14498539/how-to-overlay-downmix-two-audio-files-using-ffmpeg\">https://stackoverflow.com/questions/14498539/how-to-overlay-downmix-two-audio-files-using-ffmpeg</a> is useful if you wanna use <strong>ffmpeg</strong> which is a great tool (so far I'm calling it from command line and probably from python if it works). It doesn't seem to work wery well though so I have no model based on mixed sounds for instant. But I believe It might give some uplift !</p>",
          "rawMarkdown": "Thanks for your words @hidehisaarai1213 .\n\nYes I've already tried mixing some bird songs. This stackoverflow https://stackoverflow.com/questions/14498539/how-to-overlay-downmix-two-audio-files-using-ffmpeg is useful if you wanna use **ffmpeg** which is a great tool (so far I'm calling it from command line and probably from python if it works). It doesn't seem to work wery well though so I have no model based on mixed sounds for instant. But I believe It might give some uplift !",
          "votes": 2
        }
      ]
    },
    {
      "id": 952252,
      "postDate": "2020-07-30T18:52:56.417Z",
      "content": "<p>I might not have the correct answer for your question. but, i have a doubt with this implementation. you are concatenation(stacking) various waves one after another. in this case you will not get any overlap between different species(even though there are several bg species). \nas per the competition description, we are supposed to work on data with various species calling simultaniously. <br>\nmaybe a better approach should be mixing then into same timeframe .</p>",
      "rawMarkdown": "I might not have the correct answer for your question. but, i have a doubt with this implementation. you are concatenation(stacking) various waves one after another. in this case you will not get any overlap between different species(even though there are several bg species). \nas per the competition description, we are supposed to work on data with various species calling simultaniously.    \nmaybe a better approach should be mixing then into same timeframe .",
      "votes": 3,
      "replies": [
        {
          "id": 952333,
          "postDate": "2020-07-30T20:39:35.110Z",
          "content": "<p>You're right ! Keeping the sounds splitted make things easier to see ... the point is that if you mix these two bird songs without normalization, the louder one will shade the calmer.</p>",
          "rawMarkdown": "You're right ! Keeping the sounds splitted make things easier to see ... the point is that if you mix these two bird songs without normalization, the louder one will shade the calmer.",
          "votes": 4
        },
        {
          "id": 952590,
          "postDate": "2020-07-31T04:22:05.343Z",
          "content": "<p>&gt; the louder one will shade the calmer.</p>\n\n<p>About this, I have a little question. As <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> says: <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/170821#951101\">Annotations for test provides label for even a faint sound event</a>. So maybe mixing without normalization not a bad way?</p>",
          "rawMarkdown": "&gt; the louder one will shade the calmer.\n\nAbout this, I have a little question. As @hidehisaarai1213 says: [Annotations for test provides label for even a faint sound event](https://www.kaggle.com/c/birdsong-recognition/discussion/170821#951101). So maybe mixing without normalization not a bad way?",
          "votes": 2
        },
        {
          "id": 952853,
          "postDate": "2020-07-31T09:13:25.267Z",
          "content": "<p><a href=\"/karlyukang\">@karlyukang</a> , NO, by normalizing you'll put both birds songs at the same <strong>\"volume\"</strong>  and augment your chance to capture even faint sounds. </p>",
          "rawMarkdown": "@karlyukang , NO, by normalizing you'll put both birds songs at the same **\"volume\"**  and augment your chance to capture even faint sounds. ",
          "votes": 4
        },
        {
          "id": 952861,
          "postDate": "2020-07-31T09:19:47.507Z",
          "content": "<p><a href=\"/kneroma\">@kneroma</a> is right</p>\n\n<p>As for the faint signal, SNR is important.</p>",
          "rawMarkdown": "@kneroma is right\n\nAs for the faint signal, SNR is important.",
          "votes": 4
        },
        {
          "id": 952892,
          "postDate": "2020-07-31T09:42:55.123Z",
          "content": "<p><a href=\"/kneroma\">@kneroma</a> <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> \nGot it. Good point !</p>",
          "rawMarkdown": "@kneroma @hidehisaarai1213 \nGot it. Good point !",
          "votes": 2
        },
        {
          "id": 971229,
          "postDate": "2020-08-15T09:55:11.100Z",
          "content": "<p><a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> sorry, what's SNR?</p>",
          "rawMarkdown": "@hidehisaarai1213 sorry, what's SNR?"
        },
        {
          "id": 971233,
          "postDate": "2020-08-15T10:00:44.903Z",
          "content": "<p>Stands for Signal-to-Noise-Ratio.<br>\n<a href=\"https://en.wikipedia.org/wiki/Signal-to-noise_ratio#:~:text=Signal%2Dto%2Dnoise%20ratio%20(,power%2C%20often%20expressed%20in%20decibels\" target=\"_blank\">https://en.wikipedia.org/wiki/Signal-to-noise_ratio#:~:text=Signal%2Dto%2Dnoise%20ratio%20(,power%2C%20often%20expressed%20in%20decibels</a>.</p>",
          "rawMarkdown": "Stands for Signal-to-Noise-Ratio.\nhttps://en.wikipedia.org/wiki/Signal-to-noise_ratio#:~:text=Signal%2Dto%2Dnoise%20ratio%20(,power%2C%20often%20expressed%20in%20decibels.",
          "votes": 2
        }
      ]
    },
    {
      "id": 981890,
      "postDate": "2020-08-22T19:50:37.737Z",
      "content": "<p>Apologies if this is a silly question, but what normalization procedure do you use? From the images it looks like you're subtracting the mean of each waveform and dividing by the standard deviation but I could easily be wrong. I've come to fear the word \"normalize\" lately because usually when people use it they could plausibly be referring to ~eight different operations. </p>\n<p>The image below is created by:</p>\n<ol>\n<li>convert all waveforms to melspect</li>\n<li>add melspects elementwise</li>\n<li>divide melspect values by number of waveforms</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1577010%2Fafba0fba655e9d9dc6f64e80afc582d2%2FScreenshot%20from%202020-08-23%2005-55-34.png?generation=1598126362539466&amp;alt=media\" alt=\"\"></p>\n<p>The second image is </p>\n<ol>\n<li><code>standardize_relative</code> for all waveforms</li>\n<li>convert all waveforms to melspect</li>\n<li>add melspects elementwise</li>\n<li>divide melspect values by number of waveforms</li>\n</ol>\n<p>using </p>\n<pre><code>def standardize_relative(x):\n    return standardize(x, np.mean(x), np.std(x))\n\ndef standardize(x, mean, sd):\n    return (x - mean) / sd\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1577010%2F408505a480671190242680ffc077a675%2FScreenshot%20from%202020-08-23%2005-55-25.png?generation=1598126192473182&amp;alt=media\" alt=\"\"></p>\n<p>This final image is </p>\n<ol>\n<li><code>standardize_relative</code> for all waveforms</li>\n<li>add waveforms elementwise</li>\n<li>divide waveform values by number of waveforms </li>\n<li>convert final waveform to melspect</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1577010%2F63fe6cc81dd40975b44bab38f134841a%2FScreenshot%20from%202020-08-23%2006-03-56.png?generation=1598126678346710&amp;alt=media\" alt=\"\"></p>\n<p>The second one looks pretty decent to me, any suggestions for improvement?</p>\n<p>Below is some code to easily mix multiple waveforms using the second method if anyone is interested. All overlapping/mixing occurs at the start of the clip and if clips are different lengths the end of the mixed clip will be identical to the end of the longest clip.</p>\n<pre><code>import functools\n\ndef standardize_relative(x):\n    return standardize(x, np.mean(x), np.std(x))\n\ndef min_duration(signals):\n    return min([s.shape[1] for s in signals])\n\ndef standardize(x, mean, sd):\n    return (x - mean) / sd'\n\ndef mix(signals, sr, mel_fn):\n    mels =  [mel_fn(y=standardize_relative(i), sr=sr) for i in signals]\n    mixed = []\n    height = mels[0].shape[0]\n\n    while len(mels) &gt; 0:\n        shortest_len = min_duration(mels)\n\n        base = np.full((height, shortest_len), 0.0)\n\n        for s in mels:\n            base += s[:, :shortest_len] / len(mels)\n\n        mixed.append(base)\n\n        mels = [s[:, shortest_len:] for s in mels if s.shape[1] &gt; shortest_len]\n\n    return np.concatenate(mixed, axis=1)\n\n# a, b, c are all waveforms represented as numpy arrays\nfinal_mel = mix([a, b, c], 33075, functools.partial(librosa.feature.melspectrogram, n_fft=2048, hop_length=768, n_mels=424))\n</code></pre>",
      "rawMarkdown": "Apologies if this is a silly question, but what normalization procedure do you use? From the images it looks like you're subtracting the mean of each waveform and dividing by the standard deviation but I could easily be wrong. I've come to fear the word \"normalize\" lately because usually when people use it they could plausibly be referring to ~eight different operations. \n\nThe image below is created by:\n1. convert all waveforms to melspect\n2. add melspects elementwise\n3. divide melspect values by number of waveforms\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1577010%2Fafba0fba655e9d9dc6f64e80afc582d2%2FScreenshot%20from%202020-08-23%2005-55-34.png?generation=1598126362539466&alt=media)\n\n\nThe second image is \n0. `standardize_relative` for all waveforms\n1. convert all waveforms to melspect\n2. add melspects elementwise\n3. divide melspect values by number of waveforms\n\nusing \n\n```\ndef standardize_relative(x):\n    return standardize(x, np.mean(x), np.std(x))\n\ndef standardize(x, mean, sd):\n    return (x - mean) / sd\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1577010%2F408505a480671190242680ffc077a675%2FScreenshot%20from%202020-08-23%2005-55-25.png?generation=1598126192473182&alt=media)\n\n\nThis final image is \n\n0. `standardize_relative` for all waveforms\n1.  add waveforms elementwise\n2. divide waveform values by number of waveforms \n3. convert final waveform to melspect\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1577010%2F63fe6cc81dd40975b44bab38f134841a%2FScreenshot%20from%202020-08-23%2006-03-56.png?generation=1598126678346710&alt=media)\n\n\nThe second one looks pretty decent to me, any suggestions for improvement?\n\nBelow is some code to easily mix multiple waveforms using the second method if anyone is interested. All overlapping/mixing occurs at the start of the clip and if clips are different lengths the end of the mixed clip will be identical to the end of the longest clip.\n\n```\nimport functools\n\ndef standardize_relative(x):\n    return standardize(x, np.mean(x), np.std(x))\n\ndef min_duration(signals):\n    return min([s.shape[1] for s in signals])\n\ndef standardize(x, mean, sd):\n    return (x - mean) / sd'\n\ndef mix(signals, sr, mel_fn):\n    mels =  [mel_fn(y=standardize_relative(i), sr=sr) for i in signals]\n    mixed = []\n    height = mels[0].shape[0]\n\n    while len(mels) > 0:\n        shortest_len = min_duration(mels)\n\n        base = np.full((height, shortest_len), 0.0)\n\n        for s in mels:\n            base += s[:, :shortest_len] / len(mels)\n\n        mixed.append(base)\n\n        mels = [s[:, shortest_len:] for s in mels if s.shape[1] > shortest_len]\n    \n    return np.concatenate(mixed, axis=1)\n\n# a, b, c are all waveforms represented as numpy arrays\nfinal_mel = mix([a, b, c], 33075, functools.partial(librosa.feature.melspectrogram, n_fft=2048, hop_length=768, n_mels=424))\n```\n\n",
      "votes": 4
    },
    {
      "id": 952660,
      "postDate": "2020-07-31T05:53:19.713Z",
      "content": "<p>how a good topic👍 </p>",
      "rawMarkdown": "how a good topic👍 ",
      "replies": [
        {
          "id": 952848,
          "postDate": "2020-07-31T09:11:13.243Z",
          "content": "<p>Thanks <a href=\"/leila73\">@leila73</a> </p>",
          "rawMarkdown": "Thanks @leila73 "
        }
      ]
    },
    {
      "id": 952412,
      "postDate": "2020-07-30T22:27:25.003Z",
      "content": "<p>Good work!</p>",
      "rawMarkdown": "Good work!"
    },
    {
      "id": 952421,
      "postDate": "2020-07-30T22:38:43.713Z",
      "rawMarkdown": "",
      "votes": -4,
      "isDeleted": true,
      "replies": [
        {
          "id": 952458,
          "postDate": "2020-07-30T23:56:04.200Z",
          "content": "<p>How do you combine mel sepcs ?</p>\n\n<p><strong>PS</strong>:  MelSpecs(x1+x2) != MelSepcs(x1) + MelSpecs(x2) !</p>",
          "rawMarkdown": "How do you combine mel sepcs ?\n\n**PS**:  MelSpecs(x1+x2) != MelSepcs(x1) + MelSpecs(x2) !",
          "votes": 1
        },
        {
          "id": 952534,
          "postDate": "2020-07-31T03:01:27.440Z",
          "rawMarkdown": "",
          "votes": -18,
          "isDeleted": true
        },
        {
          "id": 952847,
          "postDate": "2020-07-31T09:10:41.007Z",
          "content": "<p>You're a very funny guy, keep it on dear 😂 </p>",
          "rawMarkdown": "You're a very funny guy, keep it on dear 😂 ",
          "votes": 6
        },
        {
          "id": 953219,
          "postDate": "2020-07-31T15:59:21.207Z",
          "rawMarkdown": "",
          "votes": -8,
          "isDeleted": true
        },
        {
          "id": 980964,
          "postDate": "2020-08-22T03:59:42.973Z",
          "content": "<p>I think people are downvoting you because you sound a little bit arrogant, just a heads up.</p>",
          "rawMarkdown": "I think people are downvoting you because you sound a little bit arrogant, just a heads up.",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 972051,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-08-16T07:38:36.700000",
      "content": "<p>as an illustration:</p>\n<p>top to bottom:<br>\nmelspec(noise)<br>\nmelspec(bird)<br>\nmelspec(noise+bird)<br>\nmelspec(noise)+melspec(bird)</p>\n<p>left: orginal spectrum, right: after my denoise algorithm</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F924b03b72af0252f68c00859d640c100%2FSelection_035.png?generation=1597563514500825&amp;alt=media\" alt=\"\"></p>",
      "votes": 6,
      "replies": [
        {
          "id": 972511,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-08-16T15:49:46.410000",
          "content": "<p>finally some  validation results of original (left) vs 50% noise(right)</p>\n<p>noise creation = melspec( 0.5 x normalise(bird) +  0.5 x normalise(noise)) <br>\ngraylevel = framewise probability (trained with sigmoid loss). each frame width is about 150 ms</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F19a53ebbf14867ffc02328f5ef0d5e93%2FSelection_025.png?generation=1597592929981383&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fea8a5ea25c3fb38c4a12d411d347a181%2FSelection_028.png?generation=1597592958181521&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F579efb5326e485c4fe7280c2b4135c27%2FSelection_029.png?generation=1597592984235530&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": [
            {
              "id": 972557,
              "author_name": "CoreyJamesLevinson",
              "author_url": "",
              "post_date": "2020-08-16T16:26:44.327000",
              "content": "<p>Heng, from your pictures, it looks like adding noise makes the model worse… I can agree with those results because I tried adding Gaussian Noise, Pink Noise, and also background noise from other soundscape recordings and my LB score worsened, even after changing thresholds around.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 972560,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2020-08-16T16:30:32.387000",
              "content": "<p>the model was trained without noise.<br>\ni believed training with noise would improve results.</p>\n<p>i am trying self-supervised contrastive loss for noise training. results should be out in a few days.</p>\n<p>loss = binary cross entropy (wave, label) + contrastive loss  (wave+noise1, wave+noise1)</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 952562,
      "author_name": "shinmura0",
      "author_url": "",
      "post_date": "2020-07-31T03:45:32.993000",
      "content": "<p><a href=\"https://arxiv.org/abs/1711.10282\">BC learning</a> may help you.\nOn page 4, It proposes equation (2), which is mixed to account for sound pressure level(G1 and G2).\nAnd in this case, r=0.5.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 952551,
      "author_name": "Hidehisa Arai",
      "author_url": "",
      "post_date": "2020-07-31T03:31:27.417000",
      "content": "<p>Quite intuitive explanation!</p>\n\n<p>I haven't tried this yet but maybe <a href=\"https://scaper.readthedocs.io/en/latest/installation.html\">Scaper</a> is useful for mixing audio clips (this is also shared by <a href=\"/hengck23\">@hengck23</a> in this <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/169538\">post</a>)</p>\n\n<p>By the way, have you already tried mixing audio clips? If you don't mind can you share us how it works? About myself, I haven't tried yet but I think it worth trying.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 952860,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2020-07-31T09:19:38.820000",
          "content": "<p>Thanks for your words <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> .</p>\n\n<p>Yes I've already tried mixing some bird songs. This stackoverflow <a href=\"https://stackoverflow.com/questions/14498539/how-to-overlay-downmix-two-audio-files-using-ffmpeg\">https://stackoverflow.com/questions/14498539/how-to-overlay-downmix-two-audio-files-using-ffmpeg</a> is useful if you wanna use <strong>ffmpeg</strong> which is a great tool (so far I'm calling it from command line and probably from python if it works). It doesn't seem to work wery well though so I have no model based on mixed sounds for instant. But I believe It might give some uplift !</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 952252,
      "author_name": "yuvaramsingh",
      "author_url": "",
      "post_date": "2020-07-30T18:52:56.417000",
      "content": "<p>I might not have the correct answer for your question. but, i have a doubt with this implementation. you are concatenation(stacking) various waves one after another. in this case you will not get any overlap between different species(even though there are several bg species). \nas per the competition description, we are supposed to work on data with various species calling simultaniously. <br>\nmaybe a better approach should be mixing then into same timeframe .</p>",
      "votes": 3,
      "replies": [
        {
          "id": 952333,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2020-07-30T20:39:35.110000",
          "content": "<p>You're right ! Keeping the sounds splitted make things easier to see ... the point is that if you mix these two bird songs without normalization, the louder one will shade the calmer.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 952590,
          "author_name": "Yu Kang",
          "author_url": "",
          "post_date": "2020-07-31T04:22:05.343000",
          "content": "<p>&gt; the louder one will shade the calmer.</p>\n\n<p>About this, I have a little question. As <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> says: <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/170821#951101\">Annotations for test provides label for even a faint sound event</a>. So maybe mixing without normalization not a bad way?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 952853,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2020-07-31T09:13:25.267000",
          "content": "<p><a href=\"/karlyukang\">@karlyukang</a> , NO, by normalizing you'll put both birds songs at the same <strong>\"volume\"</strong>  and augment your chance to capture even faint sounds. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 952861,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-07-31T09:19:47.507000",
          "content": "<p><a href=\"/kneroma\">@kneroma</a> is right</p>\n\n<p>As for the faint signal, SNR is important.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 952892,
          "author_name": "Yu Kang",
          "author_url": "",
          "post_date": "2020-07-31T09:42:55.123000",
          "content": "<p><a href=\"/kneroma\">@kneroma</a> <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> \nGot it. Good point !</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 971229,
          "author_name": "Marco Gorelli",
          "author_url": "",
          "post_date": "2020-08-15T09:55:11.100000",
          "content": "<p><a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> sorry, what's SNR?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 971233,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-08-15T10:00:44.903000",
          "content": "<p>Stands for Signal-to-Noise-Ratio.<br>\n<a href=\"https://en.wikipedia.org/wiki/Signal-to-noise_ratio#:~:text=Signal%2Dto%2Dnoise%20ratio%20(,power%2C%20often%20expressed%20in%20decibels\" target=\"_blank\">https://en.wikipedia.org/wiki/Signal-to-noise_ratio#:~:text=Signal%2Dto%2Dnoise%20ratio%20(,power%2C%20often%20expressed%20in%20decibels</a>.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 981890,
      "author_name": "Louka Ewington-Pitsos",
      "author_url": "",
      "post_date": "2020-08-22T19:50:37.737000",
      "content": "<p>Apologies if this is a silly question, but what normalization procedure do you use? From the images it looks like you're subtracting the mean of each waveform and dividing by the standard deviation but I could easily be wrong. I've come to fear the word \"normalize\" lately because usually when people use it they could plausibly be referring to ~eight different operations. </p>\n<p>The image below is created by:</p>\n<ol>\n<li>convert all waveforms to melspect</li>\n<li>add melspects elementwise</li>\n<li>divide melspect values by number of waveforms</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1577010%2Fafba0fba655e9d9dc6f64e80afc582d2%2FScreenshot%20from%202020-08-23%2005-55-34.png?generation=1598126362539466&amp;alt=media\" alt=\"\"></p>\n<p>The second image is </p>\n<ol>\n<li><code>standardize_relative</code> for all waveforms</li>\n<li>convert all waveforms to melspect</li>\n<li>add melspects elementwise</li>\n<li>divide melspect values by number of waveforms</li>\n</ol>\n<p>using </p>\n<pre><code>def standardize_relative(x):\n    return standardize(x, np.mean(x), np.std(x))\n\ndef standardize(x, mean, sd):\n    return (x - mean) / sd\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1577010%2F408505a480671190242680ffc077a675%2FScreenshot%20from%202020-08-23%2005-55-25.png?generation=1598126192473182&amp;alt=media\" alt=\"\"></p>\n<p>This final image is </p>\n<ol>\n<li><code>standardize_relative</code> for all waveforms</li>\n<li>add waveforms elementwise</li>\n<li>divide waveform values by number of waveforms </li>\n<li>convert final waveform to melspect</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1577010%2F63fe6cc81dd40975b44bab38f134841a%2FScreenshot%20from%202020-08-23%2006-03-56.png?generation=1598126678346710&amp;alt=media\" alt=\"\"></p>\n<p>The second one looks pretty decent to me, any suggestions for improvement?</p>\n<p>Below is some code to easily mix multiple waveforms using the second method if anyone is interested. All overlapping/mixing occurs at the start of the clip and if clips are different lengths the end of the mixed clip will be identical to the end of the longest clip.</p>\n<pre><code>import functools\n\ndef standardize_relative(x):\n    return standardize(x, np.mean(x), np.std(x))\n\ndef min_duration(signals):\n    return min([s.shape[1] for s in signals])\n\ndef standardize(x, mean, sd):\n    return (x - mean) / sd'\n\ndef mix(signals, sr, mel_fn):\n    mels =  [mel_fn(y=standardize_relative(i), sr=sr) for i in signals]\n    mixed = []\n    height = mels[0].shape[0]\n\n    while len(mels) &gt; 0:\n        shortest_len = min_duration(mels)\n\n        base = np.full((height, shortest_len), 0.0)\n\n        for s in mels:\n            base += s[:, :shortest_len] / len(mels)\n\n        mixed.append(base)\n\n        mels = [s[:, shortest_len:] for s in mels if s.shape[1] &gt; shortest_len]\n\n    return np.concatenate(mixed, axis=1)\n\n# a, b, c are all waveforms represented as numpy arrays\nfinal_mel = mix([a, b, c], 33075, functools.partial(librosa.feature.melspectrogram, n_fft=2048, hop_length=768, n_mels=424))\n</code></pre>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 952660,
      "author_name": "leilaSoleymani",
      "author_url": "",
      "post_date": "2020-07-31T05:53:19.713000",
      "content": "<p>how a good topic👍 </p>",
      "votes": 0,
      "replies": [
        {
          "id": 952848,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2020-07-31T09:11:13.243000",
          "content": "<p>Thanks <a href=\"/leila73\">@leila73</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 952412,
      "author_name": "Salman Ibne Eunus",
      "author_url": "",
      "post_date": "2020-07-30T22:27:25.003000",
      "content": "<p>Good work!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 952421,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-30T22:38:43.713000",
      "content": "",
      "votes": -4,
      "replies": [
        {
          "id": 952458,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2020-07-30T23:56:04.200000",
          "content": "<p>How do you combine mel sepcs ?</p>\n\n<p><strong>PS</strong>:  MelSpecs(x1+x2) != MelSepcs(x1) + MelSpecs(x2) !</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 952534,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-31T03:01:27.440000",
          "content": "",
          "votes": -18,
          "replies": []
        },
        {
          "id": 952847,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2020-07-31T09:10:41.007000",
          "content": "<p>You're a very funny guy, keep it on dear 😂 </p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 953219,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-31T15:59:21.207000",
          "content": "",
          "votes": -8,
          "replies": []
        },
        {
          "id": 980964,
          "author_name": "Louka Ewington-Pitsos",
          "author_url": "",
          "post_date": "2020-08-22T03:59:42.973000",
          "content": "<p>I think people are downvoting you because you sound a little bit arrogant, just a heads up.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "952012": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1498852%2F0559ecb0ee51031face0724230c51090%2FXC178525_PER54_20190116_un_normalized.PNG?generation=1596121738571626&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1498852%2F3f1112de680c8ad46839bec0d1c52145%2FXC178525_PER54_20190116_normalized.PNG?generation=1596122490291744&amp;alt=media)\n",
    "972051": "as an illustration:\n\ntop to bottom:\nmelspec(noise)\nmelspec(bird)\nmelspec(noise+bird)\nmelspec(noise)+melspec(bird)\n\nleft: orginal spectrum, right: after my denoise algorithm\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F924b03b72af0252f68c00859d640c100%2FSelection_035.png?generation=1597563514500825&alt=media)",
    "952562": "[BC learning](https://arxiv.org/abs/1711.10282) may help you.\nOn page 4, It proposes equation (2), which is mixed to account for sound pressure level(G1 and G2).\nAnd in this case, r=0.5.",
    "952551": "Quite intuitive explanation!\n\nI haven't tried this yet but maybe [Scaper](https://scaper.readthedocs.io/en/latest/installation.html) is useful for mixing audio clips (this is also shared by @hengck23 in this [post](https://www.kaggle.com/c/birdsong-recognition/discussion/169538))\n\nBy the way, have you already tried mixing audio clips? If you don't mind can you share us how it works? About myself, I haven't tried yet but I think it worth trying.",
    "952252": "I might not have the correct answer for your question. but, i have a doubt with this implementation. you are concatenation(stacking) various waves one after another. in this case you will not get any overlap between different species(even though there are several bg species). \nas per the competition description, we are supposed to work on data with various species calling simultaniously.    \nmaybe a better approach should be mixing then into same timeframe .",
    "981890": "Apologies if this is a silly question, but what normalization procedure do you use? From the images it looks like you're subtracting the mean of each waveform and dividing by the standard deviation but I could easily be wrong. I've come to fear the word \"normalize\" lately because usually when people use it they could plausibly be referring to ~eight different operations. \n\nThe image below is created by:\n1. convert all waveforms to melspect\n2. add melspects elementwise\n3. divide melspect values by number of waveforms\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1577010%2Fafba0fba655e9d9dc6f64e80afc582d2%2FScreenshot%20from%202020-08-23%2005-55-34.png?generation=1598126362539466&alt=media)\n\n\nThe second image is \n0. `standardize_relative` for all waveforms\n1. convert all waveforms to melspect\n2. add melspects elementwise\n3. divide melspect values by number of waveforms\n\nusing \n\n```\ndef standardize_relative(x):\n    return standardize(x, np.mean(x), np.std(x))\n\ndef standardize(x, mean, sd):\n    return (x - mean) / sd\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1577010%2F408505a480671190242680ffc077a675%2FScreenshot%20from%202020-08-23%2005-55-25.png?generation=1598126192473182&alt=media)\n\n\nThis final image is \n\n0. `standardize_relative` for all waveforms\n1.  add waveforms elementwise\n2. divide waveform values by number of waveforms \n3. convert final waveform to melspect\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1577010%2F63fe6cc81dd40975b44bab38f134841a%2FScreenshot%20from%202020-08-23%2006-03-56.png?generation=1598126678346710&alt=media)\n\n\nThe second one looks pretty decent to me, any suggestions for improvement?\n\nBelow is some code to easily mix multiple waveforms using the second method if anyone is interested. All overlapping/mixing occurs at the start of the clip and if clips are different lengths the end of the mixed clip will be identical to the end of the longest clip.\n\n```\nimport functools\n\ndef standardize_relative(x):\n    return standardize(x, np.mean(x), np.std(x))\n\ndef min_duration(signals):\n    return min([s.shape[1] for s in signals])\n\ndef standardize(x, mean, sd):\n    return (x - mean) / sd'\n\ndef mix(signals, sr, mel_fn):\n    mels =  [mel_fn(y=standardize_relative(i), sr=sr) for i in signals]\n    mixed = []\n    height = mels[0].shape[0]\n\n    while len(mels) > 0:\n        shortest_len = min_duration(mels)\n\n        base = np.full((height, shortest_len), 0.0)\n\n        for s in mels:\n            base += s[:, :shortest_len] / len(mels)\n\n        mixed.append(base)\n\n        mels = [s[:, shortest_len:] for s in mels if s.shape[1] > shortest_len]\n    \n    return np.concatenate(mixed, axis=1)\n\n# a, b, c are all waveforms represented as numpy arrays\nfinal_mel = mix([a, b, c], 33075, functools.partial(librosa.feature.melspectrogram, n_fft=2048, hop_length=768, n_mels=424))\n```\n\n",
    "952660": "how a good topic👍 ",
    "952412": "Good work!",
    "952421": ""
  }
}