{
  "id": 275346,
  "title": "did anyone get whitening to work?",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/275346",
  "author_name": "",
  "post_date": "2021-09-30T02:57:04.812949200Z",
  "votes": 6,
  "comment_count": 10,
  "views": 0,
  "content": "<p>i spend the last week on this. Did anyone really get this to work?<br>\nI cannot understand why it doesn't work …. this has been widely used in the GW communities.</p>",
  "messages": [
    {
      "id": "1528893",
      "postDate": "09/30/2021 02:57:04",
      "content": "<p>i spend the last week on this. Did anyone really get this to work?<br>\nI cannot understand why it doesn't work …. this has been widely used in the GW communities.</p>",
      "rawMarkdown": "i spend the last week on this. Did anyone really get this to work?\nI cannot understand why it doesn't work .... this has been widely used in the GW communities.",
      "votes": null
    },
    {
      "id": "1528903",
      "postDate": "09/30/2021 03:03:07",
      "content": "<p>It works well for me. Well, it doesnt if you do it wrong of course. It doesn't most probably because you are messing up your activation distributions from noisy low data (num of points, 4096) samples.</p>\n<p>I did very simple fft - PSD -ifft type of thing, but with PSD taken mean batchwise for each channel, [1, 3, 4096]. </p>",
      "rawMarkdown": "It works well for me. Well, it doesnt if you do it wrong of course. It doesn't most probably because you are messing up your activation distributions from noisy low data (num of points, 4096) samples.\n\nI did very simple fft - PSD -ifft type of thing, but with PSD taken mean batchwise for each channel, [1, 3, 4096].",
      "votes": null
    },
    {
      "id": "1528909",
      "postDate": "09/30/2021 03:10:01",
      "content": "<p>thanks for the comment. how much improvement does it bring? some teams reported poorer accuracy in the forum.</p>\n<p>\" PSD taken mean batchwise …\"<br>\nso it is batch-dependent. won't that affect your testing?</p>",
      "rawMarkdown": "thanks for the comment. how much improvement does it bring? some teams reported poorer accuracy in the forum.\n\n\" PSD taken mean batchwise ...\"\nso it is batch-dependent. won't that affect your testing?",
      "votes": null
    },
    {
      "id": "1528920",
      "postDate": "09/30/2021 03:21:20",
      "content": "<p>It should, and in final submissions I fixed that with precomputed value. I did it very early in competition, and results without it \\ or with \"unfixed\" PSD whitening were so bad i did not submit any of it. I definitely think it helps.</p>",
      "rawMarkdown": "It should, and in final submissions I fixed that with precomputed value. I did it very early in competition, and results without it \\ or with \"unfixed\" PSD whitening were so bad i did not submit any of it. I definitely think it helps.",
      "votes": null
    },
    {
      "id": "1530300",
      "postDate": "10/01/2021 04:58:55",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Similar whitening approach gave us ~0.003 boost early on (8650 range) and ~0.001+ boost later (8800 range)</p>",
      "rawMarkdown": "hengck23 Similar whitening approach gave us ~0.003 boost early on (8650 range) and ~0.001+ boost later (8800 range)",
      "votes": null
    },
    {
      "id": "1530341",
      "postDate": "10/01/2021 05:33:55",
      "content": "<p>The trick is using mean over the train set PSD for normalization… as found by my teammates</p>",
      "rawMarkdown": "The trick is using mean over the train set PSD for normalization... as found by my teammates",
      "votes": null
    },
    {
      "id": "1530354",
      "postDate": "10/01/2021 05:45:28",
      "content": "<p>In theory whitening should not work here. When you are doing whitening properly you are basically applying filter with length around few seconds, so entire signal should be messed up. Period. </p>\n<p>BUT.</p>\n<p>Here we had generated noise. PSD was same for entire dataset and PSD was smooth. In reality only few hundred samples where corrupted badly, so you can make it work \"good enough for practical application\".<br>\n<a href=\"https://www.kaggle.com/bakeryproducts\" target=\"_blank\">@bakeryproducts</a> -style - just do it and apply window to eliminate insane border effects, than normalize. Key here is to pre-calculate PSD on noise-only samples, than you are good to go. Don't write this in whitepaper tho.<br>\nAlso you can just find in dataset sequence of samples to smoothly continue signal and then do whitening. I personally just train linear AR on signal, predicted few samples and search for close match in dataset. Visually yielded much cleaner result, but did not put it in our pipeline, already was no use for us at a time.</p>",
      "rawMarkdown": "In theory whitening should not work here. When you are doing whitening properly you are basically applying filter with length around few seconds, so entire signal should be messed up. Period. \n\nBUT.\n\nHere we had generated noise. PSD was same for entire dataset and PSD was smooth. In reality only few hundred samples where corrupted badly, so you can make it work \"good enough for practical application\".\n@bakeryproducts -style - just do it and apply window to eliminate insane border effects, than normalize. Key here is to pre-calculate PSD on noise-only samples, than you are good to go. Don't write this in whitepaper tho.\nAlso you can just find in dataset sequence of samples to smoothly continue signal and then do whitening. I personally just train linear AR on signal, predicted few samples and search for close match in dataset. Visually yielded much cleaner result, but did not put it in our pipeline, already was no use for us at a time.",
      "votes": null
    },
    {
      "id": "1530380",
      "postDate": "10/01/2021 06:10:44",
      "content": "<p>thanks for the information.</p>\n<p>\" I personally just train linear AR on signal\"<br>\nI would like to try that since I read and self-study on the topic.<br>\nDo you have code on that?</p>",
      "rawMarkdown": "thanks for the information.\n\n\" I personally just train linear AR on signal\"\nI would like to try that since I read and self-study on the topic.\nDo you have code on that?",
      "votes": null
    },
    {
      "id": "1532870",
      "postDate": "10/03/2021 13:00:04",
      "content": "<p>It depends on how you define whitening.  </p>\n<p>Using the usual whitening methods did not help us.  Dividing by average noise helped.</p>",
      "rawMarkdown": "It depends on how you define whitening.  \n\nUsing the usual whitening methods did not help us.  Dividing by average noise helped.",
      "votes": null
    },
    {
      "id": "1532965",
      "postDate": "10/03/2021 14:38:21",
      "content": "<p>Median over the noise samples might be more robust to Glitchs. As I understand, we want to keep the glitches in the whitened signal and then the model should be able to differentiate between the Glitchs vs the GW signal. There is a median correction that needs to be applied (using _median_bias from scipy/signal/spectral.py) and then one could smooth along the frequency access but a special treatment is needed around the few peaks of the PSD to avoid smearing those -- see <a href=\"https://arxiv.org/pdf/2108.10588.pdf\" target=\"_blank\">https://arxiv.org/pdf/2108.10588.pdf</a></p>",
      "rawMarkdown": "Median over the noise samples might be more robust to Glitchs. As I understand, we want to keep the glitches in the whitened signal and then the model should be able to differentiate between the Glitchs vs the GW signal. There is a median correction that needs to be applied (using _median_bias from scipy/signal/spectral.py) and then one could smooth along the frequency access but a special treatment is needed around the few peaks of the PSD to avoid smearing those -- see https://arxiv.org/pdf/2108.10588.pdf",
      "votes": null
    },
    {
      "id": "1559906",
      "postDate": "10/27/2021 08:07:47",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1528903,
      "author_name": "bakeryproducts",
      "author_url": "",
      "post_date": "09/30/2021 03:03:07",
      "content": "<p>It works well for me. Well, it doesnt if you do it wrong of course. It doesn't most probably because you are messing up your activation distributions from noisy low data (num of points, 4096) samples.</p>\n<p>I did very simple fft - PSD -ifft type of thing, but with PSD taken mean batchwise for each channel, [1, 3, 4096]. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1528909,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/30/2021 03:10:01",
          "content": "<p>thanks for the comment. how much improvement does it bring? some teams reported poorer accuracy in the forum.</p>\n<p>\" PSD taken mean batchwise …\"<br>\nso it is batch-dependent. won't that affect your testing?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1528920,
          "author_name": "bakeryproducts",
          "author_url": "",
          "post_date": "09/30/2021 03:21:20",
          "content": "<p>It should, and in final submissions I fixed that with precomputed value. I did it very early in competition, and results without it \\ or with \"unfixed\" PSD whitening were so bad i did not submit any of it. I definitely think it helps.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1530300,
          "author_name": "richx86",
          "author_url": "",
          "post_date": "10/01/2021 04:58:55",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Similar whitening approach gave us ~0.003 boost early on (8650 range) and ~0.001+ boost later (8800 range)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1530341,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/01/2021 05:33:55",
          "content": "<p>The trick is using mean over the train set PSD for normalization… as found by my teammates</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1532965,
          "author_name": "ggrizzly",
          "author_url": "",
          "post_date": "10/03/2021 14:38:21",
          "content": "<p>Median over the noise samples might be more robust to Glitchs. As I understand, we want to keep the glitches in the whitened signal and then the model should be able to differentiate between the Glitchs vs the GW signal. There is a median correction that needs to be applied (using _median_bias from scipy/signal/spectral.py) and then one could smooth along the frequency access but a special treatment is needed around the few peaks of the PSD to avoid smearing those -- see <a href=\"https://arxiv.org/pdf/2108.10588.pdf\" target=\"_blank\">https://arxiv.org/pdf/2108.10588.pdf</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1530354,
      "author_name": "denisbsu",
      "author_url": "",
      "post_date": "10/01/2021 05:45:28",
      "content": "<p>In theory whitening should not work here. When you are doing whitening properly you are basically applying filter with length around few seconds, so entire signal should be messed up. Period. </p>\n<p>BUT.</p>\n<p>Here we had generated noise. PSD was same for entire dataset and PSD was smooth. In reality only few hundred samples where corrupted badly, so you can make it work \"good enough for practical application\".<br>\n<a href=\"https://www.kaggle.com/bakeryproducts\" target=\"_blank\">@bakeryproducts</a> -style - just do it and apply window to eliminate insane border effects, than normalize. Key here is to pre-calculate PSD on noise-only samples, than you are good to go. Don't write this in whitepaper tho.<br>\nAlso you can just find in dataset sequence of samples to smoothly continue signal and then do whitening. I personally just train linear AR on signal, predicted few samples and search for close match in dataset. Visually yielded much cleaner result, but did not put it in our pipeline, already was no use for us at a time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1530380,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/01/2021 06:10:44",
          "content": "<p>thanks for the information.</p>\n<p>\" I personally just train linear AR on signal\"<br>\nI would like to try that since I read and self-study on the topic.<br>\nDo you have code on that?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1532870,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "10/03/2021 13:00:04",
      "content": "<p>It depends on how you define whitening.  </p>\n<p>Using the usual whitening methods did not help us.  Dividing by average noise helped.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1559906,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 08:07:47",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1528893": "i spend the last week on this. Did anyone really get this to work?\nI cannot understand why it doesn't work .... this has been widely used in the GW communities.",
    "1528903": "It works well for me. Well, it doesnt if you do it wrong of course. It doesn't most probably because you are messing up your activation distributions from noisy low data (num of points, 4096) samples.\n\nI did very simple fft - PSD -ifft type of thing, but with PSD taken mean batchwise for each channel, [1, 3, 4096].",
    "1528909": "thanks for the comment. how much improvement does it bring? some teams reported poorer accuracy in the forum.\n\n\" PSD taken mean batchwise ...\"\nso it is batch-dependent. won't that affect your testing?",
    "1528920": "It should, and in final submissions I fixed that with precomputed value. I did it very early in competition, and results without it \\ or with \"unfixed\" PSD whitening were so bad i did not submit any of it. I definitely think it helps.",
    "1530300": "hengck23 Similar whitening approach gave us ~0.003 boost early on (8650 range) and ~0.001+ boost later (8800 range)",
    "1530341": "The trick is using mean over the train set PSD for normalization... as found by my teammates",
    "1530354": "In theory whitening should not work here. When you are doing whitening properly you are basically applying filter with length around few seconds, so entire signal should be messed up. Period. \n\nBUT.\n\nHere we had generated noise. PSD was same for entire dataset and PSD was smooth. In reality only few hundred samples where corrupted badly, so you can make it work \"good enough for practical application\".\n@bakeryproducts -style - just do it and apply window to eliminate insane border effects, than normalize. Key here is to pre-calculate PSD on noise-only samples, than you are good to go. Don't write this in whitepaper tho.\nAlso you can just find in dataset sequence of samples to smoothly continue signal and then do whitening. I personally just train linear AR on signal, predicted few samples and search for close match in dataset. Visually yielded much cleaner result, but did not put it in our pipeline, already was no use for us at a time.",
    "1530380": "thanks for the information.\n\n\" I personally just train linear AR on signal\"\nI would like to try that since I read and self-study on the topic.\nDo you have code on that?",
    "1532870": "It depends on how you define whitening.  \n\nUsing the usual whitening methods did not help us.  Dividing by average noise helped.",
    "1532965": "Median over the noise samples might be more robust to Glitchs. As I understand, we want to keep the glitches in the whitened signal and then the model should be able to differentiate between the Glitchs vs the GW signal. There is a median correction that needs to be applied (using _median_bias from scipy/signal/spectral.py) and then one could smooth along the frequency access but a special treatment is needed around the few peaks of the PSD to avoid smearing those -- see https://arxiv.org/pdf/2108.10588.pdf",
    "1559906": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}