{
  "id": 267778,
  "title": "Call the cops - too much noise",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/267778",
  "author_name": "عثمان",
  "post_date": "2021-08-24T16:38:35.097000",
  "votes": 30,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I've spent the better part of the past 24h working through various repos online attempting to understanding the whitening process. At the end of the day, it basically boils down to a two step process:</p>\n<ol>\n<li>Applying some form of windowing (hann, tukey, etc) to the waves from each detector, generating their respective PSDs, and stashing away an interpolation from freq-&gt;psd. </li>\n<li>Taking the raw waves and windowing them again. Some on the forums use a different windowing operation here while others seem to just reuse whatever they did in step 1 above. In the literature and various published studies' codes, researchers tend to use the same windowing method. In either case, dividing the \"real\" frequency-domain signal by the sqrt of the PSD-appropriately interpolated and normed, and then converting it back to time-domain.</li>\n</ol>\n<p>This whole process is usually followed by bandpassing the expected ~30-350 or so Hz range for our GW events.</p>\n<p>Whitening as a preprocessing step hasn't helped any of my model's RoC-AUC. I did notice the loss drops faster in the initial epoch; but also plateaus faster as well, causing it to fail to reach desired minima. Perhaps these models might blend well, who knows. But insofar their individual performance leaves something to be desired.</p>\n<p>To those leading the pack (top50) - a number of you commented 2-3 weeks back that whitening hasn't benefited you. If I might inquire, is that still the case today? I ask because at least for me, sometimes I make a post based on my current understanding… but then things change as I experiment more, and I don't necessarily always go back and update every post unless someone comments or asks a question specifically. I think this is common for most people. For example, earlier on in the competition, some people mistakenly assumed we were using real detector signals (noise) with simulated GW events. But now we know that is not the case, even those those EDAs remain today.</p>\n<p>Another thing I've noticed in the literature is that the first step of PSD calculation seems to be done on an excess amount of data (as many seconds of event-free strain data as the researchers can get their hands on) in order to build a \"more accurate PSD\", which is then used with smaller samples in the second step of the whitening process. Eyeballing some of these PSD diagrams, is that necessary? I see a huge peak within the first ~50 or so values. It's not even a peak, it's a smooth curve. It's just that after that, the rest of the ~2048 values is essentially === 0.</p>\n<p><img src=\"https://i.imgur.com/2bYYPIC.png\" alt=\"PSD\"></p>\n<p>In my modeling, I've been calculating PSDs on a per-instance basis. Including those instances with GW events. Is it possible this might be a cause for decreased performance? Also, is it a valid operation to, e.g. take the PSDs per-detector for all non-GW events and average them-again per detector? Or is that a flawed calculation (akin to taking the mean standard deviation per image instance over a dataset in order to calculate the std of the entire dataset)?</p>",
  "messages": [
    {
      "id": 1489007,
      "postDate": "2021-08-24T16:38:35.097Z",
      "content": "<p>I've spent the better part of the past 24h working through various repos online attempting to understanding the whitening process. At the end of the day, it basically boils down to a two step process:</p>\n<ol>\n<li>Applying some form of windowing (hann, tukey, etc) to the waves from each detector, generating their respective PSDs, and stashing away an interpolation from freq-&gt;psd. </li>\n<li>Taking the raw waves and windowing them again. Some on the forums use a different windowing operation here while others seem to just reuse whatever they did in step 1 above. In the literature and various published studies' codes, researchers tend to use the same windowing method. In either case, dividing the \"real\" frequency-domain signal by the sqrt of the PSD-appropriately interpolated and normed, and then converting it back to time-domain.</li>\n</ol>\n<p>This whole process is usually followed by bandpassing the expected ~30-350 or so Hz range for our GW events.</p>\n<p>Whitening as a preprocessing step hasn't helped any of my model's RoC-AUC. I did notice the loss drops faster in the initial epoch; but also plateaus faster as well, causing it to fail to reach desired minima. Perhaps these models might blend well, who knows. But insofar their individual performance leaves something to be desired.</p>\n<p>To those leading the pack (top50) - a number of you commented 2-3 weeks back that whitening hasn't benefited you. If I might inquire, is that still the case today? I ask because at least for me, sometimes I make a post based on my current understanding… but then things change as I experiment more, and I don't necessarily always go back and update every post unless someone comments or asks a question specifically. I think this is common for most people. For example, earlier on in the competition, some people mistakenly assumed we were using real detector signals (noise) with simulated GW events. But now we know that is not the case, even those those EDAs remain today.</p>\n<p>Another thing I've noticed in the literature is that the first step of PSD calculation seems to be done on an excess amount of data (as many seconds of event-free strain data as the researchers can get their hands on) in order to build a \"more accurate PSD\", which is then used with smaller samples in the second step of the whitening process. Eyeballing some of these PSD diagrams, is that necessary? I see a huge peak within the first ~50 or so values. It's not even a peak, it's a smooth curve. It's just that after that, the rest of the ~2048 values is essentially === 0.</p>\n<p><img src=\"https://i.imgur.com/2bYYPIC.png\" alt=\"PSD\"></p>\n<p>In my modeling, I've been calculating PSDs on a per-instance basis. Including those instances with GW events. Is it possible this might be a cause for decreased performance? Also, is it a valid operation to, e.g. take the PSDs per-detector for all non-GW events and average them-again per detector? Or is that a flawed calculation (akin to taking the mean standard deviation per image instance over a dataset in order to calculate the std of the entire dataset)?</p>",
      "rawMarkdown": "I've spent the better part of the past 24h working through various repos online attempting to understanding the whitening process. At the end of the day, it basically boils down to a two step process:\n\n1. Applying some form of windowing (hann, tukey, etc) to the waves from each detector, generating their respective PSDs, and stashing away an interpolation from freq->psd. \n2. Taking the raw waves and windowing them again. Some on the forums use a different windowing operation here while others seem to just reuse whatever they did in step 1 above. In the literature and various published studies' codes, researchers tend to use the same windowing method. In either case, dividing the \"real\" frequency-domain signal by the sqrt of the PSD-appropriately interpolated and normed, and then converting it back to time-domain.\n\nThis whole process is usually followed by bandpassing the expected ~30-350 or so Hz range for our GW events.\n\nWhitening as a preprocessing step hasn't helped any of my model's RoC-AUC. I did notice the loss drops faster in the initial epoch; but also plateaus faster as well, causing it to fail to reach desired minima. Perhaps these models might blend well, who knows. But insofar their individual performance leaves something to be desired.\n\nTo those leading the pack (top50) - a number of you commented 2-3 weeks back that whitening hasn't benefited you. If I might inquire, is that still the case today? I ask because at least for me, sometimes I make a post based on my current understanding... but then things change as I experiment more, and I don't necessarily always go back and update every post unless someone comments or asks a question specifically. I think this is common for most people. For example, earlier on in the competition, some people mistakenly assumed we were using real detector signals (noise) with simulated GW events. But now we know that is not the case, even those those EDAs remain today.\n\nAnother thing I've noticed in the literature is that the first step of PSD calculation seems to be done on an excess amount of data (as many seconds of event-free strain data as the researchers can get their hands on) in order to build a \"more accurate PSD\", which is then used with smaller samples in the second step of the whitening process. Eyeballing some of these PSD diagrams, is that necessary? I see a huge peak within the first ~50 or so values. It's not even a peak, it's a smooth curve. It's just that after that, the rest of the ~2048 values is essentially === 0.\n\n![PSD](https://i.imgur.com/2bYYPIC.png)\n\nIn my modeling, I've been calculating PSDs on a per-instance basis. Including those instances with GW events. Is it possible this might be a cause for decreased performance? Also, is it a valid operation to, e.g. take the PSDs per-detector for all non-GW events and average them-again per detector? Or is that a flawed calculation (akin to taking the mean standard deviation per image instance over a dataset in order to calculate the std of the entire dataset)?",
      "votes": 30
    },
    {
      "id": 1489030,
      "postDate": "2021-08-24T16:59:05.410Z",
      "content": "<p>The other thing to consider is the sample rate of the data which is lower than what most of the GW publications use (I think the lowest I've seen is 8192Hz in this <a href=\"https://journals.aps.org/prl/pdf/10.1103/PhysRevLett.120.141103\" target=\"_blank\">paper</a>. <strong>Edit</strong>: 4096Hz from this <a href=\"https://www.gw-openscience.org/LVT151012data/LOSC_Event_tutorial_LVT151012.html\" target=\"_blank\">demo</a>). I did some tests with simulated data <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/252396#1483747\" target=\"_blank\">here</a> </p>",
      "rawMarkdown": "The other thing to consider is the sample rate of the data which is lower than what most of the GW publications use (I think the lowest I've seen is 8192Hz in this [paper](https://journals.aps.org/prl/pdf/10.1103/PhysRevLett.120.141103). **Edit**: 4096Hz from this [demo](https://www.gw-openscience.org/LVT151012data/LOSC_Event_tutorial_LVT151012.html)). I did some tests with simulated data [here](https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/252396#1483747) ",
      "votes": 6,
      "replies": [
        {
          "id": 1490895,
          "postDate": "2021-08-26T01:25:07.850Z",
          "content": "<p>Even in that demo the PSD and ASD are made using +30sec excess of data around the event.</p>",
          "rawMarkdown": "Even in that demo the PSD and ASD are made using +30sec excess of data around the event.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1496986,
      "postDate": "2021-08-30T19:40:15.073Z",
      "content": "<p>there is one thing you can try.</p>\n<p>play the samples as sound waves and see if you can hear the difference between the noisy positive and negative train samples.</p>\n<p>if we can see the difference in CQT, chances are that image CNN would work.</p>\n<p>if we can hear the difference, maybe just 1d cnn like wavenet would work</p>",
      "rawMarkdown": "there is one thing you can try.\n\nplay the samples as sound waves and see if you can hear the difference between the noisy positive and negative train samples.\n\nif we can see the difference in CQT, chances are that image CNN would work.\n\nif we can hear the difference, maybe just 1d cnn like wavenet would work\n",
      "votes": 3
    },
    {
      "id": 1493335,
      "postDate": "2021-08-27T19:23:02.980Z",
      "content": "<p>the maths and equation are here:</p>\n<p><a href=\"https://matheo.uliege.be/bitstream/2268.2/9211/4/Master_thesis_Janquart.pdf\" target=\"_blank\">https://matheo.uliege.be/bitstream/2268.2/9211/4/Master_thesis_Janquart.pdf</a></p>",
      "rawMarkdown": "the maths and equation are here:\n\nhttps://matheo.uliege.be/bitstream/2268.2/9211/4/Master_thesis_Janquart.pdf",
      "votes": 4,
      "replies": [
        {
          "id": 1493350,
          "postDate": "2021-08-27T19:38:54.787Z",
          "content": "<p>There goes my weekend.</p>",
          "rawMarkdown": "There goes my weekend.",
          "votes": 3
        },
        {
          "id": 1493373,
          "postDate": "2021-08-27T20:20:40.603Z",
          "content": "<p>Now, we need a summary of this master thesis. Thanks for the link <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>. 👌 </p>",
          "rawMarkdown": "Now, we need a summary of this master thesis. Thanks for the link @hengck23. 👌 "
        },
        {
          "id": 1493896,
          "postDate": "2021-08-28T08:30:28.420Z",
          "content": "<p>Notice that there are collection of code repos that come with the master thesis: <a href=\"https://github.com/lemnis12?tab=repositories\" target=\"_blank\">https://github.com/lemnis12?tab=repositories</a>. Might be useful.</p>",
          "rawMarkdown": "Notice that there are collection of code repos that come with the master thesis: https://github.com/lemnis12?tab=repositories. Might be useful.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1496982,
      "postDate": "2021-08-30T19:31:57.283Z",
      "content": "<p>Firstly, the PSD should be viewed in a logarithmic scale. Furthermore, there are unofficial PSD's for the three detectors <a href=\"https://github.com/wvanzeist/riroriro_tutorials/tree/main/noise_spectra\" target=\"_blank\">available here</a>. Which originates from <a href=\"https://dcc.ligo.org/LIGO-T1500293/public\" target=\"_blank\">here</a>.</p>\n<p>Here is the Virgo PSD for example:<br>\n<a href=\"https://postimg.cc/D8fKQbXW\" target=\"_blank\"><img src=\"https://i.postimg.cc/Nfmfc6Gx/virgo-psd.png\" alt=\"virgo-psd.png\"></a></p>",
      "rawMarkdown": "Firstly, the PSD should be viewed in a logarithmic scale. Furthermore, there are unofficial PSD's for the three detectors [available here](https://github.com/wvanzeist/riroriro_tutorials/tree/main/noise_spectra). Which originates from [here](https://dcc.ligo.org/LIGO-T1500293/public).\n\nHere is the Virgo PSD for example:\n[![virgo-psd.png](https://i.postimg.cc/Nfmfc6Gx/virgo-psd.png)](https://postimg.cc/D8fKQbXW)",
      "votes": 1,
      "replies": [
        {
          "id": 1497020,
          "postDate": "2021-08-30T20:41:37.407Z",
          "content": "<p>Thanks for the direction <a href=\"https://www.kaggle.com/mistag\" target=\"_blank\">@mistag</a></p>\n<p>Is it legitimate (I'm asking the question as I queue up the experiment) to use a pre-baked PSD to denoise a segment? I ask because it feel like me taking an hour long recording from my house and then attempting to use that do denoise audio. Even if the recording was 1 month long, at the end of the day, the noise profile is still a very 'compressed' representation and likely won't do a very good job at denoising any specific samples (me cooking, me cleaning, my typing, my raging about LB freefall…).</p>\n<p>I feel like the PSD should be built about / around the sample we're trying to denoise, otherwise its utility is only great for things like getting rid of power grid noises and their harmonics or other constant / very periodic sources?</p>\n<p>The other thing, I'm not sure how many of the periodic sources they actually injected into the simulated audio. It'd be great if they generated detector noise simply by \"colouring\" white noise with these PSD's though…………………………………… &nbsp;&nbsp;🤣</p>",
          "rawMarkdown": "Thanks for the direction @mistag\n\nIs it legitimate (I'm asking the question as I queue up the experiment) to use a pre-baked PSD to denoise a segment? I ask because it feel like me taking an hour long recording from my house and then attempting to use that do denoise audio. Even if the recording was 1 month long, at the end of the day, the noise profile is still a very 'compressed' representation and likely won't do a very good job at denoising any specific samples (me cooking, me cleaning, my typing, my raging about LB freefall...).\n\nI feel like the PSD should be built about / around the sample we're trying to denoise, otherwise its utility is only great for things like getting rid of power grid noises and their harmonics or other constant / very periodic sources?\n\nThe other thing, I'm not sure how many of the periodic sources they actually injected into the simulated audio. It'd be great if they generated detector noise simply by \"colouring\" white noise with these PSD's though..........................................   🤣"
        },
        {
          "id": 1497047,
          "postDate": "2021-08-30T21:42:32.530Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1498473,
          "postDate": "2021-09-01T02:25:48.643Z",
          "content": "<p><a href=\"https://www.kaggle.com/mistag\" target=\"_blank\">@mistag</a> have you compared the PSD for the real data vs the competition. I ask because your image above is  quite different from sample. <a href=\"https://drive.google.com/file/d/15VwjrAY2bG770zuixVIsAKPciU3UT_m6/view?usp=sharing\" target=\"_blank\">'00001f4945'</a><br>\nwhich is target=0, i.e., the large noise lines are missing</p>",
          "rawMarkdown": "@mistag have you compared the PSD for the real data vs the competition. I ask because your image above is  quite different from sample. ['00001f4945'](https://drive.google.com/file/d/15VwjrAY2bG770zuixVIsAKPciU3UT_m6/view?usp=sharing)\nwhich is target=0, i.e., the large noise lines are missing\n"
        },
        {
          "id": 1498611,
          "postDate": "2021-09-01T05:12:38.410Z",
          "content": "<p>Not, yet - but planning to. The PSD represents the variance of energy across frequencies for a larger data set. The PSD for a single file will look different.</p>",
          "rawMarkdown": "Not, yet - but planning to. The PSD represents the variance of energy across frequencies for a larger data set. The PSD for a single file will look different."
        }
      ]
    },
    {
      "id": 1498444,
      "postDate": "2021-09-01T01:52:38.597Z",
      "content": "<p>I too started with this and did not get it to work  and teamed up with some other who also tried and failed.</p>\n<p>I'm not sure, but I understand that the data is all synthetic, i.e. the GW event is calculated and added to synthetic noise. What's not clear to me is how good a job the did in reproducing the noise.</p>",
      "rawMarkdown": "I too started with this and did not get it to work  and teamed up with some other who also tried and failed.\n\nI'm not sure, but I understand that the data is all synthetic, i.e. the GW event is calculated and added to synthetic noise. What's not clear to me is how good a job the did in reproducing the noise.",
      "votes": 2,
      "replies": [
        {
          "id": 1498620,
          "postDate": "2021-09-01T05:18:02.480Z",
          "content": "<p>We do not know how the host generated the noise. It is easy to generate synthetic noise signals from a PSD by multiplying the PSD with a random numbers with variance 1 and then add random phase before doing inverse FFT. But the host maby did something completely different…</p>",
          "rawMarkdown": "We do not know how the host generated the noise. It is easy to generate synthetic noise signals from a PSD by multiplying the PSD with a random numbers with variance 1 and then add random phase before doing inverse FFT. But the host maby did something completely different..."
        }
      ]
    },
    {
      "id": 1496422,
      "postDate": "2021-08-30T11:15:10.163Z",
      "content": "<p>Thanks for this very through analysis =))</p>",
      "rawMarkdown": "Thanks for this very through analysis =))",
      "votes": 2
    },
    {
      "id": 1496987,
      "postDate": "2021-08-30T19:46:10.563Z",
      "content": "<p>Please don't shot me <br>\nfor asking if you're OK =))</p>",
      "rawMarkdown": "Please don't shot me \nfor asking if you're OK =))",
      "votes": 1
    },
    {
      "id": 1559949,
      "postDate": "2021-10-27T08:36:38.787Z",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
    },
    {
      "id": 1493277,
      "postDate": "2021-08-27T18:30:08.163Z",
      "content": "<p>Thanks for this very through analysis. I can't answer any of your questions for now but will get back if I learn new things. </p>\n<p>For those that are new to the field, <strong>PSD</strong> stands for power <a href=\"https://en.wikipedia.org/wiki/Spectral_density\" target=\"_blank\">spectral density</a>. </p>\n<p>Also, from now on, I will use imgur to host my images since google drive no longer works. 👌</p>",
      "rawMarkdown": "Thanks for this very through analysis. I can't answer any of your questions for now but will get back if I learn new things. \n\nFor those that are new to the field, **PSD** stands for power [spectral density](https://en.wikipedia.org/wiki/Spectral_density). \n\nAlso, from now on, I will use imgur to host my images since google drive no longer works. 👌"
    }
  ],
  "comments": [
    {
      "id": 1489030,
      "author_name": "datasaurus",
      "author_url": "",
      "post_date": "2021-08-24T16:59:05.410000",
      "content": "<p>The other thing to consider is the sample rate of the data which is lower than what most of the GW publications use (I think the lowest I've seen is 8192Hz in this <a href=\"https://journals.aps.org/prl/pdf/10.1103/PhysRevLett.120.141103\" target=\"_blank\">paper</a>. <strong>Edit</strong>: 4096Hz from this <a href=\"https://www.gw-openscience.org/LVT151012data/LOSC_Event_tutorial_LVT151012.html\" target=\"_blank\">demo</a>). I did some tests with simulated data <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/252396#1483747\" target=\"_blank\">here</a> </p>",
      "votes": 6,
      "replies": [
        {
          "id": 1490895,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2021-08-26T01:25:07.850000",
          "content": "<p>Even in that demo the PSD and ASD are made using +30sec excess of data around the event.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1496986,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-08-30T19:40:15.073000",
      "content": "<p>there is one thing you can try.</p>\n<p>play the samples as sound waves and see if you can hear the difference between the noisy positive and negative train samples.</p>\n<p>if we can see the difference in CQT, chances are that image CNN would work.</p>\n<p>if we can hear the difference, maybe just 1d cnn like wavenet would work</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1493335,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-08-27T19:23:02.980000",
      "content": "<p>the maths and equation are here:</p>\n<p><a href=\"https://matheo.uliege.be/bitstream/2268.2/9211/4/Master_thesis_Janquart.pdf\" target=\"_blank\">https://matheo.uliege.be/bitstream/2268.2/9211/4/Master_thesis_Janquart.pdf</a></p>",
      "votes": 4,
      "replies": [
        {
          "id": 1493350,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2021-08-27T19:38:54.787000",
          "content": "<p>There goes my weekend.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1493373,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2021-08-27T20:20:40.603000",
          "content": "<p>Now, we need a summary of this master thesis. Thanks for the link <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>. 👌 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1493896,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2021-08-28T08:30:28.420000",
          "content": "<p>Notice that there are collection of code repos that come with the master thesis: <a href=\"https://github.com/lemnis12?tab=repositories\" target=\"_blank\">https://github.com/lemnis12?tab=repositories</a>. Might be useful.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1496982,
      "author_name": "Geir Drange",
      "author_url": "",
      "post_date": "2021-08-30T19:31:57.283000",
      "content": "<p>Firstly, the PSD should be viewed in a logarithmic scale. Furthermore, there are unofficial PSD's for the three detectors <a href=\"https://github.com/wvanzeist/riroriro_tutorials/tree/main/noise_spectra\" target=\"_blank\">available here</a>. Which originates from <a href=\"https://dcc.ligo.org/LIGO-T1500293/public\" target=\"_blank\">here</a>.</p>\n<p>Here is the Virgo PSD for example:<br>\n<a href=\"https://postimg.cc/D8fKQbXW\" target=\"_blank\"><img src=\"https://i.postimg.cc/Nfmfc6Gx/virgo-psd.png\" alt=\"virgo-psd.png\"></a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1497020,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2021-08-30T20:41:37.407000",
          "content": "<p>Thanks for the direction <a href=\"https://www.kaggle.com/mistag\" target=\"_blank\">@mistag</a></p>\n<p>Is it legitimate (I'm asking the question as I queue up the experiment) to use a pre-baked PSD to denoise a segment? I ask because it feel like me taking an hour long recording from my house and then attempting to use that do denoise audio. Even if the recording was 1 month long, at the end of the day, the noise profile is still a very 'compressed' representation and likely won't do a very good job at denoising any specific samples (me cooking, me cleaning, my typing, my raging about LB freefall…).</p>\n<p>I feel like the PSD should be built about / around the sample we're trying to denoise, otherwise its utility is only great for things like getting rid of power grid noises and their harmonics or other constant / very periodic sources?</p>\n<p>The other thing, I'm not sure how many of the periodic sources they actually injected into the simulated audio. It'd be great if they generated detector noise simply by \"colouring\" white noise with these PSD's though…………………………………… &nbsp;&nbsp;🤣</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1497047,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-08-30T21:42:32.530000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1498473,
          "author_name": "Fractal Feelings",
          "author_url": "",
          "post_date": "2021-09-01T02:25:48.643000",
          "content": "<p><a href=\"https://www.kaggle.com/mistag\" target=\"_blank\">@mistag</a> have you compared the PSD for the real data vs the competition. I ask because your image above is  quite different from sample. <a href=\"https://drive.google.com/file/d/15VwjrAY2bG770zuixVIsAKPciU3UT_m6/view?usp=sharing\" target=\"_blank\">'00001f4945'</a><br>\nwhich is target=0, i.e., the large noise lines are missing</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1498611,
          "author_name": "Geir Drange",
          "author_url": "",
          "post_date": "2021-09-01T05:12:38.410000",
          "content": "<p>Not, yet - but planning to. The PSD represents the variance of energy across frequencies for a larger data set. The PSD for a single file will look different.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1498444,
      "author_name": "Fractal Feelings",
      "author_url": "",
      "post_date": "2021-09-01T01:52:38.597000",
      "content": "<p>I too started with this and did not get it to work  and teamed up with some other who also tried and failed.</p>\n<p>I'm not sure, but I understand that the data is all synthetic, i.e. the GW event is calculated and added to synthetic noise. What's not clear to me is how good a job the did in reproducing the noise.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1498620,
          "author_name": "Geir Drange",
          "author_url": "",
          "post_date": "2021-09-01T05:18:02.480000",
          "content": "<p>We do not know how the host generated the noise. It is easy to generate synthetic noise signals from a PSD by multiplying the PSD with a random numbers with variance 1 and then add random phase before doing inverse FFT. But the host maby did something completely different…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1496422,
      "author_name": "fireflies",
      "author_url": "",
      "post_date": "2021-08-30T11:15:10.163000",
      "content": "<p>Thanks for this very through analysis =))</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1496987,
      "author_name": "fireflies",
      "author_url": "",
      "post_date": "2021-08-30T19:46:10.563000",
      "content": "<p>Please don't shot me <br>\nfor asking if you're OK =))</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1559949,
      "author_name": "ChristopherZerafa",
      "author_url": "",
      "post_date": "2021-10-27T08:36:38.787000",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1493277,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2021-08-27T18:30:08.163000",
      "content": "<p>Thanks for this very through analysis. I can't answer any of your questions for now but will get back if I learn new things. </p>\n<p>For those that are new to the field, <strong>PSD</strong> stands for power <a href=\"https://en.wikipedia.org/wiki/Spectral_density\" target=\"_blank\">spectral density</a>. </p>\n<p>Also, from now on, I will use imgur to host my images since google drive no longer works. 👌</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1489007": "I've spent the better part of the past 24h working through various repos online attempting to understanding the whitening process. At the end of the day, it basically boils down to a two step process:\n\n1. Applying some form of windowing (hann, tukey, etc) to the waves from each detector, generating their respective PSDs, and stashing away an interpolation from freq->psd. \n2. Taking the raw waves and windowing them again. Some on the forums use a different windowing operation here while others seem to just reuse whatever they did in step 1 above. In the literature and various published studies' codes, researchers tend to use the same windowing method. In either case, dividing the \"real\" frequency-domain signal by the sqrt of the PSD-appropriately interpolated and normed, and then converting it back to time-domain.\n\nThis whole process is usually followed by bandpassing the expected ~30-350 or so Hz range for our GW events.\n\nWhitening as a preprocessing step hasn't helped any of my model's RoC-AUC. I did notice the loss drops faster in the initial epoch; but also plateaus faster as well, causing it to fail to reach desired minima. Perhaps these models might blend well, who knows. But insofar their individual performance leaves something to be desired.\n\nTo those leading the pack (top50) - a number of you commented 2-3 weeks back that whitening hasn't benefited you. If I might inquire, is that still the case today? I ask because at least for me, sometimes I make a post based on my current understanding... but then things change as I experiment more, and I don't necessarily always go back and update every post unless someone comments or asks a question specifically. I think this is common for most people. For example, earlier on in the competition, some people mistakenly assumed we were using real detector signals (noise) with simulated GW events. But now we know that is not the case, even those those EDAs remain today.\n\nAnother thing I've noticed in the literature is that the first step of PSD calculation seems to be done on an excess amount of data (as many seconds of event-free strain data as the researchers can get their hands on) in order to build a \"more accurate PSD\", which is then used with smaller samples in the second step of the whitening process. Eyeballing some of these PSD diagrams, is that necessary? I see a huge peak within the first ~50 or so values. It's not even a peak, it's a smooth curve. It's just that after that, the rest of the ~2048 values is essentially === 0.\n\n![PSD](https://i.imgur.com/2bYYPIC.png)\n\nIn my modeling, I've been calculating PSDs on a per-instance basis. Including those instances with GW events. Is it possible this might be a cause for decreased performance? Also, is it a valid operation to, e.g. take the PSDs per-detector for all non-GW events and average them-again per detector? Or is that a flawed calculation (akin to taking the mean standard deviation per image instance over a dataset in order to calculate the std of the entire dataset)?",
    "1489030": "The other thing to consider is the sample rate of the data which is lower than what most of the GW publications use (I think the lowest I've seen is 8192Hz in this [paper](https://journals.aps.org/prl/pdf/10.1103/PhysRevLett.120.141103). **Edit**: 4096Hz from this [demo](https://www.gw-openscience.org/LVT151012data/LOSC_Event_tutorial_LVT151012.html)). I did some tests with simulated data [here](https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/252396#1483747) ",
    "1496986": "there is one thing you can try.\n\nplay the samples as sound waves and see if you can hear the difference between the noisy positive and negative train samples.\n\nif we can see the difference in CQT, chances are that image CNN would work.\n\nif we can hear the difference, maybe just 1d cnn like wavenet would work\n",
    "1493335": "the maths and equation are here:\n\nhttps://matheo.uliege.be/bitstream/2268.2/9211/4/Master_thesis_Janquart.pdf",
    "1496982": "Firstly, the PSD should be viewed in a logarithmic scale. Furthermore, there are unofficial PSD's for the three detectors [available here](https://github.com/wvanzeist/riroriro_tutorials/tree/main/noise_spectra). Which originates from [here](https://dcc.ligo.org/LIGO-T1500293/public).\n\nHere is the Virgo PSD for example:\n[![virgo-psd.png](https://i.postimg.cc/Nfmfc6Gx/virgo-psd.png)](https://postimg.cc/D8fKQbXW)",
    "1498444": "I too started with this and did not get it to work  and teamed up with some other who also tried and failed.\n\nI'm not sure, but I understand that the data is all synthetic, i.e. the GW event is calculated and added to synthetic noise. What's not clear to me is how good a job the did in reproducing the noise.",
    "1496422": "Thanks for this very through analysis =))",
    "1496987": "Please don't shot me \nfor asking if you're OK =))",
    "1559949": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
    "1493277": "Thanks for this very through analysis. I can't answer any of your questions for now but will get back if I learn new things. \n\nFor those that are new to the field, **PSD** stands for power [spectral density](https://en.wikipedia.org/wiki/Spectral_density). \n\nAlso, from now on, I will use imgur to host my images since google drive no longer works. 👌"
  }
}