{
  "id": 275507,
  "title": "Top 1 solution: DSP part",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/275507",
  "author_name": "",
  "post_date": "2021-09-30T16:49:59.383329900Z",
  "votes": 71,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Deep Learning Part: <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275476\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275476</a></p>\n<p><strong>General DSP</strong><br>\nWe were laser-focused on the 1D approach from the very beginning. The best thing the classic approach had to offer was Matched Filters with a huge theoretical basis of why it’s the best thing ever, so it made sense to mimic it. Obviously it was not enough.</p>\n<p>First problem we had to solve was that the NN could not even start learning without filtering. There were lots of examples of filtering in GW papers so as almost all teams here we started with the Butterworth filter. We chose a double application (filtfilt) of 5th order high-pass filter with critical frequency at 20.43Hz. Experiments showed that we can freely decrease order down to 2 and critical frequency down to 15Hz as soon as NN training got some traction, but there was no LB benefit and we decided not to complicate our pipeline.</p>\n<p>While researching the need for any additional filtering PSD of data was obtained using a modified Welch method (we had only 2 seconds pieces with no overlap and we used those with the Hanning window). PSD was obviously not even close to PSD of any public LIGO runs and too smooth for real data. We realized that everything in the data was generated.</p>\n<p>Pieces were too short to use the next step most of GW data processing pipelines converged to - whitening - we decided just to skip it. In the 1D approach NN should derive filters from data and we had no tangible way to “preload” it with a calculated Matched Filter bank, so why bother. Thing most other teams missed here was that whitening is equivalent to applying an arbitrary frequency-domain filter and would mess up portions of a signal. In general pipeline first and last few seconds of a signal was just discarded, but few seconds are all we have here. Later in competition we found a way to “expand” a signal enough to apply whitening, but at that point we felt no need to use it, as separate encoders kicked in.</p>\n<p><strong>Make Some Noise</strong><br>\nAs soon as we realized that noise was generated, we proceeded to generate our own noise using PSD from data - more data is always beneficial in competition where your main enemy is overfitting. Results here were mixed, just noise was obviously not enough, but it helped. And we got a few observations here.</p>\n<ul>\n<li>Channels 0 and 1 had 100% matched PSD - we can swap them freely. Made it to our augmentations for a while.</li>\n<li>Signal was probably generated on higher frequency with idealized PSD. We checked and just generation with observed PSD was good enough, so we decided not to waste time reverse-engineering it further.</li>\n<li>There were few noise injections, most notably 300Hz on both LIGO detectors (Channels 0 and 1). End up in an additional notch filter.</li>\n<li>We compared PSD with and without signal and figured out PSD of pure signals - was helpful later.</li>\n</ul>\n<p><strong>Make It All</strong><br>\nWhy stop at just noise? We proceed to generate merger signals as well. First few attempts was just playing with a public notebook <a href=\"https://www.kaggle.com/mistag/reverse-engineering-create-clean-gw-signals\" target=\"_blank\">https://www.kaggle.com/mistag/reverse-engineering-create-clean-gw-signals</a> but we were able to figure out a few things here as well. Just created a grid of signals with parametrization from the notebook and predicted it with a NN, trained on competition data. </p>\n<ul>\n<li>There were only BH-BH mergers in data and no primordial BHs.</li>\n<li>Minimum chirp mass was estimated as 5, maximum - 50.</li>\n<li>(Almost) All events were in the second half on a signal.</li>\n<li>Proposed method was too slow to generate significant amounts of data.</li>\n</ul>\n<p>So we were trying to find a better way to generate signals. And we found it right on the Overview page of the competition. We managed to figure out the correct way to generate signals in PyCBC/LALSuite just by reading the docs and experimenting. Page <a href=\"https://pycbc.org/pycbc/latest/html/waveform.html\" target=\"_blank\">https://pycbc.org/pycbc/latest/html/waveform.html</a> contains almost complete example on GW generation, but with some quirks. We used SEOBNRv4_opt approximant for speed and functionality, but it was not able to work with our sample frequency, so generation was done at higher frequency and then downsampled. We believe that we were able to reproduce the data generation pipeline with sufficient quality, but we still were lacking on parameters’ distributions. An attempt to generate data with just eyeballed parameters’ ranges and distributions was not really beneficial. NN was quick to converge but the result was useless on competition data.</p>\n<p>Here was the point where all little notes from before got us the answer. We got initial ranges for chirp mass from playing with the notebook earlier. Event timings were known from there as well. PSDs difference yielded first estimations on distance (500, 3000). With set ranges we were able to generate a dataset. As parameters for generated signals were known for us, we were able to include additional targets in our learning process. </p>\n<p>We engaged in an iterative process of selecting the next parameter for distribution improvement, training NN with additional regression for the target, adjusting the distribution of selected parameter and generating a new dataset. In a few iterations we got decent distribution for chirp mass and distance. At that point we had a virtually unlimited source of data and we were able to create pretrain. As we had all the details, we were able to calculate SNR for all our samples. SNR, added to regression targets, helped a little.</p>\n<p>We figured out that quite a lot of our samples are indistinguishable from noise (if noise present) by SNR. We believe that it is the case for competition data too. We noted the same “inversion of labels” behavior on our generated data. It looks like we were close enough to the theoretical limit of what was possible in given circumstances, so we have to commend other top teams for giving us a tough fight.</p>\n<p>Please, comment if you have any questions. <br>\nGitHub: <a href=\"https://github.com/denisbsu/g2net-generation\" target=\"_blank\">https://github.com/denisbsu/g2net-generation</a></p>",
  "messages": [
    {
      "id": "1529789",
      "postDate": "09/30/2021 16:49:59",
      "content": "<p>Deep Learning Part: <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275476\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275476</a></p>\n<p><strong>General DSP</strong><br>\nWe were laser-focused on the 1D approach from the very beginning. The best thing the classic approach had to offer was Matched Filters with a huge theoretical basis of why it’s the best thing ever, so it made sense to mimic it. Obviously it was not enough.</p>\n<p>First problem we had to solve was that the NN could not even start learning without filtering. There were lots of examples of filtering in GW papers so as almost all teams here we started with the Butterworth filter. We chose a double application (filtfilt) of 5th order high-pass filter with critical frequency at 20.43Hz. Experiments showed that we can freely decrease order down to 2 and critical frequency down to 15Hz as soon as NN training got some traction, but there was no LB benefit and we decided not to complicate our pipeline.</p>\n<p>While researching the need for any additional filtering PSD of data was obtained using a modified Welch method (we had only 2 seconds pieces with no overlap and we used those with the Hanning window). PSD was obviously not even close to PSD of any public LIGO runs and too smooth for real data. We realized that everything in the data was generated.</p>\n<p>Pieces were too short to use the next step most of GW data processing pipelines converged to - whitening - we decided just to skip it. In the 1D approach NN should derive filters from data and we had no tangible way to “preload” it with a calculated Matched Filter bank, so why bother. Thing most other teams missed here was that whitening is equivalent to applying an arbitrary frequency-domain filter and would mess up portions of a signal. In general pipeline first and last few seconds of a signal was just discarded, but few seconds are all we have here. Later in competition we found a way to “expand” a signal enough to apply whitening, but at that point we felt no need to use it, as separate encoders kicked in.</p>\n<p><strong>Make Some Noise</strong><br>\nAs soon as we realized that noise was generated, we proceeded to generate our own noise using PSD from data - more data is always beneficial in competition where your main enemy is overfitting. Results here were mixed, just noise was obviously not enough, but it helped. And we got a few observations here.</p>\n<ul>\n<li>Channels 0 and 1 had 100% matched PSD - we can swap them freely. Made it to our augmentations for a while.</li>\n<li>Signal was probably generated on higher frequency with idealized PSD. We checked and just generation with observed PSD was good enough, so we decided not to waste time reverse-engineering it further.</li>\n<li>There were few noise injections, most notably 300Hz on both LIGO detectors (Channels 0 and 1). End up in an additional notch filter.</li>\n<li>We compared PSD with and without signal and figured out PSD of pure signals - was helpful later.</li>\n</ul>\n<p><strong>Make It All</strong><br>\nWhy stop at just noise? We proceed to generate merger signals as well. First few attempts was just playing with a public notebook <a href=\"https://www.kaggle.com/mistag/reverse-engineering-create-clean-gw-signals\" target=\"_blank\">https://www.kaggle.com/mistag/reverse-engineering-create-clean-gw-signals</a> but we were able to figure out a few things here as well. Just created a grid of signals with parametrization from the notebook and predicted it with a NN, trained on competition data. </p>\n<ul>\n<li>There were only BH-BH mergers in data and no primordial BHs.</li>\n<li>Minimum chirp mass was estimated as 5, maximum - 50.</li>\n<li>(Almost) All events were in the second half on a signal.</li>\n<li>Proposed method was too slow to generate significant amounts of data.</li>\n</ul>\n<p>So we were trying to find a better way to generate signals. And we found it right on the Overview page of the competition. We managed to figure out the correct way to generate signals in PyCBC/LALSuite just by reading the docs and experimenting. Page <a href=\"https://pycbc.org/pycbc/latest/html/waveform.html\" target=\"_blank\">https://pycbc.org/pycbc/latest/html/waveform.html</a> contains almost complete example on GW generation, but with some quirks. We used SEOBNRv4_opt approximant for speed and functionality, but it was not able to work with our sample frequency, so generation was done at higher frequency and then downsampled. We believe that we were able to reproduce the data generation pipeline with sufficient quality, but we still were lacking on parameters’ distributions. An attempt to generate data with just eyeballed parameters’ ranges and distributions was not really beneficial. NN was quick to converge but the result was useless on competition data.</p>\n<p>Here was the point where all little notes from before got us the answer. We got initial ranges for chirp mass from playing with the notebook earlier. Event timings were known from there as well. PSDs difference yielded first estimations on distance (500, 3000). With set ranges we were able to generate a dataset. As parameters for generated signals were known for us, we were able to include additional targets in our learning process. </p>\n<p>We engaged in an iterative process of selecting the next parameter for distribution improvement, training NN with additional regression for the target, adjusting the distribution of selected parameter and generating a new dataset. In a few iterations we got decent distribution for chirp mass and distance. At that point we had a virtually unlimited source of data and we were able to create pretrain. As we had all the details, we were able to calculate SNR for all our samples. SNR, added to regression targets, helped a little.</p>\n<p>We figured out that quite a lot of our samples are indistinguishable from noise (if noise present) by SNR. We believe that it is the case for competition data too. We noted the same “inversion of labels” behavior on our generated data. It looks like we were close enough to the theoretical limit of what was possible in given circumstances, so we have to commend other top teams for giving us a tough fight.</p>\n<p>Please, comment if you have any questions. <br>\nGitHub: <a href=\"https://github.com/denisbsu/g2net-generation\" target=\"_blank\">https://github.com/denisbsu/g2net-generation</a></p>",
      "rawMarkdown": "Deep Learning Part: https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275476\n\n**General DSP**\nWe were laser-focused on the 1D approach from the very beginning. The best thing the classic approach had to offer was Matched Filters with a huge theoretical basis of why it’s the best thing ever, so it made sense to mimic it. Obviously it was not enough.\n\nFirst problem we had to solve was that the NN could not even start learning without filtering. There were lots of examples of filtering in GW papers so as almost all teams here we started with the Butterworth filter. We chose a double application (filtfilt) of 5th order high-pass filter with critical frequency at 20.43Hz. Experiments showed that we can freely decrease order down to 2 and critical frequency down to 15Hz as soon as NN training got some traction, but there was no LB benefit and we decided not to complicate our pipeline.\n\nWhile researching the need for any additional filtering PSD of data was obtained using a modified Welch method (we had only 2 seconds pieces with no overlap and we used those with the Hanning window). PSD was obviously not even close to PSD of any public LIGO runs and too smooth for real data. We realized that everything in the data was generated.\n\nPieces were too short to use the next step most of GW data processing pipelines converged to - whitening - we decided just to skip it. In the 1D approach NN should derive filters from data and we had no tangible way to “preload” it with a calculated Matched Filter bank, so why bother. Thing most other teams missed here was that whitening is equivalent to applying an arbitrary frequency-domain filter and would mess up portions of a signal. In general pipeline first and last few seconds of a signal was just discarded, but few seconds are all we have here. Later in competition we found a way to “expand” a signal enough to apply whitening, but at that point we felt no need to use it, as separate encoders kicked in.\n\n\n**Make Some Noise**\nAs soon as we realized that noise was generated, we proceeded to generate our own noise using PSD from data - more data is always beneficial in competition where your main enemy is overfitting. Results here were mixed, just noise was obviously not enough, but it helped. And we got a few observations here.\n- Channels 0 and 1 had 100% matched PSD - we can swap them freely. Made it to our augmentations for a while.\n- Signal was probably generated on higher frequency with idealized PSD. We checked and just generation with observed PSD was good enough, so we decided not to waste time reverse-engineering it further.\n- There were few noise injections, most notably 300Hz on both LIGO detectors (Channels 0 and 1). End up in an additional notch filter.\n- We compared PSD with and without signal and figured out PSD of pure signals - was helpful later.\n\n**Make It All**\nWhy stop at just noise? We proceed to generate merger signals as well. First few attempts was just playing with a public notebook https://www.kaggle.com/mistag/reverse-engineering-create-clean-gw-signals but we were able to figure out a few things here as well. Just created a grid of signals with parametrization from the notebook and predicted it with a NN, trained on competition data. \n- There were only BH-BH mergers in data and no primordial BHs.\n- Minimum chirp mass was estimated as 5, maximum - 50.\n- (Almost) All events were in the second half on a signal.\n- Proposed method was too slow to generate significant amounts of data.\n\nSo we were trying to find a better way to generate signals. And we found it right on the Overview page of the competition. We managed to figure out the correct way to generate signals in PyCBC/LALSuite just by reading the docs and experimenting. Page https://pycbc.org/pycbc/latest/html/waveform.html contains almost complete example on GW generation, but with some quirks. We used SEOBNRv4_opt approximant for speed and functionality, but it was not able to work with our sample frequency, so generation was done at higher frequency and then downsampled. We believe that we were able to reproduce the data generation pipeline with sufficient quality, but we still were lacking on parameters’ distributions. An attempt to generate data with just eyeballed parameters’ ranges and distributions was not really beneficial. NN was quick to converge but the result was useless on competition data.\n\nHere was the point where all little notes from before got us the answer. We got initial ranges for chirp mass from playing with the notebook earlier. Event timings were known from there as well. PSDs difference yielded first estimations on distance (500, 3000). With set ranges we were able to generate a dataset. As parameters for generated signals were known for us, we were able to include additional targets in our learning process. \n\nWe engaged in an iterative process of selecting the next parameter for distribution improvement, training NN with additional regression for the target, adjusting the distribution of selected parameter and generating a new dataset. In a few iterations we got decent distribution for chirp mass and distance. At that point we had a virtually unlimited source of data and we were able to create pretrain. As we had all the details, we were able to calculate SNR for all our samples. SNR, added to regression targets, helped a little.\n\n We figured out that quite a lot of our samples are indistinguishable from noise (if noise present) by SNR. We believe that it is the case for competition data too. We noted the same “inversion of labels” behavior on our generated data. It looks like we were close enough to the theoretical limit of what was possible in given circumstances, so we have to commend other top teams for giving us a tough fight.\n\n\nPlease, comment if you have any questions. \nGitHub: https://github.com/denisbsu/g2net-generation",
      "votes": null
    },
    {
      "id": "1529810",
      "postDate": "09/30/2021 17:04:38",
      "content": "<p>Great work!</p>\n<p>Once again, retro-engineering the data generator led to a competition win.  And only one team did it.</p>\n<p>Will you open source your data generator?</p>",
      "rawMarkdown": "Great work!\n\nOnce again, retro-engineering the data generator led to a competition win.  And only one team did it.\n\nWill you open source your data generator?",
      "votes": null
    },
    {
      "id": "1529823",
      "postDate": "09/30/2021 17:16:19",
      "content": "<blockquote>\n  <p>and we had no tangible way to “preload” it with a calculated Matched Filter bank</p>\n</blockquote>\n<p>What do you mean by this? If you're speaking of ~250k GWs filters perhaps. But if it's the same number as whatever # of filters your first convolutional layer has, then… that should be quite straightforward?</p>",
      "rawMarkdown": "> and we had no tangible way to “preload” it with a calculated Matched Filter bank\n\nWhat do you mean by this? If you're speaking of ~250k GWs filters perhaps. But if it's the same number as whatever # of filters your first convolutional layer has, then... that should be quite straightforward?",
      "votes": null
    },
    {
      "id": "1529836",
      "postDate": "09/30/2021 17:31:21",
      "content": "<p>It's 1-4k filters in a filter bank to cover 5-50 range with decent accuracy. But the main problem it's that raw filters are <em>long</em> - easily longer that our piece itself, but most of them are 512-2048.</p>",
      "rawMarkdown": "It's 1-4k filters in a filter bank to cover 5-50 range with decent accuracy. But the main problem it's that raw filters are _long_ - easily longer that our piece itself, but most of them are 512-2048.",
      "votes": null
    },
    {
      "id": "1529839",
      "postDate": "09/30/2021 17:32:15",
      "content": "<p>Yes, that's a great idea. I'll share it in about half an hour.</p>",
      "rawMarkdown": "Yes, that's a great idea. I'll share it in about half an hour.",
      "votes": null
    },
    {
      "id": "1529840",
      "postDate": "09/30/2021 17:34:19",
      "content": "<p>while you don't share here is mine from 4 week ago: <a href=\"https://www.kaggle.com/titericz/simulated-gw/notebook\" target=\"_blank\">https://www.kaggle.com/titericz/simulated-gw/notebook</a></p>",
      "rawMarkdown": "while you don't share here is mine from 4 week ago: https://www.kaggle.com/titericz/simulated-gw/notebook",
      "votes": null
    },
    {
      "id": "1529846",
      "postDate": "09/30/2021 17:37:01",
      "content": "<p>Understood</p>",
      "rawMarkdown": "Understood",
      "votes": null
    },
    {
      "id": "1529857",
      "postDate": "09/30/2021 17:43:33",
      "content": "<p>Here we go: <a href=\"https://github.com/denisbsu/g2net-generation\" target=\"_blank\">https://github.com/denisbsu/g2net-generation</a></p>",
      "rawMarkdown": "Here we go: https://github.com/denisbsu/g2net-generation",
      "votes": null
    },
    {
      "id": "1529909",
      "postDate": "09/30/2021 18:48:19",
      "content": "<p>I think you suggested it before, but hosts should just open source the data generator from the start.</p>",
      "rawMarkdown": "I think you suggested it before, but hosts should just open source the data generator from the start.",
      "votes": null
    },
    {
      "id": "1529939",
      "postDate": "09/30/2021 19:20:28",
      "content": "<p>maybe on next G2Net competition</p>",
      "rawMarkdown": "maybe on next G2Net competition",
      "votes": null
    },
    {
      "id": "1529944",
      "postDate": "09/30/2021 19:27:01",
      "content": "<p>very good work.</p>\n<p>thanks for the post on resnet34.<br>\nthis is a turning point for me in this competition</p>",
      "rawMarkdown": "very good work.\n\nthanks for the post on resnet34.\nthis is a turning point for me in this competition",
      "votes": null
    },
    {
      "id": "1530330",
      "postDate": "10/01/2021 05:24:52",
      "content": "<p>Congratulations for the win and thanks for the great write-up! At some point(~0.8850LB), it was obvious for us that you guys were generating synthetic data, but still we couldn't make the idea work as well. </p>",
      "rawMarkdown": "Congratulations for the win and thanks for the great write-up! At some point(~0.8850LB), it was obvious for us that you guys were generating synthetic data, but still we couldn't make the idea work as well.",
      "votes": null
    },
    {
      "id": "1559895",
      "postDate": "10/27/2021 08:06:48",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    },
    {
      "id": "2504098",
      "postDate": "10/29/2023 17:02:33",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/denisbsu\" target=\"_blank\">@denisbsu</a>, where one can get your complete code for this competition ? </p>",
      "rawMarkdown": "Hi @denisbsu, where one can get your complete code for this competition ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1529810,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "09/30/2021 17:04:38",
      "content": "<p>Great work!</p>\n<p>Once again, retro-engineering the data generator led to a competition win.  And only one team did it.</p>\n<p>Will you open source your data generator?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1529839,
          "author_name": "denisbsu",
          "author_url": "",
          "post_date": "09/30/2021 17:32:15",
          "content": "<p>Yes, that's a great idea. I'll share it in about half an hour.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1529840,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "09/30/2021 17:34:19",
          "content": "<p>while you don't share here is mine from 4 week ago: <a href=\"https://www.kaggle.com/titericz/simulated-gw/notebook\" target=\"_blank\">https://www.kaggle.com/titericz/simulated-gw/notebook</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1529857,
          "author_name": "denisbsu",
          "author_url": "",
          "post_date": "09/30/2021 17:43:33",
          "content": "<p>Here we go: <a href=\"https://github.com/denisbsu/g2net-generation\" target=\"_blank\">https://github.com/denisbsu/g2net-generation</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1529909,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "09/30/2021 18:48:19",
          "content": "<p>I think you suggested it before, but hosts should just open source the data generator from the start.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1529939,
          "author_name": "killimi",
          "author_url": "",
          "post_date": "09/30/2021 19:20:28",
          "content": "<p>maybe on next G2Net competition</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1529823,
      "author_name": "authman",
      "author_url": "",
      "post_date": "09/30/2021 17:16:19",
      "content": "<blockquote>\n  <p>and we had no tangible way to “preload” it with a calculated Matched Filter bank</p>\n</blockquote>\n<p>What do you mean by this? If you're speaking of ~250k GWs filters perhaps. But if it's the same number as whatever # of filters your first convolutional layer has, then… that should be quite straightforward?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1529836,
          "author_name": "denisbsu",
          "author_url": "",
          "post_date": "09/30/2021 17:31:21",
          "content": "<p>It's 1-4k filters in a filter bank to cover 5-50 range with decent accuracy. But the main problem it's that raw filters are <em>long</em> - easily longer that our piece itself, but most of them are 512-2048.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1529846,
          "author_name": "authman",
          "author_url": "",
          "post_date": "09/30/2021 17:37:01",
          "content": "<p>Understood</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1529944,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/30/2021 19:27:01",
      "content": "<p>very good work.</p>\n<p>thanks for the post on resnet34.<br>\nthis is a turning point for me in this competition</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1530330,
      "author_name": "richx86",
      "author_url": "",
      "post_date": "10/01/2021 05:24:52",
      "content": "<p>Congratulations for the win and thanks for the great write-up! At some point(~0.8850LB), it was obvious for us that you guys were generating synthetic data, but still we couldn't make the idea work as well. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1559895,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 08:06:48",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2504098,
      "author_name": "astroboysanju",
      "author_url": "",
      "post_date": "10/29/2023 17:02:33",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/denisbsu\" target=\"_blank\">@denisbsu</a>, where one can get your complete code for this competition ? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1529789": "Deep Learning Part: https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275476\n\n**General DSP**\nWe were laser-focused on the 1D approach from the very beginning. The best thing the classic approach had to offer was Matched Filters with a huge theoretical basis of why it’s the best thing ever, so it made sense to mimic it. Obviously it was not enough.\n\nFirst problem we had to solve was that the NN could not even start learning without filtering. There were lots of examples of filtering in GW papers so as almost all teams here we started with the Butterworth filter. We chose a double application (filtfilt) of 5th order high-pass filter with critical frequency at 20.43Hz. Experiments showed that we can freely decrease order down to 2 and critical frequency down to 15Hz as soon as NN training got some traction, but there was no LB benefit and we decided not to complicate our pipeline.\n\nWhile researching the need for any additional filtering PSD of data was obtained using a modified Welch method (we had only 2 seconds pieces with no overlap and we used those with the Hanning window). PSD was obviously not even close to PSD of any public LIGO runs and too smooth for real data. We realized that everything in the data was generated.\n\nPieces were too short to use the next step most of GW data processing pipelines converged to - whitening - we decided just to skip it. In the 1D approach NN should derive filters from data and we had no tangible way to “preload” it with a calculated Matched Filter bank, so why bother. Thing most other teams missed here was that whitening is equivalent to applying an arbitrary frequency-domain filter and would mess up portions of a signal. In general pipeline first and last few seconds of a signal was just discarded, but few seconds are all we have here. Later in competition we found a way to “expand” a signal enough to apply whitening, but at that point we felt no need to use it, as separate encoders kicked in.\n\n\n**Make Some Noise**\nAs soon as we realized that noise was generated, we proceeded to generate our own noise using PSD from data - more data is always beneficial in competition where your main enemy is overfitting. Results here were mixed, just noise was obviously not enough, but it helped. And we got a few observations here.\n- Channels 0 and 1 had 100% matched PSD - we can swap them freely. Made it to our augmentations for a while.\n- Signal was probably generated on higher frequency with idealized PSD. We checked and just generation with observed PSD was good enough, so we decided not to waste time reverse-engineering it further.\n- There were few noise injections, most notably 300Hz on both LIGO detectors (Channels 0 and 1). End up in an additional notch filter.\n- We compared PSD with and without signal and figured out PSD of pure signals - was helpful later.\n\n**Make It All**\nWhy stop at just noise? We proceed to generate merger signals as well. First few attempts was just playing with a public notebook https://www.kaggle.com/mistag/reverse-engineering-create-clean-gw-signals but we were able to figure out a few things here as well. Just created a grid of signals with parametrization from the notebook and predicted it with a NN, trained on competition data. \n- There were only BH-BH mergers in data and no primordial BHs.\n- Minimum chirp mass was estimated as 5, maximum - 50.\n- (Almost) All events were in the second half on a signal.\n- Proposed method was too slow to generate significant amounts of data.\n\nSo we were trying to find a better way to generate signals. And we found it right on the Overview page of the competition. We managed to figure out the correct way to generate signals in PyCBC/LALSuite just by reading the docs and experimenting. Page https://pycbc.org/pycbc/latest/html/waveform.html contains almost complete example on GW generation, but with some quirks. We used SEOBNRv4_opt approximant for speed and functionality, but it was not able to work with our sample frequency, so generation was done at higher frequency and then downsampled. We believe that we were able to reproduce the data generation pipeline with sufficient quality, but we still were lacking on parameters’ distributions. An attempt to generate data with just eyeballed parameters’ ranges and distributions was not really beneficial. NN was quick to converge but the result was useless on competition data.\n\nHere was the point where all little notes from before got us the answer. We got initial ranges for chirp mass from playing with the notebook earlier. Event timings were known from there as well. PSDs difference yielded first estimations on distance (500, 3000). With set ranges we were able to generate a dataset. As parameters for generated signals were known for us, we were able to include additional targets in our learning process. \n\nWe engaged in an iterative process of selecting the next parameter for distribution improvement, training NN with additional regression for the target, adjusting the distribution of selected parameter and generating a new dataset. In a few iterations we got decent distribution for chirp mass and distance. At that point we had a virtually unlimited source of data and we were able to create pretrain. As we had all the details, we were able to calculate SNR for all our samples. SNR, added to regression targets, helped a little.\n\n We figured out that quite a lot of our samples are indistinguishable from noise (if noise present) by SNR. We believe that it is the case for competition data too. We noted the same “inversion of labels” behavior on our generated data. It looks like we were close enough to the theoretical limit of what was possible in given circumstances, so we have to commend other top teams for giving us a tough fight.\n\n\nPlease, comment if you have any questions. \nGitHub: https://github.com/denisbsu/g2net-generation",
    "1529810": "Great work!\n\nOnce again, retro-engineering the data generator led to a competition win.  And only one team did it.\n\nWill you open source your data generator?",
    "1529823": "> and we had no tangible way to “preload” it with a calculated Matched Filter bank\n\nWhat do you mean by this? If you're speaking of ~250k GWs filters perhaps. But if it's the same number as whatever # of filters your first convolutional layer has, then... that should be quite straightforward?",
    "1529836": "It's 1-4k filters in a filter bank to cover 5-50 range with decent accuracy. But the main problem it's that raw filters are _long_ - easily longer that our piece itself, but most of them are 512-2048.",
    "1529839": "Yes, that's a great idea. I'll share it in about half an hour.",
    "1529840": "while you don't share here is mine from 4 week ago: https://www.kaggle.com/titericz/simulated-gw/notebook",
    "1529846": "Understood",
    "1529857": "Here we go: https://github.com/denisbsu/g2net-generation",
    "1529909": "I think you suggested it before, but hosts should just open source the data generator from the start.",
    "1529939": "maybe on next G2Net competition",
    "1529944": "very good work.\n\nthanks for the post on resnet34.\nthis is a turning point for me in this competition",
    "1530330": "Congratulations for the win and thanks for the great write-up! At some point(~0.8850LB), it was obvious for us that you guys were generating synthetic data, but still we couldn't make the idea work as well.",
    "1559895": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
    "2504098": "Hi @denisbsu, where one can get your complete code for this competition ?"
  },
  "source": "meta"
}