{
  "id": 275390,
  "title": "A Self supervised approach for detecting gravitational waves",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/275390",
  "author_name": "",
  "post_date": "2021-09-30T08:02:55.063087600Z",
  "votes": 18,
  "comment_count": 5,
  "views": 0,
  "content": "<p>First of all, my sincere thanks to the organizers for hosting such an interesting competition. I really enjoyed playing around with the dataset and experimenting with various models. I will describe one of my interesting experiments using semi-supervised learning below. The presented solution scores 0.851 on LB. I came up with this approach a few days ago, so further improvement is definitely possible.</p>\n<p><strong>Context:</strong><br>\nI was really bored with fitting CNNs on CQT spectrograms using supervised learning. So, i decided to try a semi-supervised training approach. One of the main components of semi-supervised training is the augmentation scheme. <br>\nAs many have pointed out, it is difficult to come up with a good augmentation for this dataset. While thinking how to augment the data such that the signal characteristic is preserved, I realized that we already have an augmented dataset :-)<br>\nWhenever there is a gravitational wave, all three detectors must detect it. So, we can imagine that we have access to augmented versions of the same signal!!</p>\n<p><strong>Training details:</strong><br>\nI used the recently proposed Barlow Twins method for semi-supervised training. I would highly recommend going through the <a href=\"https://arxiv.org/pdf/2103.03230.pdf\" target=\"_blank\">paper</a> to understand the details. The basic idea is to have two models which see different versions of data (in our case the gravitational wave signal) and use the barlow twin's objective function to learn embeddings. While training on imagenet, people generally use cropping, flipping, blurring, random contrast etc<br>\nto create different versions of the same image. In our case, we can feed the data from Hanford/Livingston into Net 1 and data from the Virgo detector into Net 2. Since the noise in Virgo detector is quite different from Hanford/Livingston, we can imagine it as an augmented sample of the same underlying signal. Most semi-supervised work is focused on images. So, usually Net 1 and Net 2 are 2D CNN like resnet and the inputs are images. However, for faster training I decided to train using waveforms. Since the input was 1D so as an initial experiment, I tried the <a href=\"https://www.kaggle.com/scaomath/g2net-1d-cnn-gem-pool-pytorch-train-inference\" target=\"_blank\">publicly available 1D CNN</a>.  Outputs of both the networks are normalized and along the batch dimension and we calculate the invariance and redundancy reduction terms. Using these terms,<br>\nwe get the final loss used for training the networks. In the original paper, they used LARS optimizer but I found that AdamW also works. (I haven't tried LARS optimizer yet, so maybe it will work even better.)</p>\n<p><strong>Evaluation details:</strong><br>\nAfter training both CNNs, we are able to get embeddings of the input waveform. We need to train a Fully connected (FC) layer to make final prediction. We obtain embeddings for all three detectors. Embeddings for Hanford/Livingston are obtained from Net 1 and Net 2 provides the<br>\nVirgo embedding. The concatenation of all the embeddings is used as input to the FC layer. <br>\nWe freeze the backbone layer weights (cnn1 and cnn2) and train the FC layer only on a subset of training dataset (i used 10%). It is sufficient to train FC layer for 4-5 epochs. Once the FC training is complete, we can use the model to detect GW.</p>\n<p><strong>Things to try:</strong></p>\n<ul>\n<li>Try this semi-supervised approach with 2D CNN model. In this case, instead of using direct waveforms we can use the extracted CQT features.</li>\n<li>Different 1d CNN architectures.</li>\n<li>Augmentations</li>\n</ul>\n<p>I am planning to share the training &amp; evaluation code later.</p>",
  "messages": [
    {
      "id": "1529186",
      "postDate": "09/30/2021 08:02:55",
      "content": "<p>First of all, my sincere thanks to the organizers for hosting such an interesting competition. I really enjoyed playing around with the dataset and experimenting with various models. I will describe one of my interesting experiments using semi-supervised learning below. The presented solution scores 0.851 on LB. I came up with this approach a few days ago, so further improvement is definitely possible.</p>\n<p><strong>Context:</strong><br>\nI was really bored with fitting CNNs on CQT spectrograms using supervised learning. So, i decided to try a semi-supervised training approach. One of the main components of semi-supervised training is the augmentation scheme. <br>\nAs many have pointed out, it is difficult to come up with a good augmentation for this dataset. While thinking how to augment the data such that the signal characteristic is preserved, I realized that we already have an augmented dataset :-)<br>\nWhenever there is a gravitational wave, all three detectors must detect it. So, we can imagine that we have access to augmented versions of the same signal!!</p>\n<p><strong>Training details:</strong><br>\nI used the recently proposed Barlow Twins method for semi-supervised training. I would highly recommend going through the <a href=\"https://arxiv.org/pdf/2103.03230.pdf\" target=\"_blank\">paper</a> to understand the details. The basic idea is to have two models which see different versions of data (in our case the gravitational wave signal) and use the barlow twin's objective function to learn embeddings. While training on imagenet, people generally use cropping, flipping, blurring, random contrast etc<br>\nto create different versions of the same image. In our case, we can feed the data from Hanford/Livingston into Net 1 and data from the Virgo detector into Net 2. Since the noise in Virgo detector is quite different from Hanford/Livingston, we can imagine it as an augmented sample of the same underlying signal. Most semi-supervised work is focused on images. So, usually Net 1 and Net 2 are 2D CNN like resnet and the inputs are images. However, for faster training I decided to train using waveforms. Since the input was 1D so as an initial experiment, I tried the <a href=\"https://www.kaggle.com/scaomath/g2net-1d-cnn-gem-pool-pytorch-train-inference\" target=\"_blank\">publicly available 1D CNN</a>.  Outputs of both the networks are normalized and along the batch dimension and we calculate the invariance and redundancy reduction terms. Using these terms,<br>\nwe get the final loss used for training the networks. In the original paper, they used LARS optimizer but I found that AdamW also works. (I haven't tried LARS optimizer yet, so maybe it will work even better.)</p>\n<p><strong>Evaluation details:</strong><br>\nAfter training both CNNs, we are able to get embeddings of the input waveform. We need to train a Fully connected (FC) layer to make final prediction. We obtain embeddings for all three detectors. Embeddings for Hanford/Livingston are obtained from Net 1 and Net 2 provides the<br>\nVirgo embedding. The concatenation of all the embeddings is used as input to the FC layer. <br>\nWe freeze the backbone layer weights (cnn1 and cnn2) and train the FC layer only on a subset of training dataset (i used 10%). It is sufficient to train FC layer for 4-5 epochs. Once the FC training is complete, we can use the model to detect GW.</p>\n<p><strong>Things to try:</strong></p>\n<ul>\n<li>Try this semi-supervised approach with 2D CNN model. In this case, instead of using direct waveforms we can use the extracted CQT features.</li>\n<li>Different 1d CNN architectures.</li>\n<li>Augmentations</li>\n</ul>\n<p>I am planning to share the training &amp; evaluation code later.</p>",
      "rawMarkdown": "First of all, my sincere thanks to the organizers for hosting such an interesting competition. I really enjoyed playing around with the dataset and experimenting with various models. I will describe one of my interesting experiments using semi-supervised learning below. The presented solution scores 0.851 on LB. I came up with this approach a few days ago, so further improvement is definitely possible.\n\n**Context:**\nI was really bored with fitting CNNs on CQT spectrograms using supervised learning. So, i decided to try a semi-supervised training approach. One of the main components of semi-supervised training is the augmentation scheme. \nAs many have pointed out, it is difficult to come up with a good augmentation for this dataset. While thinking how to augment the data such that the signal characteristic is preserved, I realized that we already have an augmented dataset :-)\nWhenever there is a gravitational wave, all three detectors must detect it. So, we can imagine that we have access to augmented versions of the same signal!!\n\n**Training details:**\nI used the recently proposed Barlow Twins method for semi-supervised training. I would highly recommend going through the [paper](https://arxiv.org/pdf/2103.03230.pdf) to understand the details. The basic idea is to have two models which see different versions of data (in our case the gravitational wave signal) and use the barlow twin's objective function to learn embeddings. While training on imagenet, people generally use cropping, flipping, blurring, random contrast etc\nto create different versions of the same image. In our case, we can feed the data from Hanford/Livingston into Net 1 and data from the Virgo detector into Net 2. Since the noise in Virgo detector is quite different from Hanford/Livingston, we can imagine it as an augmented sample of the same underlying signal. Most semi-supervised work is focused on images. So, usually Net 1 and Net 2 are 2D CNN like resnet and the inputs are images. However, for faster training I decided to train using waveforms. Since the input was 1D so as an initial experiment, I tried the [publicly available 1D CNN](https://www.kaggle.com/scaomath/g2net-1d-cnn-gem-pool-pytorch-train-inference).  Outputs of both the networks are normalized and along the batch dimension and we calculate the invariance and redundancy reduction terms. Using these terms,\nwe get the final loss used for training the networks. In the original paper, they used LARS optimizer but I found that AdamW also works. (I haven't tried LARS optimizer yet, so maybe it will work even better.)\n\n**Evaluation details:**\nAfter training both CNNs, we are able to get embeddings of the input waveform. We need to train a Fully connected (FC) layer to make final prediction. We obtain embeddings for all three detectors. Embeddings for Hanford/Livingston are obtained from Net 1 and Net 2 provides the\nVirgo embedding. The concatenation of all the embeddings is used as input to the FC layer. \nWe freeze the backbone layer weights (cnn1 and cnn2) and train the FC layer only on a subset of training dataset (i used 10%). It is sufficient to train FC layer for 4-5 epochs. Once the FC training is complete, we can use the model to detect GW.\n\n**Things to try:**\n- Try this semi-supervised approach with 2D CNN model. In this case, instead of using direct waveforms we can use the extracted CQT features.\n- Different 1d CNN architectures.\n- Augmentations\n\nI am planning to share the training & evaluation code later.",
      "votes": null
    },
    {
      "id": "1529257",
      "postDate": "09/30/2021 08:54:38",
      "content": "<p>Nice !! Thank you for sharing.  Looking forward to the training and evaluation code as these are new approach for me 👍 </p>",
      "rawMarkdown": "Nice !! Thank you for sharing.  Looking forward to the training and evaluation code as these are new approach for me 👍",
      "votes": null
    },
    {
      "id": "1529687",
      "postDate": "09/30/2021 15:42:04",
      "content": "<p>Wow, quite interesting approach…nice try!!<br>\nDid you also try other SSL approaches? Also, did you try putting two augmented samples into Siamese architecture instead of comparing the embeddings of Virgo vs Hanford/Livingston?</p>",
      "rawMarkdown": "Wow, quite interesting approach...nice try!!\nDid you also try other SSL approaches? Also, did you try putting two augmented samples into Siamese architecture instead of comparing the embeddings of Virgo vs Hanford/Livingston?",
      "votes": null
    },
    {
      "id": "1530161",
      "postDate": "10/01/2021 01:34:01",
      "content": "<p>Thanks. I haven't tried other approaches like SimCLR, BYOL yet. It would be interesting to try them!<br>\nBarlow twins performs very similar to these approaches on ImageNet (refer Table 1 of paper), so I am expecting them to perform similarly here. But it will be interesting to try!<br>\nI tried adding noise from pycbc. I am not sure to what extent it's helpful. I haven't tried any adding any other augmentation yet. The performance may improve with suitable augmentation.</p>",
      "rawMarkdown": "Thanks. I haven't tried other approaches like SimCLR, BYOL yet. It would be interesting to try them!\nBarlow twins performs very similar to these approaches on ImageNet (refer Table 1 of paper), so I am expecting them to perform similarly here. But it will be interesting to try!\nI tried adding noise from pycbc. I am not sure to what extent it's helpful. I haven't tried any adding any other augmentation yet. The performance may improve with suitable augmentation.",
      "votes": null
    },
    {
      "id": "1532542",
      "postDate": "10/03/2021 05:37:54",
      "content": "<p>UPDATE:<br>\nIf anyone is interested they can find the notebook here <a href=\"https://www.kaggle.com/meaninglesslives/self-supervised-method-for-gravitation-wave-det\" target=\"_blank\">https://www.kaggle.com/meaninglesslives/self-supervised-method-for-gravitation-wave-det</a><br>\nThe notebook is complete in itself. It first trains the backbone model which takes the waveforms as input and outputs embedding for the waveforms. Next the backbone weights are frozen and we train a FC layer. Finally we do test set predictions.</p>\n<p>Github Repo: <a href=\"https://github.com/sidml/Self-Supervised-Learning-for-Gravitational-Waves\" target=\"_blank\">https://github.com/sidml/Self-Supervised-Learning-for-Gravitational-Waves</a></p>",
      "rawMarkdown": "UPDATE:\nIf anyone is interested they can find the notebook here https://www.kaggle.com/meaninglesslives/self-supervised-method-for-gravitation-wave-det\nThe notebook is complete in itself. It first trains the backbone model which takes the waveforms as input and outputs embedding for the waveforms. Next the backbone weights are frozen and we train a FC layer. Finally we do test set predictions.\n\nGithub Repo: https://github.com/sidml/Self-Supervised-Learning-for-Gravitational-Waves",
      "votes": null
    },
    {
      "id": "1559904",
      "postDate": "10/27/2021 08:07:38",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1529257,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "09/30/2021 08:54:38",
      "content": "<p>Nice !! Thank you for sharing.  Looking forward to the training and evaluation code as these are new approach for me 👍 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1529687,
      "author_name": "hidehisaarai1213",
      "author_url": "",
      "post_date": "09/30/2021 15:42:04",
      "content": "<p>Wow, quite interesting approach…nice try!!<br>\nDid you also try other SSL approaches? Also, did you try putting two augmented samples into Siamese architecture instead of comparing the embeddings of Virgo vs Hanford/Livingston?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1530161,
          "author_name": "meaninglesslives",
          "author_url": "",
          "post_date": "10/01/2021 01:34:01",
          "content": "<p>Thanks. I haven't tried other approaches like SimCLR, BYOL yet. It would be interesting to try them!<br>\nBarlow twins performs very similar to these approaches on ImageNet (refer Table 1 of paper), so I am expecting them to perform similarly here. But it will be interesting to try!<br>\nI tried adding noise from pycbc. I am not sure to what extent it's helpful. I haven't tried any adding any other augmentation yet. The performance may improve with suitable augmentation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1532542,
      "author_name": "meaninglesslives",
      "author_url": "",
      "post_date": "10/03/2021 05:37:54",
      "content": "<p>UPDATE:<br>\nIf anyone is interested they can find the notebook here <a href=\"https://www.kaggle.com/meaninglesslives/self-supervised-method-for-gravitation-wave-det\" target=\"_blank\">https://www.kaggle.com/meaninglesslives/self-supervised-method-for-gravitation-wave-det</a><br>\nThe notebook is complete in itself. It first trains the backbone model which takes the waveforms as input and outputs embedding for the waveforms. Next the backbone weights are frozen and we train a FC layer. Finally we do test set predictions.</p>\n<p>Github Repo: <a href=\"https://github.com/sidml/Self-Supervised-Learning-for-Gravitational-Waves\" target=\"_blank\">https://github.com/sidml/Self-Supervised-Learning-for-Gravitational-Waves</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1559904,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 08:07:38",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1529186": "First of all, my sincere thanks to the organizers for hosting such an interesting competition. I really enjoyed playing around with the dataset and experimenting with various models. I will describe one of my interesting experiments using semi-supervised learning below. The presented solution scores 0.851 on LB. I came up with this approach a few days ago, so further improvement is definitely possible.\n\n**Context:**\nI was really bored with fitting CNNs on CQT spectrograms using supervised learning. So, i decided to try a semi-supervised training approach. One of the main components of semi-supervised training is the augmentation scheme. \nAs many have pointed out, it is difficult to come up with a good augmentation for this dataset. While thinking how to augment the data such that the signal characteristic is preserved, I realized that we already have an augmented dataset :-)\nWhenever there is a gravitational wave, all three detectors must detect it. So, we can imagine that we have access to augmented versions of the same signal!!\n\n**Training details:**\nI used the recently proposed Barlow Twins method for semi-supervised training. I would highly recommend going through the [paper](https://arxiv.org/pdf/2103.03230.pdf) to understand the details. The basic idea is to have two models which see different versions of data (in our case the gravitational wave signal) and use the barlow twin's objective function to learn embeddings. While training on imagenet, people generally use cropping, flipping, blurring, random contrast etc\nto create different versions of the same image. In our case, we can feed the data from Hanford/Livingston into Net 1 and data from the Virgo detector into Net 2. Since the noise in Virgo detector is quite different from Hanford/Livingston, we can imagine it as an augmented sample of the same underlying signal. Most semi-supervised work is focused on images. So, usually Net 1 and Net 2 are 2D CNN like resnet and the inputs are images. However, for faster training I decided to train using waveforms. Since the input was 1D so as an initial experiment, I tried the [publicly available 1D CNN](https://www.kaggle.com/scaomath/g2net-1d-cnn-gem-pool-pytorch-train-inference).  Outputs of both the networks are normalized and along the batch dimension and we calculate the invariance and redundancy reduction terms. Using these terms,\nwe get the final loss used for training the networks. In the original paper, they used LARS optimizer but I found that AdamW also works. (I haven't tried LARS optimizer yet, so maybe it will work even better.)\n\n**Evaluation details:**\nAfter training both CNNs, we are able to get embeddings of the input waveform. We need to train a Fully connected (FC) layer to make final prediction. We obtain embeddings for all three detectors. Embeddings for Hanford/Livingston are obtained from Net 1 and Net 2 provides the\nVirgo embedding. The concatenation of all the embeddings is used as input to the FC layer. \nWe freeze the backbone layer weights (cnn1 and cnn2) and train the FC layer only on a subset of training dataset (i used 10%). It is sufficient to train FC layer for 4-5 epochs. Once the FC training is complete, we can use the model to detect GW.\n\n**Things to try:**\n- Try this semi-supervised approach with 2D CNN model. In this case, instead of using direct waveforms we can use the extracted CQT features.\n- Different 1d CNN architectures.\n- Augmentations\n\nI am planning to share the training & evaluation code later.",
    "1529257": "Nice !! Thank you for sharing.  Looking forward to the training and evaluation code as these are new approach for me 👍",
    "1529687": "Wow, quite interesting approach...nice try!!\nDid you also try other SSL approaches? Also, did you try putting two augmented samples into Siamese architecture instead of comparing the embeddings of Virgo vs Hanford/Livingston?",
    "1530161": "Thanks. I haven't tried other approaches like SimCLR, BYOL yet. It would be interesting to try them!\nBarlow twins performs very similar to these approaches on ImageNet (refer Table 1 of paper), so I am expecting them to perform similarly here. But it will be interesting to try!\nI tried adding noise from pycbc. I am not sure to what extent it's helpful. I haven't tried any adding any other augmentation yet. The performance may improve with suitable augmentation.",
    "1532542": "UPDATE:\nIf anyone is interested they can find the notebook here https://www.kaggle.com/meaninglesslives/self-supervised-method-for-gravitation-wave-det\nThe notebook is complete in itself. It first trains the backbone model which takes the waveforms as input and outputs embedding for the waveforms. Next the backbone weights are frozen and we train a FC layer. Finally we do test set predictions.\n\nGithub Repo: https://github.com/sidml/Self-Supervised-Learning-for-Gravitational-Waves",
    "1559904": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}