{
  "id": 275476,
  "title": "Top 1 solution: Deep Learning part",
  "url": "/competitions/g2net-gravitational-wave-detection/writeups/kdl-top-1-solution-deep-learning-part",
  "author_name": "",
  "post_date": "2021-10-02T11:22:10.123Z",
  "votes": 132,
  "comment_count": 27,
  "views": 0,
  "content": "<p>We decided to make two posts to make them more or less  focused and concise. <br>\n DSP part  <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275507\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275507</a></p>\n<h3>Baseline solution, simple Conv1D (0.88 public LB)</h3>\n<p>At the very beginning we tried  different inputs:</p>\n<ol>\n<li>CQT - could not make it past 0.87</li>\n<li>Spectrograms - with nnAudio, a bit better</li>\n<li>raw signal - much better </li>\n</ol>\n<p>After applying highpass filter at 20Hz  1D conv  on stacked channels input quickly produced good val scores. Also it was much easier to make experiments as training on a single GPU took less than 1 hour. The model was very simple, just a bunch of conv-bn-silu blocks and maxpools. For convolutions we used kernel sizes (by block) 64- &gt; 32 -&gt; 16 -&gt; 8. </p>\n<p>At that stage for augmentations we used:</p>\n<p>channel shuffle (only Hanford, Livingston)<br>\nminor time shifts between channels, 5ms</p>\n<p>This approach gave us 0.877 public LB score with a single 5fold model.</p>\n<p><strong>SGD is better than AdamW for Conv1D without synthetic data</strong></p>\n<p>Even though model capacity was quite low, the model quickly overfitted, usually after 20 epochs, which took around 30 mins. <br>\nSwitching to SGD with weight decay and nesterov momentum improved LB score to <strong>0.880</strong>. Other Optimizers like AdamW with whatever weight decay, MadGrad gave lower quality.</p>\n<h3>Improved Conv1D model (0.883 public LB)</h3>\n<p>At that time the generated synthetic dataset was not good enough and we felt that 1D model could be improved. <br>\nFollowing the Inception V3 approach we added different kernel sizes in each conv block, the starting block was with 32, 64, 128 kernel sizes. <br>\nThis gave us a boost from 0.88 to 0.881 on the LB.<br>\nNext step was to try even more kernel sizes as  it also made sense from DSP theory. Inception like block with 5 different kernels (16, 32, 64, 128, 256) allowed us to get 0.8823 single model LB score. <br>\nA small ensemble improved LB score to <strong>0.883</strong>. <br>\nAdding more kernel sizes did not improve CV/LB scores.</p>\n<pre><code># Main building block for Conv1D models\nclass ConcatBlockConv5(nn.Module):\n    def __init__(self, in_ch, out_ch, k, act=nn.SiLU):\n        super().__init__()\n        self.c1 = conv_bn_silu_block(in_ch, out_ch, k, act)\n        self.c2 = conv_bn_silu_block(in_ch, out_ch, k * 2, act)\n        self.c3 = conv_bn_silu_block(in_ch, out_ch, k // 2, act)\n        self.c4 = conv_bn_silu_block(in_ch, out_ch, k // 4, act)\n        self.c5 = conv_bn_silu_block(in_ch, out_ch, k * 4, act)\n        self.c6 = conv_bn_silu_block(in_ch * 5 + in_ch, out_ch, 1, act)\n    def forward(self, x):\n        x = torch.cat([self.c1(x), self.c2(x), self.c3(x), self.c4(x), self.c5(x), x], dim=1)\n        x = self.c6(x)\n        return x\n</code></pre>\n<p>Hyperparameters</p>\n<ul>\n<li>optimizer: SGD, wd=1e-4, nesterov momentum</li>\n<li>learning rate: 0.1 with cosine annealing</li>\n<li>batch size: 128</li>\n<li>epochs: 40 </li>\n<li>input: 3 channels of raw signal filtered with butterworth filter at 20hz</li>\n<li>augmentations: freq masking, time masking, small shifts, channel shuffle</li>\n</ul>\n<h3>Using synthetic dataset (0.886 public LB)</h3>\n<p>As soon as Denis found a more or less good approach to signal/noise generation we started experimenting with additional data. <br>\nOverall we had 2 million noise samples and 1 million pure signal samples. During training positive sample = random noise sample + random signal sample.</p>\n<p>From these experiments </p>\n<ul>\n<li>mixing synthetic data with the train dataset did not work </li>\n<li>augmentations are actually harmful in this case</li>\n<li>pretraining on synthetic data and fine tuning on the train set works great</li>\n</ul>\n<p>During pretraining stage for simplicity we used the same amount of samples in epoch as in the train set.<br>\nPre-training around 100 epochs and fine tuning 5 folds on the train set gave 8836 on the public LB for the single Conv1d model. <br>\nAs detectors, especially Virgo, have different noise distribution it makes sense to use a separate conv1d encoder for each channel. Split encoders and a linear classifier on top of concatenated features boosted the LB score to <strong>0.8842</strong>.</p>\n<p>It is clear that a fully connected layer is not the best fusion approach for  the model with separate  encoders for each channel. <br>\nThat’s where resnet34 came into play and surprisingly it worked better than other 2D models. We also predicted signal parameters during pretraining  (SNR, chirp mass, Q) which also brought minor improvements. <br>\nPretraining 200 epochs  and fine tuning 5 folds just 1 epoch gives <strong>0.8858</strong> public LB score.<br>\nAugmentations during finetuning or pretraining negatively affected CV, so the best models are without any augmentations and trained with AdamW optimizer.<br>\n<img src=\"https://i.imgur.com/t1X1mL4.png\" alt=\"model\"></p>\n<p>Input to Resnet34 looked the following way (Handford band)</p>\n<p><img src=\"https://i.imgur.com/iMc4Yrp.jpeg\" alt=\"1D features\"></p>\n<h3>Segmentation</h3>\n<p>Binary segmentation using output of Conv1D predicted good masks for strong signals but did not improve recall on weak signals. In general it could be a useful tool to analyse the data, but we did not get any boost on the LB from that.<br>\n<img src=\"https://i.imgur.com/zLPtw6S.png\" alt=\"good signal segmentation\"></p>\n<h3>Things that did not work</h3>\n<p>There were much more experiments that I won't describe (including different frontends, training approaches etc.), but most noticeable are:</p>\n<p><strong>Denoising autoencoder</strong></p>\n<p>I trained different variants of autoencoders to separate noise and signals which worked great for strong signals but produced poor results on medium to low amplitude signals.</p>\n<p><strong>OHEM collapse and reverse labels mystery</strong></p>\n<p>I tried different versions of hard example mining to improve model performance on hard samples but usually the model collapsed and started predicting the same probability for all samples. </p>\n<p>Which led to an interesting experiment:</p>\n<ul>\n<li>from full OOF predictions select positive samples with low probability and negative with high probability</li>\n<li>train on this subset but validate on proper split</li>\n<li>evaluate using predicted probabilities</li>\n<li>evaluate using  reversed predctions (1 - p) </li>\n</ul>\n<p><img src=\"https://i.imgur.com/ZbnpqeV.png\" alt=\"\"></p>\n<p>That result was really confusing and at first we thought that the dataset was mislabeled. Later even with generated synthetic data we had the same problem. <br>\nIt is clear that because of the SNR wall some positive samples can be considered as just noise, but how the model generalized to predict signal from noise samples, that’s what we could not find. </p>",
  "messages": [
    {
      "id": "1529717",
      "postDate": "09/30/2021 16:07:29",
      "content": "<p>We decided to make two posts to make them more or less  focused and concise. <br>\n DSP part  <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275507\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275507</a></p>\n<h3>Baseline solution, simple Conv1D (0.88 public LB)</h3>\n<p>At the very beginning we tried  different inputs:</p>\n<ol>\n<li>CQT - could not make it past 0.87</li>\n<li>Spectrograms - with nnAudio, a bit better</li>\n<li>raw signal - much better </li>\n</ol>\n<p>After applying highpass filter at 20Hz  1D conv  on stacked channels input quickly produced good val scores. Also it was much easier to make experiments as training on a single GPU took less than 1 hour. The model was very simple, just a bunch of conv-bn-silu blocks and maxpools. For convolutions we used kernel sizes (by block) 64- &gt; 32 -&gt; 16 -&gt; 8. </p>\n<p>At that stage for augmentations we used:</p>\n<p>channel shuffle (only Hanford, Livingston)<br>\nminor time shifts between channels, 5ms</p>\n<p>This approach gave us 0.877 public LB score with a single 5fold model.</p>\n<p><strong>SGD is better than AdamW for Conv1D without synthetic data</strong></p>\n<p>Even though model capacity was quite low, the model quickly overfitted, usually after 20 epochs, which took around 30 mins. <br>\nSwitching to SGD with weight decay and nesterov momentum improved LB score to <strong>0.880</strong>. Other Optimizers like AdamW with whatever weight decay, MadGrad gave lower quality.</p>\n<h3>Improved Conv1D model (0.883 public LB)</h3>\n<p>At that time the generated synthetic dataset was not good enough and we felt that 1D model could be improved. <br>\nFollowing the Inception V3 approach we added different kernel sizes in each conv block, the starting block was with 32, 64, 128 kernel sizes. <br>\nThis gave us a boost from 0.88 to 0.881 on the LB.<br>\nNext step was to try even more kernel sizes as  it also made sense from DSP theory. Inception like block with 5 different kernels (16, 32, 64, 128, 256) allowed us to get 0.8823 single model LB score. <br>\nA small ensemble improved LB score to <strong>0.883</strong>. <br>\nAdding more kernel sizes did not improve CV/LB scores.</p>\n<pre><code># Main building block for Conv1D models\nclass ConcatBlockConv5(nn.Module):\n    def __init__(self, in_ch, out_ch, k, act=nn.SiLU):\n        super().__init__()\n        self.c1 = conv_bn_silu_block(in_ch, out_ch, k, act)\n        self.c2 = conv_bn_silu_block(in_ch, out_ch, k * 2, act)\n        self.c3 = conv_bn_silu_block(in_ch, out_ch, k // 2, act)\n        self.c4 = conv_bn_silu_block(in_ch, out_ch, k // 4, act)\n        self.c5 = conv_bn_silu_block(in_ch, out_ch, k * 4, act)\n        self.c6 = conv_bn_silu_block(in_ch * 5 + in_ch, out_ch, 1, act)\n    def forward(self, x):\n        x = torch.cat([self.c1(x), self.c2(x), self.c3(x), self.c4(x), self.c5(x), x], dim=1)\n        x = self.c6(x)\n        return x\n</code></pre>\n<p>Hyperparameters</p>\n<ul>\n<li>optimizer: SGD, wd=1e-4, nesterov momentum</li>\n<li>learning rate: 0.1 with cosine annealing</li>\n<li>batch size: 128</li>\n<li>epochs: 40 </li>\n<li>input: 3 channels of raw signal filtered with butterworth filter at 20hz</li>\n<li>augmentations: freq masking, time masking, small shifts, channel shuffle</li>\n</ul>\n<h3>Using synthetic dataset (0.886 public LB)</h3>\n<p>As soon as Denis found a more or less good approach to signal/noise generation we started experimenting with additional data. <br>\nOverall we had 2 million noise samples and 1 million pure signal samples. During training positive sample = random noise sample + random signal sample.</p>\n<p>From these experiments </p>\n<ul>\n<li>mixing synthetic data with the train dataset did not work </li>\n<li>augmentations are actually harmful in this case</li>\n<li>pretraining on synthetic data and fine tuning on the train set works great</li>\n</ul>\n<p>During pretraining stage for simplicity we used the same amount of samples in epoch as in the train set.<br>\nPre-training around 100 epochs and fine tuning 5 folds on the train set gave 8836 on the public LB for the single Conv1d model. <br>\nAs detectors, especially Virgo, have different noise distribution it makes sense to use a separate conv1d encoder for each channel. Split encoders and a linear classifier on top of concatenated features boosted the LB score to <strong>0.8842</strong>.</p>\n<p>It is clear that a fully connected layer is not the best fusion approach for  the model with separate  encoders for each channel. <br>\nThat’s where resnet34 came into play and surprisingly it worked better than other 2D models. We also predicted signal parameters during pretraining  (SNR, chirp mass, Q) which also brought minor improvements. <br>\nPretraining 200 epochs  and fine tuning 5 folds just 1 epoch gives <strong>0.8858</strong> public LB score.<br>\nAugmentations during finetuning or pretraining negatively affected CV, so the best models are without any augmentations and trained with AdamW optimizer.<br>\n<img src=\"https://i.imgur.com/t1X1mL4.png\" alt=\"model\"></p>\n<p>Input to Resnet34 looked the following way (Handford band)</p>\n<p><img src=\"https://i.imgur.com/iMc4Yrp.jpeg\" alt=\"1D features\"></p>\n<h3>Segmentation</h3>\n<p>Binary segmentation using output of Conv1D predicted good masks for strong signals but did not improve recall on weak signals. In general it could be a useful tool to analyse the data, but we did not get any boost on the LB from that.<br>\n<img src=\"https://i.imgur.com/zLPtw6S.png\" alt=\"good signal segmentation\"></p>\n<h3>Things that did not work</h3>\n<p>There were much more experiments that I won't describe (including different frontends, training approaches etc.), but most noticeable are:</p>\n<p><strong>Denoising autoencoder</strong></p>\n<p>I trained different variants of autoencoders to separate noise and signals which worked great for strong signals but produced poor results on medium to low amplitude signals.</p>\n<p><strong>OHEM collapse and reverse labels mystery</strong></p>\n<p>I tried different versions of hard example mining to improve model performance on hard samples but usually the model collapsed and started predicting the same probability for all samples. </p>\n<p>Which led to an interesting experiment:</p>\n<ul>\n<li>from full OOF predictions select positive samples with low probability and negative with high probability</li>\n<li>train on this subset but validate on proper split</li>\n<li>evaluate using predicted probabilities</li>\n<li>evaluate using  reversed predctions (1 - p) </li>\n</ul>\n<p><img src=\"https://i.imgur.com/ZbnpqeV.png\" alt=\"\"></p>\n<p>That result was really confusing and at first we thought that the dataset was mislabeled. Later even with generated synthetic data we had the same problem. <br>\nIt is clear that because of the SNR wall some positive samples can be considered as just noise, but how the model generalized to predict signal from noise samples, that’s what we could not find. </p>",
      "rawMarkdown": "We decided to make two posts to make them more or less  focused and concise. \n DSP part  https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275507\n\n### Baseline solution, simple Conv1D (0.88 public LB)\nAt the very beginning we tried  different inputs:\n\n1. CQT - could not make it past 0.87\n2. Spectrograms - with nnAudio, a bit better\n3. raw signal - much better \n\nAfter applying highpass filter at 20Hz  1D conv  on stacked channels input quickly produced good val scores. Also it was much easier to make experiments as training on a single GPU took less than 1 hour. The model was very simple, just a bunch of conv-bn-silu blocks and maxpools. For convolutions we used kernel sizes (by block) 64- > 32 -> 16 -> 8. \n\nAt that stage for augmentations we used:\n\nchannel shuffle (only Hanford, Livingston)\nminor time shifts between channels, 5ms\n\nThis approach gave us 0.877 public LB score with a single 5fold model.\n\n**SGD is better than AdamW for Conv1D without synthetic data**\n\nEven though model capacity was quite low, the model quickly overfitted, usually after 20 epochs, which took around 30 mins. \nSwitching to SGD with weight decay and nesterov momentum improved LB score to **0.880**. Other Optimizers like AdamW with whatever weight decay, MadGrad gave lower quality.\n\n### Improved Conv1D model (0.883 public LB)\nAt that time the generated synthetic dataset was not good enough and we felt that 1D model could be improved. \nFollowing the Inception V3 approach we added different kernel sizes in each conv block, the starting block was with 32, 64, 128 kernel sizes. \nThis gave us a boost from 0.88 to 0.881 on the LB.\nNext step was to try even more kernel sizes as  it also made sense from DSP theory. Inception like block with 5 different kernels (16, 32, 64, 128, 256) allowed us to get 0.8823 single model LB score. \nA small ensemble improved LB score to **0.883**. \nAdding more kernel sizes did not improve CV/LB scores.\n```\n# Main building block for Conv1D models\nclass ConcatBlockConv5(nn.Module):\n    def __init__(self, in_ch, out_ch, k, act=nn.SiLU):\n        super().__init__()\n        self.c1 = conv_bn_silu_block(in_ch, out_ch, k, act)\n        self.c2 = conv_bn_silu_block(in_ch, out_ch, k * 2, act)\n        self.c3 = conv_bn_silu_block(in_ch, out_ch, k // 2, act)\n        self.c4 = conv_bn_silu_block(in_ch, out_ch, k // 4, act)\n        self.c5 = conv_bn_silu_block(in_ch, out_ch, k * 4, act)\n        self.c6 = conv_bn_silu_block(in_ch * 5 + in_ch, out_ch, 1, act)\n    def forward(self, x):\n        x = torch.cat([self.c1(x), self.c2(x), self.c3(x), self.c4(x), self.c5(x), x], dim=1)\n        x = self.c6(x)\n        return x\n```\n\nHyperparameters\n- optimizer: SGD, wd=1e-4, nesterov momentum\n- learning rate: 0.1 with cosine annealing\n- batch size: 128\n- epochs: 40 \n- input: 3 channels of raw signal filtered with butterworth filter at 20hz\n- augmentations: freq masking, time masking, small shifts, channel shuffle\n\n\n### Using synthetic dataset (0.886 public LB)\nAs soon as Denis found a more or less good approach to signal/noise generation we started experimenting with additional data. \nOverall we had 2 million noise samples and 1 million pure signal samples. During training positive sample = random noise sample + random signal sample.\n\nFrom these experiments \n- mixing synthetic data with the train dataset did not work \n- augmentations are actually harmful in this case\n- pretraining on synthetic data and fine tuning on the train set works great\n\nDuring pretraining stage for simplicity we used the same amount of samples in epoch as in the train set.\nPre-training around 100 epochs and fine tuning 5 folds on the train set gave 8836 on the public LB for the single Conv1d model. \nAs detectors, especially Virgo, have different noise distribution it makes sense to use a separate conv1d encoder for each channel. Split encoders and a linear classifier on top of concatenated features boosted the LB score to **0.8842**.\n\nIt is clear that a fully connected layer is not the best fusion approach for  the model with separate  encoders for each channel. \nThat’s where resnet34 came into play and surprisingly it worked better than other 2D models. We also predicted signal parameters during pretraining  (SNR, chirp mass, Q) which also brought minor improvements. \nPretraining 200 epochs  and fine tuning 5 folds just 1 epoch gives **0.8858** public LB score.\nAugmentations during finetuning or pretraining negatively affected CV, so the best models are without any augmentations and trained with AdamW optimizer.\n![model](https://i.imgur.com/t1X1mL4.png)\n\nInput to Resnet34 looked the following way (Handford band)\n\n![1D features](https://i.imgur.com/iMc4Yrp.jpeg)\n\n\n### Segmentation\nBinary segmentation using output of Conv1D predicted good masks for strong signals but did not improve recall on weak signals. In general it could be a useful tool to analyse the data, but we did not get any boost on the LB from that.\n![good signal segmentation](https://i.imgur.com/zLPtw6S.png)\n\n### Things that did not work \nThere were much more experiments that I won't describe (including different frontends, training approaches etc.), but most noticeable are:\n\n**Denoising autoencoder**\n\nI trained different variants of autoencoders to separate noise and signals which worked great for strong signals but produced poor results on medium to low amplitude signals.\n\n**OHEM collapse and reverse labels mystery**\n\nI tried different versions of hard example mining to improve model performance on hard samples but usually the model collapsed and started predicting the same probability for all samples. \n\nWhich led to an interesting experiment:\n- from full OOF predictions select positive samples with low probability and negative with high probability\n- train on this subset but validate on proper split\n- evaluate using predicted probabilities\n- evaluate using  reversed predctions (1 - p) \n\n![](https://i.imgur.com/ZbnpqeV.png)\n\nThat result was really confusing and at first we thought that the dataset was mislabeled. Later even with generated synthetic data we had the same problem. \nIt is clear that because of the SNR wall some positive samples can be considered as just noise, but how the model generalized to predict signal from noise samples, that’s what we could not find.",
      "votes": null
    },
    {
      "id": "1529732",
      "postDate": "09/30/2021 16:20:57",
      "content": "<p>Solid approach, congrats.<br>\nI tried synthetic data but couldn't score better than 0.865. <br>\nCould you describe how you generated the synthetic data?</p>",
      "rawMarkdown": "Solid approach, congrats.\nI tried synthetic data but couldn't score better than 0.865. \nCould you describe how you generated the synthetic data?",
      "votes": null
    },
    {
      "id": "1529736",
      "postDate": "09/30/2021 16:24:24",
      "content": "<p>Yes it is quite tricky to make it work. <a href=\"https://www.kaggle.com/denisbsu\" target=\"_blank\">@denisbsu</a> is writing a separate post on that right now, will be ready soon  </p>",
      "rawMarkdown": "Yes it is quite tricky to make it work. @denisbsu is writing a separate post on that right now, will be ready soon",
      "votes": null
    },
    {
      "id": "1529739",
      "postDate": "09/30/2021 16:24:57",
      "content": "<p>Indeed, same thing here. We used the synthetic data in the end for pretraining but it only gives a few points gain. Very curious about how they generated and use the synthetic data. </p>",
      "rawMarkdown": "Indeed, same thing here. We used the synthetic data in the end for pretraining but it only gives a few points gain. Very curious about how they generated and use the synthetic data.",
      "votes": null
    },
    {
      "id": "1529743",
      "postDate": "09/30/2021 16:30:19",
      "content": "<p>Data centric approach!! Congrats. So PB score of 0.8864 was from ensemble? If yes, how many and all of them were resnet34?</p>",
      "rawMarkdown": "Data centric approach!! Congrats. So PB score of 0.8864 was from ensemble? If yes, how many and all of them were resnet34?",
      "votes": null
    },
    {
      "id": "1529761",
      "postDate": "09/30/2021 16:38:09",
      "content": "<p>We also trained 20 folds on the same model - it gave 0.8862, but basically anything above is exponentially hard, at that point all models make the same mistakes and ensemble brings very minor boost. <br>\nThey key was not resnet34 but the features from Conv1d.</p>",
      "rawMarkdown": "We also trained 20 folds on the same model - it gave 0.8862, but basically anything above is exponentially hard, at that point all models make the same mistakes and ensemble brings very minor boost. \nThey key was not resnet34 but the features from Conv1d.",
      "votes": null
    },
    {
      "id": "1529799",
      "postDate": "09/30/2021 16:57:35",
      "content": "<p><a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275507\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275507</a></p>",
      "rawMarkdown": "titericz https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275507",
      "votes": null
    },
    {
      "id": "1529812",
      "postDate": "09/30/2021 17:05:26",
      "content": "<p>Very interesting method, congrats!</p>\n<p>It is clear that additional data helped but it looks like you could have won without it.</p>",
      "rawMarkdown": "Very interesting method, congrats!\n\nIt is clear that additional data helped but it looks like you could have won without it.",
      "votes": null
    },
    {
      "id": "1529814",
      "postDate": "09/30/2021 17:07:36",
      "content": "<p>Brilliant solution! Well done, as always)</p>",
      "rawMarkdown": "Brilliant solution! Well done, as always)",
      "votes": null
    },
    {
      "id": "1529818",
      "postDate": "09/30/2021 17:11:24",
      "content": "<p>That would be much harder without big ensembles, on the other hand that would push us to experiment more with Conv1D architecture and augmentations.  </p>",
      "rawMarkdown": "That would be much harder without big ensembles, on the other hand that would push us to experiment more with Conv1D architecture and augmentations.",
      "votes": null
    },
    {
      "id": "1529820",
      "postDate": "09/30/2021 17:13:40",
      "content": "<p>thanks, my 1cent <a href=\"https://www.kaggle.com/titericz/simulated-gw/notebook\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "thanks, my 1cent [here](https://www.kaggle.com/titericz/simulated-gw/notebook)",
      "votes": null
    },
    {
      "id": "1529829",
      "postDate": "09/30/2021 17:18:34",
      "content": "<p>Would love to see you two distill this into a white paper.</p>",
      "rawMarkdown": "Would love to see you two distill this into a white paper.",
      "votes": null
    },
    {
      "id": "1529864",
      "postDate": "09/30/2021 17:50:30",
      "content": "<p><a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a><br>\nI used my custom optimized version of riroriro (266x faster than original in 1 thread, so 1 signal was generated in 0.3 sec, and 3200x faster in 12 threads - kind of real-time) to create synthetic data only for the 2 ligos (without virgo), instead of pycbc  - but I reached the same 0.865 auc as you did - probably for similar reasons.</p>\n<p>My 1d convnet could learn from the raw data from the start without any fft and without any bandpass or other 'human' designed preprocessing, but I started this experiment only 2 days before competition end :(</p>\n<p>Maybe I can use it for a possible next G2Net competition :)</p>",
      "rawMarkdown": "titericz\nI used my custom optimized version of riroriro (266x faster than original in 1 thread, so 1 signal was generated in 0.3 sec, and 3200x faster in 12 threads - kind of real-time) to create synthetic data only for the 2 ligos (without virgo), instead of pycbc  - but I reached the same 0.865 auc as you did - probably for similar reasons.\n\nMy 1d convnet could learn from the raw data from the start without any fft and without any bandpass or other 'human' designed preprocessing, but I started this experiment only 2 days before competition end :(\n\nMaybe I can use it for a possible next G2Net competition :)",
      "votes": null
    },
    {
      "id": "1529949",
      "postDate": "09/30/2021 19:37:06",
      "content": "<p>thank you for the write-up.<br>\nit is a brilliant solution.</p>\n<p>\"Input to Resnet34 looked the following way (Handford band)\"<br>\nHow are the features sorted? e.g. by kernel_size (64,32,16,8) or random or by frequency response?</p>\n<p>i think the frequency response is related to the length. that is why you can see the banana shape for the chirp wave in the segmentation diagram?</p>",
      "rawMarkdown": "thank you for the write-up.\nit is a brilliant solution.\n\n\"Input to Resnet34 looked the following way (Handford band)\"\nHow are the features sorted? e.g. by kernel_size (64,32,16,8) or random or by frequency response?\n\ni think the frequency response is related to the length. that is why you can see the banana shape for the chirp wave in the segmentation diagram?",
      "votes": null
    },
    {
      "id": "1529967",
      "postDate": "09/30/2021 19:54:54",
      "content": "<p>Surprisingly random is working just fine, just hoped that model learned the right represenation needed for classification :)<br>\nIf we add segmentation mask as an additional target then it learns sorting implicitly. For example you can see the effect of notch filter at 300Hz on the segmentation mask in the main post. <br>\nAlso we removed low freq signal part from the masks because signals are attenuated at low frequencies due to filtering. <br>\nSo in general model learns to detect something that is similar to GW,  visualisation of logits clear shows that. Unfortunately it doesn't matter for the CV/LB scores. <br>\n<img src=\"https://i.imgur.com/IP9Qspo.png\" alt=\"\"><br>\nSpectrograms of GW look similar to the masks just with low frquency part.</p>",
      "rawMarkdown": "Surprisingly random is working just fine, just hoped that model learned the right represenation needed for classification :)\nIf we add segmentation mask as an additional target then it learns sorting implicitly. For example you can see the effect of notch filter at 300Hz on the segmentation mask in the main post. \nAlso we removed low freq signal part from the masks because signals are attenuated at low frequencies due to filtering. \nSo in general model learns to detect something that is similar to GW,  visualisation of logits clear shows that. Unfortunately it doesn't matter for the CV/LB scores. \n![](https://i.imgur.com/IP9Qspo.png)\nSpectrograms of GW look similar to the masks just with low frquency part.",
      "votes": null
    },
    {
      "id": "1529970",
      "postDate": "09/30/2021 20:01:36",
      "content": "<p>writing a good paper would take more time than the challenge itself😧</p>",
      "rawMarkdown": "writing a good paper would take more time than the challenge itself😧",
      "votes": null
    },
    {
      "id": "1529983",
      "postDate": "09/30/2021 20:19:10",
      "content": "<p>\"Surprisingly random is working just fine,\"<br>\n\"That’s where resnet34 came into play and surprisingly it worked better than other 2D models\"</p>\n<p>so the resnet34 has to enforce some spatial consistency in the stacked response image in its own ways (e.g. 2 nearby 1dcnn encoders should activate at the same location in the stack image). we should see some line-like structure (the banana shape is deformed and dismantled). max pooling on the stacked response image would max it more obvious (different 1d features is activated for different wave) </p>\n<p>This is probably a simple convolution 2dCNN is better.</p>\n<p>\"If we add segmentation mask as an additional target then it learns sorting implicitly.\"<br>\nthis is basically a modern-day matched filter. Basically, we can trace out the ridge and recover the wavelet parameters.</p>\n<hr>\n<p>as a side note, I wonder what happens if you use your 1dCNN encoder+2dCNNdecoder on other data.<br>\nHow would the 2d decoder image look like?  what if we use it on NLP?</p>",
      "rawMarkdown": "\"Surprisingly random is working just fine,\"\n\"That’s where resnet34 came into play and surprisingly it worked better than other 2D models\"\n\nso the resnet34 has to enforce some spatial consistency in the stacked response image in its own ways (e.g. 2 nearby 1dcnn encoders should activate at the same location in the stack image). we should see some line-like structure (the banana shape is deformed and dismantled). max pooling on the stacked response image would max it more obvious (different 1d features is activated for different wave) \n\nThis is probably a simple convolution 2dCNN is better.\n\n\"If we add segmentation mask as an additional target then it learns sorting implicitly.\"\nthis is basically a modern-day matched filter. Basically, we can trace out the ridge and recover the wavelet parameters.\n\n---\n\n as a side note, I wonder what happens if you use your 1dCNN encoder+2dCNNdecoder on other data.\nHow would the 2d decoder image look like?  what if we use it on NLP?",
      "votes": null
    },
    {
      "id": "1530005",
      "postDate": "09/30/2021 20:58:35",
      "content": "<blockquote>\n  <p>what if we use it on NLP?</p>\n</blockquote>\n<p>ufortunately last time I seriously worked with NLP data was  when GloVe and word2vec + LSTM were state of the art approaches 😄</p>\n<blockquote>\n  <p>How would the 2d decoder image look like?</p>\n</blockquote>\n<p>Have not done any experiments with UNet here though I doubt that it would bring any improvements to the score. The limit here is how well Conv1D can handle signals with low SNR. <br>\nWhile playing with synthetic data we saw a lot of examples when model could properly  predict <code>1.2* signal + noise</code> as 1. with high probability and the same <code>singal + noise</code> with close to 0 probability. There was a clear threshold (SNR wall) where model could not differentiate GW from noise.</p>",
      "rawMarkdown": "> what if we use it on NLP?\n\nufortunately last time I seriously worked with NLP data was  when GloVe and word2vec + LSTM were state of the art approaches 😄\n\n> How would the 2d decoder image look like?\n\nHave not done any experiments with UNet here though I doubt that it would bring any improvements to the score. The limit here is how well Conv1D can handle signals with low SNR. \nWhile playing with synthetic data we saw a lot of examples when model could properly  predict `1.2* signal + noise` as 1. with high probability and the same `singal + noise` with close to 0 probability. There was a clear threshold (SNR wall) where model could not differentiate GW from noise.",
      "votes": null
    },
    {
      "id": "1530069",
      "postDate": "09/30/2021 22:42:32",
      "content": "<p>Beautiful solution, congratulations!</p>",
      "rawMarkdown": "Beautiful solution, congratulations!",
      "votes": null
    },
    {
      "id": "1530514",
      "postDate": "10/01/2021 07:51:53",
      "content": "<p>\"OHEM collapse and reverse labels mystery\"</p>\n<p>check my post related to this (comparison of zebra augmentation with others):<br>\n<a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275335\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275335</a></p>",
      "rawMarkdown": "\"OHEM collapse and reverse labels mystery\"\n\ncheck my post related to this (comparison of zebra augmentation with others):\nhttps://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275335",
      "votes": null
    },
    {
      "id": "1531524",
      "postDate": "10/02/2021 05:17:08",
      "content": "<p>Wow. Beautiful solution indeed. Congratulations 👏🎉</p>",
      "rawMarkdown": "Wow. Beautiful solution indeed. Congratulations 👏🎉",
      "votes": null
    },
    {
      "id": "1531738",
      "postDate": "10/02/2021 10:10:17",
      "content": "<p>Such a wonderful solution!</p>",
      "rawMarkdown": "Such a wonderful solution!",
      "votes": null
    },
    {
      "id": "1532617",
      "postDate": "10/03/2021 07:43:01",
      "content": "<p>Awesome, great to see new stuff, well posted <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a> </p>",
      "rawMarkdown": "Awesome, great to see new stuff, well posted @selimsef",
      "votes": null
    },
    {
      "id": "1533070",
      "postDate": "10/03/2021 16:22:28",
      "content": "<p>Wow, this is awesome </p>",
      "rawMarkdown": "Wow, this is awesome",
      "votes": null
    },
    {
      "id": "1559907",
      "postDate": "10/27/2021 08:07:58",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    },
    {
      "id": "1680231",
      "postDate": "02/07/2022 17:03:02",
      "content": "<p>hello, it' there is a github for the solution with the deep learning as there is for DSP part ? <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a></p>",
      "rawMarkdown": "hello, it' there is a github for the solution with the deep learning as there is for DSP part ? @selimsef",
      "votes": null
    },
    {
      "id": "1691527",
      "postDate": "02/15/2022 13:09:03",
      "content": "<p>Unfortunately no, I have not uploaded this one to github</p>",
      "rawMarkdown": "Unfortunately no, I have not uploaded this one to github",
      "votes": null
    },
    {
      "id": "2504103",
      "postDate": "10/29/2023 17:04:28",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a> , is your whole code of this problem publically available? where one can get your complete solution for this competition ? </p>",
      "rawMarkdown": "Hi @selimsef , is your whole code of this problem publically available? where one can get your complete solution for this competition ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1529732,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "09/30/2021 16:20:57",
      "content": "<p>Solid approach, congrats.<br>\nI tried synthetic data but couldn't score better than 0.865. <br>\nCould you describe how you generated the synthetic data?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1529736,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "09/30/2021 16:24:24",
          "content": "<p>Yes it is quite tricky to make it work. <a href=\"https://www.kaggle.com/denisbsu\" target=\"_blank\">@denisbsu</a> is writing a separate post on that right now, will be ready soon  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1529739,
          "author_name": "vincentwang25",
          "author_url": "",
          "post_date": "09/30/2021 16:24:57",
          "content": "<p>Indeed, same thing here. We used the synthetic data in the end for pretraining but it only gives a few points gain. Very curious about how they generated and use the synthetic data. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1529799,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "09/30/2021 16:57:35",
          "content": "<p><a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275507\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275507</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1529820,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "09/30/2021 17:13:40",
          "content": "<p>thanks, my 1cent <a href=\"https://www.kaggle.com/titericz/simulated-gw/notebook\" target=\"_blank\">here</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1529864,
          "author_name": "killimi",
          "author_url": "",
          "post_date": "09/30/2021 17:50:30",
          "content": "<p><a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a><br>\nI used my custom optimized version of riroriro (266x faster than original in 1 thread, so 1 signal was generated in 0.3 sec, and 3200x faster in 12 threads - kind of real-time) to create synthetic data only for the 2 ligos (without virgo), instead of pycbc  - but I reached the same 0.865 auc as you did - probably for similar reasons.</p>\n<p>My 1d convnet could learn from the raw data from the start without any fft and without any bandpass or other 'human' designed preprocessing, but I started this experiment only 2 days before competition end :(</p>\n<p>Maybe I can use it for a possible next G2Net competition :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1529743,
      "author_name": "bibek777",
      "author_url": "",
      "post_date": "09/30/2021 16:30:19",
      "content": "<p>Data centric approach!! Congrats. So PB score of 0.8864 was from ensemble? If yes, how many and all of them were resnet34?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1529761,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "09/30/2021 16:38:09",
          "content": "<p>We also trained 20 folds on the same model - it gave 0.8862, but basically anything above is exponentially hard, at that point all models make the same mistakes and ensemble brings very minor boost. <br>\nThey key was not resnet34 but the features from Conv1d.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1529812,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "09/30/2021 17:05:26",
      "content": "<p>Very interesting method, congrats!</p>\n<p>It is clear that additional data helped but it looks like you could have won without it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1529818,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "09/30/2021 17:11:24",
          "content": "<p>That would be much harder without big ensembles, on the other hand that would push us to experiment more with Conv1D architecture and augmentations.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1529814,
      "author_name": "trytolose",
      "author_url": "",
      "post_date": "09/30/2021 17:07:36",
      "content": "<p>Brilliant solution! Well done, as always)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1529829,
      "author_name": "authman",
      "author_url": "",
      "post_date": "09/30/2021 17:18:34",
      "content": "<p>Would love to see you two distill this into a white paper.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1529970,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "09/30/2021 20:01:36",
          "content": "<p>writing a good paper would take more time than the challenge itself😧</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1529949,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/30/2021 19:37:06",
      "content": "<p>thank you for the write-up.<br>\nit is a brilliant solution.</p>\n<p>\"Input to Resnet34 looked the following way (Handford band)\"<br>\nHow are the features sorted? e.g. by kernel_size (64,32,16,8) or random or by frequency response?</p>\n<p>i think the frequency response is related to the length. that is why you can see the banana shape for the chirp wave in the segmentation diagram?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1529967,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "09/30/2021 19:54:54",
          "content": "<p>Surprisingly random is working just fine, just hoped that model learned the right represenation needed for classification :)<br>\nIf we add segmentation mask as an additional target then it learns sorting implicitly. For example you can see the effect of notch filter at 300Hz on the segmentation mask in the main post. <br>\nAlso we removed low freq signal part from the masks because signals are attenuated at low frequencies due to filtering. <br>\nSo in general model learns to detect something that is similar to GW,  visualisation of logits clear shows that. Unfortunately it doesn't matter for the CV/LB scores. <br>\n<img src=\"https://i.imgur.com/IP9Qspo.png\" alt=\"\"><br>\nSpectrograms of GW look similar to the masks just with low frquency part.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1529983,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/30/2021 20:19:10",
          "content": "<p>\"Surprisingly random is working just fine,\"<br>\n\"That’s where resnet34 came into play and surprisingly it worked better than other 2D models\"</p>\n<p>so the resnet34 has to enforce some spatial consistency in the stacked response image in its own ways (e.g. 2 nearby 1dcnn encoders should activate at the same location in the stack image). we should see some line-like structure (the banana shape is deformed and dismantled). max pooling on the stacked response image would max it more obvious (different 1d features is activated for different wave) </p>\n<p>This is probably a simple convolution 2dCNN is better.</p>\n<p>\"If we add segmentation mask as an additional target then it learns sorting implicitly.\"<br>\nthis is basically a modern-day matched filter. Basically, we can trace out the ridge and recover the wavelet parameters.</p>\n<hr>\n<p>as a side note, I wonder what happens if you use your 1dCNN encoder+2dCNNdecoder on other data.<br>\nHow would the 2d decoder image look like?  what if we use it on NLP?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1530005,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "09/30/2021 20:58:35",
          "content": "<blockquote>\n  <p>what if we use it on NLP?</p>\n</blockquote>\n<p>ufortunately last time I seriously worked with NLP data was  when GloVe and word2vec + LSTM were state of the art approaches 😄</p>\n<blockquote>\n  <p>How would the 2d decoder image look like?</p>\n</blockquote>\n<p>Have not done any experiments with UNet here though I doubt that it would bring any improvements to the score. The limit here is how well Conv1D can handle signals with low SNR. <br>\nWhile playing with synthetic data we saw a lot of examples when model could properly  predict <code>1.2* signal + noise</code> as 1. with high probability and the same <code>singal + noise</code> with close to 0 probability. There was a clear threshold (SNR wall) where model could not differentiate GW from noise.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1530069,
      "author_name": "analokamus",
      "author_url": "",
      "post_date": "09/30/2021 22:42:32",
      "content": "<p>Beautiful solution, congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1530514,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/01/2021 07:51:53",
      "content": "<p>\"OHEM collapse and reverse labels mystery\"</p>\n<p>check my post related to this (comparison of zebra augmentation with others):<br>\n<a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275335\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275335</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1531524,
      "author_name": "kalilurrahman",
      "author_url": "",
      "post_date": "10/02/2021 05:17:08",
      "content": "<p>Wow. Beautiful solution indeed. Congratulations 👏🎉</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1531738,
      "author_name": "shashankrathi",
      "author_url": "",
      "post_date": "10/02/2021 10:10:17",
      "content": "<p>Such a wonderful solution!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1532617,
      "author_name": "",
      "author_url": "",
      "post_date": "10/03/2021 07:43:01",
      "content": "<p>Awesome, great to see new stuff, well posted <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1533070,
      "author_name": "anandhuh",
      "author_url": "",
      "post_date": "10/03/2021 16:22:28",
      "content": "<p>Wow, this is awesome </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1559907,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 08:07:58",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1680231,
      "author_name": "patrickmichel",
      "author_url": "",
      "post_date": "02/07/2022 17:03:02",
      "content": "<p>hello, it' there is a github for the solution with the deep learning as there is for DSP part ? <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1691527,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "02/15/2022 13:09:03",
          "content": "<p>Unfortunately no, I have not uploaded this one to github</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2504103,
      "author_name": "astroboysanju",
      "author_url": "",
      "post_date": "10/29/2023 17:04:28",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a> , is your whole code of this problem publically available? where one can get your complete solution for this competition ? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1529717": "We decided to make two posts to make them more or less  focused and concise. \n DSP part  https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275507\n\n### Baseline solution, simple Conv1D (0.88 public LB)\nAt the very beginning we tried  different inputs:\n\n1. CQT - could not make it past 0.87\n2. Spectrograms - with nnAudio, a bit better\n3. raw signal - much better \n\nAfter applying highpass filter at 20Hz  1D conv  on stacked channels input quickly produced good val scores. Also it was much easier to make experiments as training on a single GPU took less than 1 hour. The model was very simple, just a bunch of conv-bn-silu blocks and maxpools. For convolutions we used kernel sizes (by block) 64- > 32 -> 16 -> 8. \n\nAt that stage for augmentations we used:\n\nchannel shuffle (only Hanford, Livingston)\nminor time shifts between channels, 5ms\n\nThis approach gave us 0.877 public LB score with a single 5fold model.\n\n**SGD is better than AdamW for Conv1D without synthetic data**\n\nEven though model capacity was quite low, the model quickly overfitted, usually after 20 epochs, which took around 30 mins. \nSwitching to SGD with weight decay and nesterov momentum improved LB score to **0.880**. Other Optimizers like AdamW with whatever weight decay, MadGrad gave lower quality.\n\n### Improved Conv1D model (0.883 public LB)\nAt that time the generated synthetic dataset was not good enough and we felt that 1D model could be improved. \nFollowing the Inception V3 approach we added different kernel sizes in each conv block, the starting block was with 32, 64, 128 kernel sizes. \nThis gave us a boost from 0.88 to 0.881 on the LB.\nNext step was to try even more kernel sizes as  it also made sense from DSP theory. Inception like block with 5 different kernels (16, 32, 64, 128, 256) allowed us to get 0.8823 single model LB score. \nA small ensemble improved LB score to **0.883**. \nAdding more kernel sizes did not improve CV/LB scores.\n```\n# Main building block for Conv1D models\nclass ConcatBlockConv5(nn.Module):\n    def __init__(self, in_ch, out_ch, k, act=nn.SiLU):\n        super().__init__()\n        self.c1 = conv_bn_silu_block(in_ch, out_ch, k, act)\n        self.c2 = conv_bn_silu_block(in_ch, out_ch, k * 2, act)\n        self.c3 = conv_bn_silu_block(in_ch, out_ch, k // 2, act)\n        self.c4 = conv_bn_silu_block(in_ch, out_ch, k // 4, act)\n        self.c5 = conv_bn_silu_block(in_ch, out_ch, k * 4, act)\n        self.c6 = conv_bn_silu_block(in_ch * 5 + in_ch, out_ch, 1, act)\n    def forward(self, x):\n        x = torch.cat([self.c1(x), self.c2(x), self.c3(x), self.c4(x), self.c5(x), x], dim=1)\n        x = self.c6(x)\n        return x\n```\n\nHyperparameters\n- optimizer: SGD, wd=1e-4, nesterov momentum\n- learning rate: 0.1 with cosine annealing\n- batch size: 128\n- epochs: 40 \n- input: 3 channels of raw signal filtered with butterworth filter at 20hz\n- augmentations: freq masking, time masking, small shifts, channel shuffle\n\n\n### Using synthetic dataset (0.886 public LB)\nAs soon as Denis found a more or less good approach to signal/noise generation we started experimenting with additional data. \nOverall we had 2 million noise samples and 1 million pure signal samples. During training positive sample = random noise sample + random signal sample.\n\nFrom these experiments \n- mixing synthetic data with the train dataset did not work \n- augmentations are actually harmful in this case\n- pretraining on synthetic data and fine tuning on the train set works great\n\nDuring pretraining stage for simplicity we used the same amount of samples in epoch as in the train set.\nPre-training around 100 epochs and fine tuning 5 folds on the train set gave 8836 on the public LB for the single Conv1d model. \nAs detectors, especially Virgo, have different noise distribution it makes sense to use a separate conv1d encoder for each channel. Split encoders and a linear classifier on top of concatenated features boosted the LB score to **0.8842**.\n\nIt is clear that a fully connected layer is not the best fusion approach for  the model with separate  encoders for each channel. \nThat’s where resnet34 came into play and surprisingly it worked better than other 2D models. We also predicted signal parameters during pretraining  (SNR, chirp mass, Q) which also brought minor improvements. \nPretraining 200 epochs  and fine tuning 5 folds just 1 epoch gives **0.8858** public LB score.\nAugmentations during finetuning or pretraining negatively affected CV, so the best models are without any augmentations and trained with AdamW optimizer.\n![model](https://i.imgur.com/t1X1mL4.png)\n\nInput to Resnet34 looked the following way (Handford band)\n\n![1D features](https://i.imgur.com/iMc4Yrp.jpeg)\n\n\n### Segmentation\nBinary segmentation using output of Conv1D predicted good masks for strong signals but did not improve recall on weak signals. In general it could be a useful tool to analyse the data, but we did not get any boost on the LB from that.\n![good signal segmentation](https://i.imgur.com/zLPtw6S.png)\n\n### Things that did not work \nThere were much more experiments that I won't describe (including different frontends, training approaches etc.), but most noticeable are:\n\n**Denoising autoencoder**\n\nI trained different variants of autoencoders to separate noise and signals which worked great for strong signals but produced poor results on medium to low amplitude signals.\n\n**OHEM collapse and reverse labels mystery**\n\nI tried different versions of hard example mining to improve model performance on hard samples but usually the model collapsed and started predicting the same probability for all samples. \n\nWhich led to an interesting experiment:\n- from full OOF predictions select positive samples with low probability and negative with high probability\n- train on this subset but validate on proper split\n- evaluate using predicted probabilities\n- evaluate using  reversed predctions (1 - p) \n\n![](https://i.imgur.com/ZbnpqeV.png)\n\nThat result was really confusing and at first we thought that the dataset was mislabeled. Later even with generated synthetic data we had the same problem. \nIt is clear that because of the SNR wall some positive samples can be considered as just noise, but how the model generalized to predict signal from noise samples, that’s what we could not find.",
    "1529732": "Solid approach, congrats.\nI tried synthetic data but couldn't score better than 0.865. \nCould you describe how you generated the synthetic data?",
    "1529736": "Yes it is quite tricky to make it work. @denisbsu is writing a separate post on that right now, will be ready soon",
    "1529739": "Indeed, same thing here. We used the synthetic data in the end for pretraining but it only gives a few points gain. Very curious about how they generated and use the synthetic data.",
    "1529743": "Data centric approach!! Congrats. So PB score of 0.8864 was from ensemble? If yes, how many and all of them were resnet34?",
    "1529761": "We also trained 20 folds on the same model - it gave 0.8862, but basically anything above is exponentially hard, at that point all models make the same mistakes and ensemble brings very minor boost. \nThey key was not resnet34 but the features from Conv1d.",
    "1529799": "titericz https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275507",
    "1529812": "Very interesting method, congrats!\n\nIt is clear that additional data helped but it looks like you could have won without it.",
    "1529814": "Brilliant solution! Well done, as always)",
    "1529818": "That would be much harder without big ensembles, on the other hand that would push us to experiment more with Conv1D architecture and augmentations.",
    "1529820": "thanks, my 1cent [here](https://www.kaggle.com/titericz/simulated-gw/notebook)",
    "1529829": "Would love to see you two distill this into a white paper.",
    "1529864": "titericz\nI used my custom optimized version of riroriro (266x faster than original in 1 thread, so 1 signal was generated in 0.3 sec, and 3200x faster in 12 threads - kind of real-time) to create synthetic data only for the 2 ligos (without virgo), instead of pycbc  - but I reached the same 0.865 auc as you did - probably for similar reasons.\n\nMy 1d convnet could learn from the raw data from the start without any fft and without any bandpass or other 'human' designed preprocessing, but I started this experiment only 2 days before competition end :(\n\nMaybe I can use it for a possible next G2Net competition :)",
    "1529949": "thank you for the write-up.\nit is a brilliant solution.\n\n\"Input to Resnet34 looked the following way (Handford band)\"\nHow are the features sorted? e.g. by kernel_size (64,32,16,8) or random or by frequency response?\n\ni think the frequency response is related to the length. that is why you can see the banana shape for the chirp wave in the segmentation diagram?",
    "1529967": "Surprisingly random is working just fine, just hoped that model learned the right represenation needed for classification :)\nIf we add segmentation mask as an additional target then it learns sorting implicitly. For example you can see the effect of notch filter at 300Hz on the segmentation mask in the main post. \nAlso we removed low freq signal part from the masks because signals are attenuated at low frequencies due to filtering. \nSo in general model learns to detect something that is similar to GW,  visualisation of logits clear shows that. Unfortunately it doesn't matter for the CV/LB scores. \n![](https://i.imgur.com/IP9Qspo.png)\nSpectrograms of GW look similar to the masks just with low frquency part.",
    "1529970": "writing a good paper would take more time than the challenge itself😧",
    "1529983": "\"Surprisingly random is working just fine,\"\n\"That’s where resnet34 came into play and surprisingly it worked better than other 2D models\"\n\nso the resnet34 has to enforce some spatial consistency in the stacked response image in its own ways (e.g. 2 nearby 1dcnn encoders should activate at the same location in the stack image). we should see some line-like structure (the banana shape is deformed and dismantled). max pooling on the stacked response image would max it more obvious (different 1d features is activated for different wave) \n\nThis is probably a simple convolution 2dCNN is better.\n\n\"If we add segmentation mask as an additional target then it learns sorting implicitly.\"\nthis is basically a modern-day matched filter. Basically, we can trace out the ridge and recover the wavelet parameters.\n\n---\n\n as a side note, I wonder what happens if you use your 1dCNN encoder+2dCNNdecoder on other data.\nHow would the 2d decoder image look like?  what if we use it on NLP?",
    "1530005": "> what if we use it on NLP?\n\nufortunately last time I seriously worked with NLP data was  when GloVe and word2vec + LSTM were state of the art approaches 😄\n\n> How would the 2d decoder image look like?\n\nHave not done any experiments with UNet here though I doubt that it would bring any improvements to the score. The limit here is how well Conv1D can handle signals with low SNR. \nWhile playing with synthetic data we saw a lot of examples when model could properly  predict `1.2* signal + noise` as 1. with high probability and the same `singal + noise` with close to 0 probability. There was a clear threshold (SNR wall) where model could not differentiate GW from noise.",
    "1530069": "Beautiful solution, congratulations!",
    "1530514": "\"OHEM collapse and reverse labels mystery\"\n\ncheck my post related to this (comparison of zebra augmentation with others):\nhttps://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275335",
    "1531524": "Wow. Beautiful solution indeed. Congratulations 👏🎉",
    "1531738": "Such a wonderful solution!",
    "1532617": "Awesome, great to see new stuff, well posted @selimsef",
    "1533070": "Wow, this is awesome",
    "1559907": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
    "1680231": "hello, it' there is a github for the solution with the deep learning as there is for DSP part ? @selimsef",
    "1691527": "Unfortunately no, I have not uploaded this one to github",
    "2504103": "Hi @selimsef , is your whole code of this problem publically available? where one can get your complete solution for this competition ?"
  },
  "source": "meta"
}