{
  "id": 275728,
  "title": "23rd place solution",
  "url": "/competitions/g2net-gravitational-wave-detection/writeups/afterparty-23rd-place-solution",
  "author_name": "",
  "post_date": "2021-10-01T12:34:28.207513Z",
  "votes": 10,
  "comment_count": 2,
  "views": 0,
  "content": "<p>At first, I would like to thank my team members <a href=\"https://www.kaggle.com/vladimirsydor\" target=\"_blank\">@vladimirsydor</a>, <a href=\"https://www.kaggle.com/zekamrozek\" target=\"_blank\">@zekamrozek</a>, <a href=\"https://www.kaggle.com/yakuben\" target=\"_blank\">@yakuben</a> and <a href=\"https://www.kaggle.com/uulott\" target=\"_blank\">@uulott</a> for their work and organizers and the Kaggle community for the amazing experience we had during this competition.</p>\n<h2>TL;DR</h2>\n<p>Blend of 6 2D CNN + 3 1D CNN: 1x2D EfficentNet_b0 + 3 x _b3 + 2 x _b5 + 2 x 1D ResNet with transformer + 1D basic CNN</p>\n<h2>Details</h2>\n<p><strong>Bandpass</strong> was crucial for this data, we experimented with it a lot and tried different setups, those parameters showed better results then others we tried:<br>\n                <code>fmin=20</code>,<br>\n                <code>fmax=1024</code> or <code>fmax=500</code> or <code>fmax=600</code> depending on the model,<br>\n                <code>hop_length=4</code>,<br>\n                <code>bins_per_octave=12</code>,<br>\n                <code>filter_scale=0.5</code>,<br>\n                <code>pad=10</code>.</p>\n<p>For images, we used only CQT and didn't experiment with CWT.</p>\n<p><strong>Augmentations</strong> that worked for us:</p>\n<ol>\n<li>Wave augs:</li>\n</ol>\n<ul>\n<li>Random shift channel +-30;</li>\n<li>Hard MixUp of waves: <code>new_wave = (wave1 + wave2) / 2</code>  and <code>new_target = max(target1, target2)</code>;</li>\n<li>Adding Gaussian Noise with small std.</li>\n</ul>\n<p><strong>Models</strong><br>\nFor most of the competition, we were training EfficentNets, until around 1-2 last weeks we started experimenting with 1D models.</p>\n<p>We found out, that increasing the input image would increase the performance even for EfficentNet_B0, so we experimented with it more. For us best result was when we used <code>stride=(1,2)</code> instead of default <code>(2,2)</code> with an unchanged shape of input. That allowed model to \"stretch\" the input itself across the frequency axis and give an effect of feeding higher resolution images to the network.</p>\n<p>Our 2D models suffered from spoiled gradients in BatchNorm layers on first epochs, so to overcome that we used 2 procedures.  The first idea that worked was to substitute BatchNorm for InstanceNorm throughout the whole model.  2nd  idea that worked was that we would freeze all encoder layers but BatchNorm, and train only them plus Linear head for the first 2 epochs.</p>\n<p>As was mentioned, we started working late on 1D models, mainly simple stacking Conv layers, but after merging with <a href=\"https://www.kaggle.com/uulott\" target=\"_blank\">@uulott</a>, he brought a really promising 1D ResNet model that performed on similar level as our best 2D ones (~0.879 LB).</p>\n<p>Finally, all of our models benefited when we started finetuning them on a small learning rate for multiple epochs with augmentations turned off. That boosted us past 0.88 LB and was keeping us in silver as we kept going up LB the more models we fine-tune and blend with.</p>\n<p><strong>Ensemble</strong><br>\nFor ensembling, we used straightforward hyperopt blending. Adding Any 1D model boosted all of our blends.</p>\n<p><strong>The final blend</strong> consisted of 6 2D-CNNs and 3 1D-CNNs:</p>\n<ol>\n<li><p>EfficentNet-B5, successful augs, freeze 2 epochs, normalize after CQT; </p></li>\n<li><p>Same B5, but fine-tuned;</p></li>\n<li><p>fine-tuned EfficentNet-B3, all successful augs, freeze 2 epochs, normalize after CQT;</p></li>\n<li><p>fine-tuned EfficentNet-B3, initial stride=(1,2)  GaussNoise aug, freeze 2 epochs, normalize after CQT;</p></li>\n<li><p>fine-tuned EfficentNet-B3, initial stride=(1,2)  GaussNoise aug, freeze 2 epochs, normalize after CQT;</p></li>\n<li><p>fine-tuned EfficentNet-B0 , InstanceNorm instead of BatchNorm, normalize after CQT;</p></li>\n<li><p>1D ResNet + Transformer head,  filter order 8 and CosineAnnealing;</p></li>\n<li><p>1D ResNet + Transformer head,  filter order 5 and ReduceLROnPlateau;</p></li>\n<li><p>fine-tuned 6-layer CNN1d with MLP head, MixUp and GaussNoise, linear warm-up and cosine decay.</p></li>\n</ol>\n<p><strong>P.S</strong>. Other things we tried, but it didn't make it to the final blend:</p>\n<ul>\n<li>drop confident 1s before fine-tuning;</li>\n<li>InceptionTime for 1D + Transformer head;</li>\n<li>8 layer custom 1D CNN;</li>\n<li>rexnet_100;</li>\n<li>EfficeintNet-B2 and other;</li>\n<li>resize CQT image into (512,512) or (256, 256).</li>\n</ul>\n<p>Augmentations that we stopped using :</p>\n<ul>\n<li>Channels shuffle for an image in 2D models;</li>\n<li>Horizontal and vertical flip of images;</li>\n<li>normal MixUp.</li>\n</ul>",
  "messages": [
    {
      "id": "1530787",
      "postDate": "10/01/2021 12:34:28",
      "content": "<p>At first, I would like to thank my team members <a href=\"https://www.kaggle.com/vladimirsydor\" target=\"_blank\">@vladimirsydor</a>, <a href=\"https://www.kaggle.com/zekamrozek\" target=\"_blank\">@zekamrozek</a>, <a href=\"https://www.kaggle.com/yakuben\" target=\"_blank\">@yakuben</a> and <a href=\"https://www.kaggle.com/uulott\" target=\"_blank\">@uulott</a> for their work and organizers and the Kaggle community for the amazing experience we had during this competition.</p>\n<h2>TL;DR</h2>\n<p>Blend of 6 2D CNN + 3 1D CNN: 1x2D EfficentNet_b0 + 3 x _b3 + 2 x _b5 + 2 x 1D ResNet with transformer + 1D basic CNN</p>\n<h2>Details</h2>\n<p><strong>Bandpass</strong> was crucial for this data, we experimented with it a lot and tried different setups, those parameters showed better results then others we tried:<br>\n                <code>fmin=20</code>,<br>\n                <code>fmax=1024</code> or <code>fmax=500</code> or <code>fmax=600</code> depending on the model,<br>\n                <code>hop_length=4</code>,<br>\n                <code>bins_per_octave=12</code>,<br>\n                <code>filter_scale=0.5</code>,<br>\n                <code>pad=10</code>.</p>\n<p>For images, we used only CQT and didn't experiment with CWT.</p>\n<p><strong>Augmentations</strong> that worked for us:</p>\n<ol>\n<li>Wave augs:</li>\n</ol>\n<ul>\n<li>Random shift channel +-30;</li>\n<li>Hard MixUp of waves: <code>new_wave = (wave1 + wave2) / 2</code>  and <code>new_target = max(target1, target2)</code>;</li>\n<li>Adding Gaussian Noise with small std.</li>\n</ul>\n<p><strong>Models</strong><br>\nFor most of the competition, we were training EfficentNets, until around 1-2 last weeks we started experimenting with 1D models.</p>\n<p>We found out, that increasing the input image would increase the performance even for EfficentNet_B0, so we experimented with it more. For us best result was when we used <code>stride=(1,2)</code> instead of default <code>(2,2)</code> with an unchanged shape of input. That allowed model to \"stretch\" the input itself across the frequency axis and give an effect of feeding higher resolution images to the network.</p>\n<p>Our 2D models suffered from spoiled gradients in BatchNorm layers on first epochs, so to overcome that we used 2 procedures.  The first idea that worked was to substitute BatchNorm for InstanceNorm throughout the whole model.  2nd  idea that worked was that we would freeze all encoder layers but BatchNorm, and train only them plus Linear head for the first 2 epochs.</p>\n<p>As was mentioned, we started working late on 1D models, mainly simple stacking Conv layers, but after merging with <a href=\"https://www.kaggle.com/uulott\" target=\"_blank\">@uulott</a>, he brought a really promising 1D ResNet model that performed on similar level as our best 2D ones (~0.879 LB).</p>\n<p>Finally, all of our models benefited when we started finetuning them on a small learning rate for multiple epochs with augmentations turned off. That boosted us past 0.88 LB and was keeping us in silver as we kept going up LB the more models we fine-tune and blend with.</p>\n<p><strong>Ensemble</strong><br>\nFor ensembling, we used straightforward hyperopt blending. Adding Any 1D model boosted all of our blends.</p>\n<p><strong>The final blend</strong> consisted of 6 2D-CNNs and 3 1D-CNNs:</p>\n<ol>\n<li><p>EfficentNet-B5, successful augs, freeze 2 epochs, normalize after CQT; </p></li>\n<li><p>Same B5, but fine-tuned;</p></li>\n<li><p>fine-tuned EfficentNet-B3, all successful augs, freeze 2 epochs, normalize after CQT;</p></li>\n<li><p>fine-tuned EfficentNet-B3, initial stride=(1,2)  GaussNoise aug, freeze 2 epochs, normalize after CQT;</p></li>\n<li><p>fine-tuned EfficentNet-B3, initial stride=(1,2)  GaussNoise aug, freeze 2 epochs, normalize after CQT;</p></li>\n<li><p>fine-tuned EfficentNet-B0 , InstanceNorm instead of BatchNorm, normalize after CQT;</p></li>\n<li><p>1D ResNet + Transformer head,  filter order 8 and CosineAnnealing;</p></li>\n<li><p>1D ResNet + Transformer head,  filter order 5 and ReduceLROnPlateau;</p></li>\n<li><p>fine-tuned 6-layer CNN1d with MLP head, MixUp and GaussNoise, linear warm-up and cosine decay.</p></li>\n</ol>\n<p><strong>P.S</strong>. Other things we tried, but it didn't make it to the final blend:</p>\n<ul>\n<li>drop confident 1s before fine-tuning;</li>\n<li>InceptionTime for 1D + Transformer head;</li>\n<li>8 layer custom 1D CNN;</li>\n<li>rexnet_100;</li>\n<li>EfficeintNet-B2 and other;</li>\n<li>resize CQT image into (512,512) or (256, 256).</li>\n</ul>\n<p>Augmentations that we stopped using :</p>\n<ul>\n<li>Channels shuffle for an image in 2D models;</li>\n<li>Horizontal and vertical flip of images;</li>\n<li>normal MixUp.</li>\n</ul>",
      "rawMarkdown": "At first, I would like to thank my team members @vladimirsydor, @zekamrozek, @yakuben and @uulott for their work and organizers and the Kaggle community for the amazing experience we had during this competition.\n\n##TL;DR\n Blend of 6 2D CNN + 3 1D CNN: 1x2D EfficentNet_b0 + 3 x _b3 + 2 x _b5 + 2 x 1D ResNet with transformer + 1D basic CNN\n\n## Details\n**Bandpass** was crucial for this data, we experimented with it a lot and tried different setups, those parameters showed better results then others we tried:\n                `fmin=20`,\n                `fmax=1024` or `fmax=500` or `fmax=600` depending on the model,\n                `hop_length=4`,\n                `bins_per_octave=12`,\n                `filter_scale=0.5`,\n                `pad=10`.\n\nFor images, we used only CQT and didn't experiment with CWT.\n\n**Augmentations** that worked for us:\n1. Wave augs:\n- Random shift channel +-30;\n- Hard MixUp of waves: `new_wave = (wave1 + wave2) / 2`  and `new_target = max(target1, target2)`;\n- Adding Gaussian Noise with small std.\n\n\n**Models**\nFor most of the competition, we were training EfficentNets, until around 1-2 last weeks we started experimenting with 1D models.\n\nWe found out, that increasing the input image would increase the performance even for EfficentNet_B0, so we experimented with it more. For us best result was when we used `stride=(1,2)` instead of default `(2,2)` with an unchanged shape of input. That allowed model to \"stretch\" the input itself across the frequency axis and give an effect of feeding higher resolution images to the network.\n\nOur 2D models suffered from spoiled gradients in BatchNorm layers on first epochs, so to overcome that we used 2 procedures.  The first idea that worked was to substitute BatchNorm for InstanceNorm throughout the whole model.  2nd  idea that worked was that we would freeze all encoder layers but BatchNorm, and train only them plus Linear head for the first 2 epochs.\n\nAs was mentioned, we started working late on 1D models, mainly simple stacking Conv layers, but after merging with @uulott, he brought a really promising 1D ResNet model that performed on similar level as our best 2D ones (~0.879 LB).\n\nFinally, all of our models benefited when we started finetuning them on a small learning rate for multiple epochs with augmentations turned off. That boosted us past 0.88 LB and was keeping us in silver as we kept going up LB the more models we fine-tune and blend with.\n\n**Ensemble**\nFor ensembling, we used straightforward hyperopt blending. Adding Any 1D model boosted all of our blends.\n\n\n**The final blend** consisted of 6 2D-CNNs and 3 1D-CNNs:\n\n1.  EfficentNet-B5, successful augs, freeze 2 epochs, normalize after CQT; \n2. Same B5, but fine-tuned;\n3.  fine-tuned EfficentNet-B3, all successful augs, freeze 2 epochs, normalize after CQT;\n4. fine-tuned EfficentNet-B3, initial stride=(1,2)  GaussNoise aug, freeze 2 epochs, normalize after CQT;\n5. fine-tuned EfficentNet-B3, initial stride=(1,2)  GaussNoise aug, freeze 2 epochs, normalize after CQT;\n6. fine-tuned EfficentNet-B0 , InstanceNorm instead of BatchNorm, normalize after CQT;\n\n7.  1D ResNet + Transformer head,  filter order 8 and CosineAnnealing;\n8. 1D ResNet + Transformer head,  filter order 5 and ReduceLROnPlateau;\n9. fine-tuned 6-layer CNN1d with MLP head, MixUp and GaussNoise, linear warm-up and cosine decay.\n\n**P.S**. Other things we tried, but it didn't make it to the final blend:\n- drop confident 1s before fine-tuning;\n- InceptionTime for 1D + Transformer head;\n- 8 layer custom 1D CNN;\n- rexnet_100;\n- EfficeintNet-B2 and other;\n- resize CQT image into (512,512) or (256, 256).\n\nAugmentations that we stopped using :\n- Channels shuffle for an image in 2D models;\n- Horizontal and vertical flip of images;\n- normal MixUp.",
      "votes": null
    },
    {
      "id": "1531215",
      "postDate": "10/01/2021 18:20:31",
      "content": "<p>Fun to know that low learning rate worked for you… I got stuck for 2 weeks @ CV 8750 only because I was using lr &lt; 1e-4. After trying lower epochs and lr 1e-3 I got a big jump to 8795.</p>",
      "rawMarkdown": "Fun to know that low learning rate worked for you… I got stuck for 2 weeks @ CV 8750 only because I was using lr < 1e-4. After trying lower epochs and lr 1e-3 I got a big jump to 8795.",
      "votes": null
    },
    {
      "id": "1559900",
      "postDate": "10/27/2021 08:07:12",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1531215,
      "author_name": "callmeb",
      "author_url": "",
      "post_date": "10/01/2021 18:20:31",
      "content": "<p>Fun to know that low learning rate worked for you… I got stuck for 2 weeks @ CV 8750 only because I was using lr &lt; 1e-4. After trying lower epochs and lr 1e-3 I got a big jump to 8795.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1559900,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 08:07:12",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1530787": "At first, I would like to thank my team members @vladimirsydor, @zekamrozek, @yakuben and @uulott for their work and organizers and the Kaggle community for the amazing experience we had during this competition.\n\n##TL;DR\n Blend of 6 2D CNN + 3 1D CNN: 1x2D EfficentNet_b0 + 3 x _b3 + 2 x _b5 + 2 x 1D ResNet with transformer + 1D basic CNN\n\n## Details\n**Bandpass** was crucial for this data, we experimented with it a lot and tried different setups, those parameters showed better results then others we tried:\n                `fmin=20`,\n                `fmax=1024` or `fmax=500` or `fmax=600` depending on the model,\n                `hop_length=4`,\n                `bins_per_octave=12`,\n                `filter_scale=0.5`,\n                `pad=10`.\n\nFor images, we used only CQT and didn't experiment with CWT.\n\n**Augmentations** that worked for us:\n1. Wave augs:\n- Random shift channel +-30;\n- Hard MixUp of waves: `new_wave = (wave1 + wave2) / 2`  and `new_target = max(target1, target2)`;\n- Adding Gaussian Noise with small std.\n\n\n**Models**\nFor most of the competition, we were training EfficentNets, until around 1-2 last weeks we started experimenting with 1D models.\n\nWe found out, that increasing the input image would increase the performance even for EfficentNet_B0, so we experimented with it more. For us best result was when we used `stride=(1,2)` instead of default `(2,2)` with an unchanged shape of input. That allowed model to \"stretch\" the input itself across the frequency axis and give an effect of feeding higher resolution images to the network.\n\nOur 2D models suffered from spoiled gradients in BatchNorm layers on first epochs, so to overcome that we used 2 procedures.  The first idea that worked was to substitute BatchNorm for InstanceNorm throughout the whole model.  2nd  idea that worked was that we would freeze all encoder layers but BatchNorm, and train only them plus Linear head for the first 2 epochs.\n\nAs was mentioned, we started working late on 1D models, mainly simple stacking Conv layers, but after merging with @uulott, he brought a really promising 1D ResNet model that performed on similar level as our best 2D ones (~0.879 LB).\n\nFinally, all of our models benefited when we started finetuning them on a small learning rate for multiple epochs with augmentations turned off. That boosted us past 0.88 LB and was keeping us in silver as we kept going up LB the more models we fine-tune and blend with.\n\n**Ensemble**\nFor ensembling, we used straightforward hyperopt blending. Adding Any 1D model boosted all of our blends.\n\n\n**The final blend** consisted of 6 2D-CNNs and 3 1D-CNNs:\n\n1.  EfficentNet-B5, successful augs, freeze 2 epochs, normalize after CQT; \n2. Same B5, but fine-tuned;\n3.  fine-tuned EfficentNet-B3, all successful augs, freeze 2 epochs, normalize after CQT;\n4. fine-tuned EfficentNet-B3, initial stride=(1,2)  GaussNoise aug, freeze 2 epochs, normalize after CQT;\n5. fine-tuned EfficentNet-B3, initial stride=(1,2)  GaussNoise aug, freeze 2 epochs, normalize after CQT;\n6. fine-tuned EfficentNet-B0 , InstanceNorm instead of BatchNorm, normalize after CQT;\n\n7.  1D ResNet + Transformer head,  filter order 8 and CosineAnnealing;\n8. 1D ResNet + Transformer head,  filter order 5 and ReduceLROnPlateau;\n9. fine-tuned 6-layer CNN1d with MLP head, MixUp and GaussNoise, linear warm-up and cosine decay.\n\n**P.S**. Other things we tried, but it didn't make it to the final blend:\n- drop confident 1s before fine-tuning;\n- InceptionTime for 1D + Transformer head;\n- 8 layer custom 1D CNN;\n- rexnet_100;\n- EfficeintNet-B2 and other;\n- resize CQT image into (512,512) or (256, 256).\n\nAugmentations that we stopped using :\n- Channels shuffle for an image in 2D models;\n- Horizontal and vertical flip of images;\n- normal MixUp.",
    "1531215": "Fun to know that low learning rate worked for you… I got stuck for 2 weeks @ CV 8750 only because I was using lr < 1e-4. After trying lower epochs and lr 1e-3 I got a big jump to 8795.",
    "1559900": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}