{
  "id": 275617,
  "title": "3rd place solution",
  "url": "/competitions/g2net-gravitational-wave-detection/writeups/gogogo-jarvislabs-ai-3rd-place-solution",
  "author_name": "",
  "post_date": "2021-10-01T05:43:34.410Z",
  "votes": 80,
  "comment_count": 39,
  "views": 0,
  "content": "<h3>TL;DR</h3>\n<h5>Models we used in our best submission: 15 1D Model  + 8 2D Model</h5>\n<p><strong>2D Models</strong> [best single model:  <strong>0.87875/0.8805/0.8787</strong> at CV, public, and private LB]</p>\n<ul>\n<li>CQT, CWT Transformation</li>\n<li>Efficientnet, Resnext, InceptionV3…</li>\n<li>TTA: shuffle LIGO channel</li>\n</ul>\n<p><strong>1D Models</strong> [best single model: <strong>0.8819/0.8827/0.8820</strong> at CV, public, and private LB]</p>\n<ul>\n<li>Customized architecture targeted at GW detection</li>\n<li>TTA: vflip, shuffle LIGO channels, Gaussian noise, time shift, time mask, MC dropout</li>\n</ul>\n<p><strong>Training:</strong></p>\n<ul>\n<li>Pretraining with GW</li>\n<li>Training with Pseudo Label or Soft Label</li>\n<li>BCE, Rank Loss</li>\n<li>MixUp</li>\n<li>AdamW,  RangerLars+Lookahead optimizer </li>\n</ul>\n<p><strong>Preprocessing:</strong></p>\n<ul>\n<li>Avg PSD of target 0 (Design Curves)</li>\n<li>Extending waves</li>\n<li>Whitening with Tukey window</li>\n</ul>\n<p><strong>GW Simulation</strong></p>\n<ul>\n<li>Distribution of Total Mass</li>\n<li>SNR injection ratio max(N(3.6,1),1)</li>\n</ul>\n<p><strong>Ensemble</strong></p>\n<ul>\n<li>CMA-ES Optimization</li>\n<li>Hacking Private LB</li>\n</ul>\n<h3>Introduction</h3>\n<p>Our team would like to thank organizers and kaggle for making this competition possible. Also, I want to express my gratitude to my outstanding teammates <a href=\"https://www.kaggle.com/vincentwang25\" target=\"_blank\">@vincentwang25</a>, <a href=\"https://www.kaggle.com/richx86\" target=\"_blank\">@richx86</a>, <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a>, and <a href=\"https://www.kaggle.com/yamsam\" target=\"_blank\">@yamsam</a> for their incredible contribution towards our final result. In this competition I got the last gold medal required for getting kaggle GM, and also it is the first gold medal for <a href=\"https://www.kaggle.com/vincentwang25\" target=\"_blank\">@vincentwang25</a> and <a href=\"https://www.kaggle.com/richx86\" target=\"_blank\">@richx86</a> . The provided write-up is the result of the cumulative efforts of all our team members, and I just got the privilege to post it.</p>\n<h3>Details</h3>\n<h4>2D Models</h4>\n<p>Like many of the participants, our team initially focused on 2D models. Among the various approaches we tried, we found that whitened signals, which we will discuss later, was the most effective preprocessing for signals. We created the following 2D models using the CQT and CWT images generated from the whitened signal as input.</p>\n<p>We also tried various augmentations for 2D images (mixup, cutout, …), but most of them did not work here, however mixup on the input waveform prior to CQT/CWT did help counteract overfitting. The most effective one was the augmentation of swapping LIGO signals. This worked both for training and inference (TTA, Test Time Augmentation) and we found soft pseudo labeling also improved the score. </p>\n<p>The performance of our best single 2D model is <strong>0.87875/0.8805/0.8787</strong> at CV, public, and private LB. </p>\n<ul>\n<li>EfficientNet(B3, B4, B5, B7), EfficientNetV2(M), ResNet200D, Inception-V3 (also we performed a number of initial experiments with ResNeXt models)</li>\n<li>CQT and CWT images generated based on the whitened signal</li>\n<li>Image size 128 x 128 〜 512 x 512</li>\n<li>Soft leak-free pseudo labeling from ensemble results</li>\n<li>LIGO channel swap argumentation (randomly swapping LIGO channels) for both training and TTA</li>\n<li>1D mixup prior to CQT/CWT</li>\n<li>Adding a 4th channel to the spectrogram/scalogram input which is just a linear gradient (-1, 1) along the frequency dimension used for frequency encoding (similar to positional encoding in transformers)</li>\n</ul>\n<h4>1D Models</h4>\n<p>1D models appeared to be the key component of our solution, even if we didn't make great improvement until the last 1.5 weeks of the competition. The reason why these models were not widely explored by most of the participants may be the need of using signal whitening to reach a performance comparable with 2D models (at least easily), and whitening is not straightforward for 2s long signals (see discussion below). However, 1D models are much faster to train, and they also outperform our 2D setup. The performance of our best single 1D model is <strong>0.8819/0.8827/0.8820</strong> at CV, public, and private LB. So <strong>it can reach top-7 LB after ~8 hours of training</strong>.</p>\n<p>One of the main contributions towards this result is the proper choice of the model architecture for GW detection task. Specifically, <strong>GW is not just a signal of the specific shape, but rather a correlation in signal between multiple detectors</strong>. Meanwhile, signals may be shifted by up to ~10-20 ms because of the time needed for the signal to cross the distance between detectors. So direct concatenation of signals into (3,4096) stack and then applying a joined convolution is not a good idea (our baseline V0 architecture with CV of 0.8768). Random shift between components prohibits the generation of joined features. Thus, we asked the question, why not split the network into branches for different detectors, like proposed in <a href=\"https://www.sciencedirect.com/science/article/pii/S0370269320308327\" target=\"_blank\">this paper</a>? So the first layers, extractor and following Res blocks, learn how to extract features from a signal, a kind of learnable FFT or wavelet transform. So before merging the signals the network already has a basic understanding of what is going on. We also share weights between LIGO1 and LIGO2 branches because the signals are very similar.<br>\n<img src=\"https://i.ibb.co/4PgdmLY/1Dmodel.png\" alt=\"\"></p>\n<p>Merge of the extracted features instead of the original signal mitigates the effect of the relative shift (like in Short Time Fourier Transform correlation turns into a product of two aligned cells). So simple concatenation at this stage (instead of early concatenation) followed by several Res blocks (V1 architecture) gives an improvement from <strong>0.8768 to 0.8789</strong>. However, the model after combining the signal and getting a better idea about GW, may still want to look into individual branches as a reference. So we extend our individual branches and perform the second concatenation at the next convolutional block (V2 architecture). It results in a further improvement of CV to <strong>0.8804</strong>. </p>\n<p>After the basic structure of the V2 model was defined, we performed a number of additional experiments for further model optimization. The model architecture tricks giving further improvement include the use of SiLU instead of ReLU, use of GeM instead of strided convolution, use of concatenation pooling at the end of the convolutional part. In one of our final runs, we used ResNeSt blocks (Split Attention convolution) having a comparable performance in the preliminary experiments, but it also performed slightly worse at the end of full training. Using CBAM modules and Stochastic Depth modules gave a slight boost and made the 1D models more diversified. The network width, n, is equal to 16 and 32 for our final models. One of our experiments is also performed for a combined 1D+2D models, pretrained separately and then finetuned jointly, which gave 0.8815/0.8831/0.8817 score.</p>\n<p><strong>Things that didn’t work:</strong> use of larger network depth (extra Res blocks), use of smaller/larger conv size in the extractor/first blocks, use of multi-head self-attention blocks before pooling or in the last stages of ResBlocks (i.e. Bottleneck Transformers), use of densely connected or ResNeXt blocks instead of ReBlocks, learnable CWT like extractors (FFT-&gt;multiplication by learnable weights-&gt;iFFT), WaveNet like network structures, pretrained ViT following the extractors (instead of customized resnet).</p>\n<h5>Training (1D models)</h5>\n<ul>\n<li>Pretraining with Simulated GW for 2~4 epochs (<strong>2-8bps</strong> boost)</li>\n<li>Training with Pseudo-labeling data for 4~6 epochs depending on the model</li>\n<li>Training with rank loss and low lr for 2 epoch (<strong>~1bps</strong> performance boost)</li>\n</ul>\n<p>TTA based on multiplying the signal by -1 and first/second channel swap gave <strong>~2bps</strong> boost. 64 fold MC dropout gave an additional <strong>~1bps</strong> boost. Also during training in some experiments we used 0.05 spectral dropout: drop the given percentage of FFT spectrum during whitening.</p>\n<h3>Preprocessing</h3>\n<p>Regarding whitening, direct use of packages, such as pycbc, doesn’t work mainly because of the short duration of the provided signals: only 2 seconds chunks in contrast to the almost unlimited length of data from LIGO/VIRGO detectors. To make the estimated PSD smooth, pycbc package uses an algorithm that corrupts the boundary of data, which is too costly for our dataset, whose duration of signals is only 2 seconds. We reduce the variance of estimated PSD by taking the average PSD for all negative training samples. This is the key to make whitening work (interestingly, this averaging idea came up independently to two of the team members). To further reduce the boundary effect from discontinuity and allow ourselves to use a window function that has a larger decay area (for example, Tukey window with a larger alpha), we first <strong>extend the wave while keeping the first derivative continuous at boundaries</strong>. Finally, we apply the window function to each data and <strong>use the average PSD to normalize the wave for different frequencies</strong>.</p>\n<p><strong>For 1D model:</strong></p>\n<ul>\n<li>Extend the signal to (3,8192)</li>\n<li>Tukey window with alpha 0.5</li>\n<li>Whitening with PSD based on the average FFT of all negative training examples</li>\n</ul>\n<p><strong>For 2D model:</strong></p>\n<ul>\n<li>Using the whiten data and apply CQT transform with the following parameters:</li>\n<li>CQT1992v2(sr=2048, fmin=20, fmax=1000, window='flattop', bins_per_octave=48, filter_scale=0.25, hop_length=8)</li>\n<li>Resize (128x128, 256x256 …. ) </li>\n</ul>\n<p><img src=\"https://i.ibb.co/CtL2nk6/whitening.png\" alt=\"\"></p>\n<h3>GW Simulation</h3>\n<p>This idea is coming from curriculum learning, and in <a href=\"https://arxiv.org/abs/2106.03741\" target=\"_blank\">this paper</a>, it mentioned that “We find that the deep learning algorithms can generalize <strong>low signal-to-noise ratio (SNR) signals to high SNR ones but not vice versa</strong>”, so we follow it and try to generate a signal and inject into the noise with low SNR. Even though there are around 15 parameters, we found that the most important one is the total mass and mass ratio (maybe we are wrong) because it affects the shape of GW the most through eyeballing.  So we adjust the total mass and mass ratio using different distributions and inject the signal into the noise with a given SNR following max(random.gauss(3.6,1),1) distribution. This SNR distribution is determined by checking the training loss trend: we want it to follow the trend of original data (not too hard, not too simple). This way gives us a <strong>2~8bps increase</strong> depending on the model we use.<br>\nWe also tried to follow this idea by giving Hard Positive samples from the train data more weight but due to the time constraints, we didn’t make it work. It could be a potential win here.</p>\n<h3>Ensemble</h3>\n<p>First, to confirm that train and test data are similar and do not have any hidden peculiarities we used <strong>adversarial validation</strong> that gave 0.5 AUC = train and test data are indistinguishable.</p>\n<p>We tried many different methods to ensemble the models and saw the following relative performance trend: <a href=\"https://github.com/CMA-ES/pycma\" target=\"_blank\">CME-ES</a> with Logit &gt; CME-ES with rank&gt; Scipy Optimization &gt; Neural Network &gt; other methods.  We also tried to use <code>sklearn.preprocessing.PolynomialFeatures</code> with probability prediction to do the CME-ES optimization. It brings the highest CV and LB but with a little chance of overfitting. We are glad that it turns out to be our best 0.8829 submission.</p>\n<p>Our second submission is based on a simulation of the private LB by excluding 16% of our OOF data out of CV optimization. So the produced model weights are more robust to the data noise and potentially can lead to a better performance at the private LB. We do this because we found the CV and LB correlation is very high and we also used adversarial validation to check that indeed they are similar. So We bootstrapped 16% OOF data which has a similar CV to public LB score and <strong>the optimized CV for the remaining data matched the same as the private LB (0.8828 for one of our submissions)</strong>.  </p>\n<h3>Acknowledgment</h3>\n<p>Some of the team members used JarvisLabs.ai for their GPU workstations. The cloud platform is easy to use, stable, and very affordable. The pause functionality is great. Recommend this platform to Kagglers.</p>",
  "messages": [
    {
      "id": "1530199",
      "postDate": "10/01/2021 02:48:17",
      "content": "<h3>TL;DR</h3>\n<h5>Models we used in our best submission: 15 1D Model  + 8 2D Model</h5>\n<p><strong>2D Models</strong> [best single model:  <strong>0.87875/0.8805/0.8787</strong> at CV, public, and private LB]</p>\n<ul>\n<li>CQT, CWT Transformation</li>\n<li>Efficientnet, Resnext, InceptionV3…</li>\n<li>TTA: shuffle LIGO channel</li>\n</ul>\n<p><strong>1D Models</strong> [best single model: <strong>0.8819/0.8827/0.8820</strong> at CV, public, and private LB]</p>\n<ul>\n<li>Customized architecture targeted at GW detection</li>\n<li>TTA: vflip, shuffle LIGO channels, Gaussian noise, time shift, time mask, MC dropout</li>\n</ul>\n<p><strong>Training:</strong></p>\n<ul>\n<li>Pretraining with GW</li>\n<li>Training with Pseudo Label or Soft Label</li>\n<li>BCE, Rank Loss</li>\n<li>MixUp</li>\n<li>AdamW,  RangerLars+Lookahead optimizer </li>\n</ul>\n<p><strong>Preprocessing:</strong></p>\n<ul>\n<li>Avg PSD of target 0 (Design Curves)</li>\n<li>Extending waves</li>\n<li>Whitening with Tukey window</li>\n</ul>\n<p><strong>GW Simulation</strong></p>\n<ul>\n<li>Distribution of Total Mass</li>\n<li>SNR injection ratio max(N(3.6,1),1)</li>\n</ul>\n<p><strong>Ensemble</strong></p>\n<ul>\n<li>CMA-ES Optimization</li>\n<li>Hacking Private LB</li>\n</ul>\n<h3>Introduction</h3>\n<p>Our team would like to thank organizers and kaggle for making this competition possible. Also, I want to express my gratitude to my outstanding teammates <a href=\"https://www.kaggle.com/vincentwang25\" target=\"_blank\">@vincentwang25</a>, <a href=\"https://www.kaggle.com/richx86\" target=\"_blank\">@richx86</a>, <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a>, and <a href=\"https://www.kaggle.com/yamsam\" target=\"_blank\">@yamsam</a> for their incredible contribution towards our final result. In this competition I got the last gold medal required for getting kaggle GM, and also it is the first gold medal for <a href=\"https://www.kaggle.com/vincentwang25\" target=\"_blank\">@vincentwang25</a> and <a href=\"https://www.kaggle.com/richx86\" target=\"_blank\">@richx86</a> . The provided write-up is the result of the cumulative efforts of all our team members, and I just got the privilege to post it.</p>\n<h3>Details</h3>\n<h4>2D Models</h4>\n<p>Like many of the participants, our team initially focused on 2D models. Among the various approaches we tried, we found that whitened signals, which we will discuss later, was the most effective preprocessing for signals. We created the following 2D models using the CQT and CWT images generated from the whitened signal as input.</p>\n<p>We also tried various augmentations for 2D images (mixup, cutout, …), but most of them did not work here, however mixup on the input waveform prior to CQT/CWT did help counteract overfitting. The most effective one was the augmentation of swapping LIGO signals. This worked both for training and inference (TTA, Test Time Augmentation) and we found soft pseudo labeling also improved the score. </p>\n<p>The performance of our best single 2D model is <strong>0.87875/0.8805/0.8787</strong> at CV, public, and private LB. </p>\n<ul>\n<li>EfficientNet(B3, B4, B5, B7), EfficientNetV2(M), ResNet200D, Inception-V3 (also we performed a number of initial experiments with ResNeXt models)</li>\n<li>CQT and CWT images generated based on the whitened signal</li>\n<li>Image size 128 x 128 〜 512 x 512</li>\n<li>Soft leak-free pseudo labeling from ensemble results</li>\n<li>LIGO channel swap argumentation (randomly swapping LIGO channels) for both training and TTA</li>\n<li>1D mixup prior to CQT/CWT</li>\n<li>Adding a 4th channel to the spectrogram/scalogram input which is just a linear gradient (-1, 1) along the frequency dimension used for frequency encoding (similar to positional encoding in transformers)</li>\n</ul>\n<h4>1D Models</h4>\n<p>1D models appeared to be the key component of our solution, even if we didn't make great improvement until the last 1.5 weeks of the competition. The reason why these models were not widely explored by most of the participants may be the need of using signal whitening to reach a performance comparable with 2D models (at least easily), and whitening is not straightforward for 2s long signals (see discussion below). However, 1D models are much faster to train, and they also outperform our 2D setup. The performance of our best single 1D model is <strong>0.8819/0.8827/0.8820</strong> at CV, public, and private LB. So <strong>it can reach top-7 LB after ~8 hours of training</strong>.</p>\n<p>One of the main contributions towards this result is the proper choice of the model architecture for GW detection task. Specifically, <strong>GW is not just a signal of the specific shape, but rather a correlation in signal between multiple detectors</strong>. Meanwhile, signals may be shifted by up to ~10-20 ms because of the time needed for the signal to cross the distance between detectors. So direct concatenation of signals into (3,4096) stack and then applying a joined convolution is not a good idea (our baseline V0 architecture with CV of 0.8768). Random shift between components prohibits the generation of joined features. Thus, we asked the question, why not split the network into branches for different detectors, like proposed in <a href=\"https://www.sciencedirect.com/science/article/pii/S0370269320308327\" target=\"_blank\">this paper</a>? So the first layers, extractor and following Res blocks, learn how to extract features from a signal, a kind of learnable FFT or wavelet transform. So before merging the signals the network already has a basic understanding of what is going on. We also share weights between LIGO1 and LIGO2 branches because the signals are very similar.<br>\n<img src=\"https://i.ibb.co/4PgdmLY/1Dmodel.png\" alt=\"\"></p>\n<p>Merge of the extracted features instead of the original signal mitigates the effect of the relative shift (like in Short Time Fourier Transform correlation turns into a product of two aligned cells). So simple concatenation at this stage (instead of early concatenation) followed by several Res blocks (V1 architecture) gives an improvement from <strong>0.8768 to 0.8789</strong>. However, the model after combining the signal and getting a better idea about GW, may still want to look into individual branches as a reference. So we extend our individual branches and perform the second concatenation at the next convolutional block (V2 architecture). It results in a further improvement of CV to <strong>0.8804</strong>. </p>\n<p>After the basic structure of the V2 model was defined, we performed a number of additional experiments for further model optimization. The model architecture tricks giving further improvement include the use of SiLU instead of ReLU, use of GeM instead of strided convolution, use of concatenation pooling at the end of the convolutional part. In one of our final runs, we used ResNeSt blocks (Split Attention convolution) having a comparable performance in the preliminary experiments, but it also performed slightly worse at the end of full training. Using CBAM modules and Stochastic Depth modules gave a slight boost and made the 1D models more diversified. The network width, n, is equal to 16 and 32 for our final models. One of our experiments is also performed for a combined 1D+2D models, pretrained separately and then finetuned jointly, which gave 0.8815/0.8831/0.8817 score.</p>\n<p><strong>Things that didn’t work:</strong> use of larger network depth (extra Res blocks), use of smaller/larger conv size in the extractor/first blocks, use of multi-head self-attention blocks before pooling or in the last stages of ResBlocks (i.e. Bottleneck Transformers), use of densely connected or ResNeXt blocks instead of ReBlocks, learnable CWT like extractors (FFT-&gt;multiplication by learnable weights-&gt;iFFT), WaveNet like network structures, pretrained ViT following the extractors (instead of customized resnet).</p>\n<h5>Training (1D models)</h5>\n<ul>\n<li>Pretraining with Simulated GW for 2~4 epochs (<strong>2-8bps</strong> boost)</li>\n<li>Training with Pseudo-labeling data for 4~6 epochs depending on the model</li>\n<li>Training with rank loss and low lr for 2 epoch (<strong>~1bps</strong> performance boost)</li>\n</ul>\n<p>TTA based on multiplying the signal by -1 and first/second channel swap gave <strong>~2bps</strong> boost. 64 fold MC dropout gave an additional <strong>~1bps</strong> boost. Also during training in some experiments we used 0.05 spectral dropout: drop the given percentage of FFT spectrum during whitening.</p>\n<h3>Preprocessing</h3>\n<p>Regarding whitening, direct use of packages, such as pycbc, doesn’t work mainly because of the short duration of the provided signals: only 2 seconds chunks in contrast to the almost unlimited length of data from LIGO/VIRGO detectors. To make the estimated PSD smooth, pycbc package uses an algorithm that corrupts the boundary of data, which is too costly for our dataset, whose duration of signals is only 2 seconds. We reduce the variance of estimated PSD by taking the average PSD for all negative training samples. This is the key to make whitening work (interestingly, this averaging idea came up independently to two of the team members). To further reduce the boundary effect from discontinuity and allow ourselves to use a window function that has a larger decay area (for example, Tukey window with a larger alpha), we first <strong>extend the wave while keeping the first derivative continuous at boundaries</strong>. Finally, we apply the window function to each data and <strong>use the average PSD to normalize the wave for different frequencies</strong>.</p>\n<p><strong>For 1D model:</strong></p>\n<ul>\n<li>Extend the signal to (3,8192)</li>\n<li>Tukey window with alpha 0.5</li>\n<li>Whitening with PSD based on the average FFT of all negative training examples</li>\n</ul>\n<p><strong>For 2D model:</strong></p>\n<ul>\n<li>Using the whiten data and apply CQT transform with the following parameters:</li>\n<li>CQT1992v2(sr=2048, fmin=20, fmax=1000, window='flattop', bins_per_octave=48, filter_scale=0.25, hop_length=8)</li>\n<li>Resize (128x128, 256x256 …. ) </li>\n</ul>\n<p><img src=\"https://i.ibb.co/CtL2nk6/whitening.png\" alt=\"\"></p>\n<h3>GW Simulation</h3>\n<p>This idea is coming from curriculum learning, and in <a href=\"https://arxiv.org/abs/2106.03741\" target=\"_blank\">this paper</a>, it mentioned that “We find that the deep learning algorithms can generalize <strong>low signal-to-noise ratio (SNR) signals to high SNR ones but not vice versa</strong>”, so we follow it and try to generate a signal and inject into the noise with low SNR. Even though there are around 15 parameters, we found that the most important one is the total mass and mass ratio (maybe we are wrong) because it affects the shape of GW the most through eyeballing.  So we adjust the total mass and mass ratio using different distributions and inject the signal into the noise with a given SNR following max(random.gauss(3.6,1),1) distribution. This SNR distribution is determined by checking the training loss trend: we want it to follow the trend of original data (not too hard, not too simple). This way gives us a <strong>2~8bps increase</strong> depending on the model we use.<br>\nWe also tried to follow this idea by giving Hard Positive samples from the train data more weight but due to the time constraints, we didn’t make it work. It could be a potential win here.</p>\n<h3>Ensemble</h3>\n<p>First, to confirm that train and test data are similar and do not have any hidden peculiarities we used <strong>adversarial validation</strong> that gave 0.5 AUC = train and test data are indistinguishable.</p>\n<p>We tried many different methods to ensemble the models and saw the following relative performance trend: <a href=\"https://github.com/CMA-ES/pycma\" target=\"_blank\">CME-ES</a> with Logit &gt; CME-ES with rank&gt; Scipy Optimization &gt; Neural Network &gt; other methods.  We also tried to use <code>sklearn.preprocessing.PolynomialFeatures</code> with probability prediction to do the CME-ES optimization. It brings the highest CV and LB but with a little chance of overfitting. We are glad that it turns out to be our best 0.8829 submission.</p>\n<p>Our second submission is based on a simulation of the private LB by excluding 16% of our OOF data out of CV optimization. So the produced model weights are more robust to the data noise and potentially can lead to a better performance at the private LB. We do this because we found the CV and LB correlation is very high and we also used adversarial validation to check that indeed they are similar. So We bootstrapped 16% OOF data which has a similar CV to public LB score and <strong>the optimized CV for the remaining data matched the same as the private LB (0.8828 for one of our submissions)</strong>.  </p>\n<h3>Acknowledgment</h3>\n<p>Some of the team members used JarvisLabs.ai for their GPU workstations. The cloud platform is easy to use, stable, and very affordable. The pause functionality is great. Recommend this platform to Kagglers.</p>",
      "rawMarkdown": "### TL;DR\n##### Models we used in our best submission: 15 1D Model  + 8 2D Model\n**2D Models** [best single model:  **0.87875/0.8805/0.8787** at CV, public, and private LB]\n- CQT, CWT Transformation\n- Efficientnet, Resnext, InceptionV3...\n- TTA: shuffle LIGO channel\n\n**1D Models** [best single model: **0.8819/0.8827/0.8820** at CV, public, and private LB]\n- Customized architecture targeted at GW detection\n- TTA: vflip, shuffle LIGO channels, Gaussian noise, time shift, time mask, MC dropout\n\n**Training:**\n- Pretraining with GW\n- Training with Pseudo Label or Soft Label\n- BCE, Rank Loss\n- MixUp\n- AdamW,  RangerLars+Lookahead optimizer \n\n**Preprocessing:**\n- Avg PSD of target 0 (Design Curves)\n- Extending waves\n- Whitening with Tukey window\n\n**GW Simulation**\n- Distribution of Total Mass\n- SNR injection ratio max(N(3.6,1),1)\n\n**Ensemble**\n- CMA-ES Optimization\n- Hacking Private LB\n\n### Introduction\nOur team would like to thank organizers and kaggle for making this competition possible. Also, I want to express my gratitude to my outstanding teammates @vincentwang25, @richx86, @anjum48, and @yamsam for their incredible contribution towards our final result. In this competition I got the last gold medal required for getting kaggle GM, and also it is the first gold medal for @vincentwang25 and @richx86 . The provided write-up is the result of the cumulative efforts of all our team members, and I just got the privilege to post it.\n\n### Details \n#### 2D Models\nLike many of the participants, our team initially focused on 2D models. Among the various approaches we tried, we found that whitened signals, which we will discuss later, was the most effective preprocessing for signals. We created the following 2D models using the CQT and CWT images generated from the whitened signal as input.\n\nWe also tried various augmentations for 2D images (mixup, cutout, ...), but most of them did not work here, however mixup on the input waveform prior to CQT/CWT did help counteract overfitting. The most effective one was the augmentation of swapping LIGO signals. This worked both for training and inference (TTA, Test Time Augmentation) and we found soft pseudo labeling also improved the score. \n\nThe performance of our best single 2D model is **0.87875/0.8805/0.8787** at CV, public, and private LB. \n- EfficientNet(B3, B4, B5, B7), EfficientNetV2(M), ResNet200D, Inception-V3 (also we performed a number of initial experiments with ResNeXt models)\n- CQT and CWT images generated based on the whitened signal\n- Image size 128 x 128 〜 512 x 512\n- Soft leak-free pseudo labeling from ensemble results\n- LIGO channel swap argumentation (randomly swapping LIGO channels) for both training and TTA\n- 1D mixup prior to CQT/CWT\n- Adding a 4th channel to the spectrogram/scalogram input which is just a linear gradient (-1, 1) along the frequency dimension used for frequency encoding (similar to positional encoding in transformers)\n#### 1D Models\n1D models appeared to be the key component of our solution, even if we didn't make great improvement until the last 1.5 weeks of the competition. The reason why these models were not widely explored by most of the participants may be the need of using signal whitening to reach a performance comparable with 2D models (at least easily), and whitening is not straightforward for 2s long signals (see discussion below). However, 1D models are much faster to train, and they also outperform our 2D setup. The performance of our best single 1D model is **0.8819/0.8827/0.8820** at CV, public, and private LB. So **it can reach top-7 LB after ~8 hours of training**.\n\nOne of the main contributions towards this result is the proper choice of the model architecture for GW detection task. Specifically, **GW is not just a signal of the specific shape, but rather a correlation in signal between multiple detectors**. Meanwhile, signals may be shifted by up to ~10-20 ms because of the time needed for the signal to cross the distance between detectors. So direct concatenation of signals into (3,4096) stack and then applying a joined convolution is not a good idea (our baseline V0 architecture with CV of 0.8768). Random shift between components prohibits the generation of joined features. Thus, we asked the question, why not split the network into branches for different detectors, like proposed in [this paper](https://www.sciencedirect.com/science/article/pii/S0370269320308327)? So the first layers, extractor and following Res blocks, learn how to extract features from a signal, a kind of learnable FFT or wavelet transform. So before merging the signals the network already has a basic understanding of what is going on. We also share weights between LIGO1 and LIGO2 branches because the signals are very similar.\n![](https://i.ibb.co/4PgdmLY/1Dmodel.png)\n\nMerge of the extracted features instead of the original signal mitigates the effect of the relative shift (like in Short Time Fourier Transform correlation turns into a product of two aligned cells). So simple concatenation at this stage (instead of early concatenation) followed by several Res blocks (V1 architecture) gives an improvement from **0.8768 to 0.8789**. However, the model after combining the signal and getting a better idea about GW, may still want to look into individual branches as a reference. So we extend our individual branches and perform the second concatenation at the next convolutional block (V2 architecture). It results in a further improvement of CV to **0.8804**. \n\nAfter the basic structure of the V2 model was defined, we performed a number of additional experiments for further model optimization. The model architecture tricks giving further improvement include the use of SiLU instead of ReLU, use of GeM instead of strided convolution, use of concatenation pooling at the end of the convolutional part. In one of our final runs, we used ResNeSt blocks (Split Attention convolution) having a comparable performance in the preliminary experiments, but it also performed slightly worse at the end of full training. Using CBAM modules and Stochastic Depth modules gave a slight boost and made the 1D models more diversified. The network width, n, is equal to 16 and 32 for our final models. One of our experiments is also performed for a combined 1D+2D models, pretrained separately and then finetuned jointly, which gave 0.8815/0.8831/0.8817 score.\n\n**Things that didn’t work:** use of larger network depth (extra Res blocks), use of smaller/larger conv size in the extractor/first blocks, use of multi-head self-attention blocks before pooling or in the last stages of ResBlocks (i.e. Bottleneck Transformers), use of densely connected or ResNeXt blocks instead of ReBlocks, learnable CWT like extractors (FFT->multiplication by learnable weights->iFFT), WaveNet like network structures, pretrained ViT following the extractors (instead of customized resnet).\n\n##### Training (1D models)\n- Pretraining with Simulated GW for 2~4 epochs (**2-8bps** boost)\n- Training with Pseudo-labeling data for 4~6 epochs depending on the model\n- Training with rank loss and low lr for 2 epoch (**~1bps** performance boost)\n\nTTA based on multiplying the signal by -1 and first/second channel swap gave **~2bps** boost. 64 fold MC dropout gave an additional **~1bps** boost. Also during training in some experiments we used 0.05 spectral dropout: drop the given percentage of FFT spectrum during whitening.\n\n### Preprocessing\nRegarding whitening, direct use of packages, such as pycbc, doesn’t work mainly because of the short duration of the provided signals: only 2 seconds chunks in contrast to the almost unlimited length of data from LIGO/VIRGO detectors. To make the estimated PSD smooth, pycbc package uses an algorithm that corrupts the boundary of data, which is too costly for our dataset, whose duration of signals is only 2 seconds. We reduce the variance of estimated PSD by taking the average PSD for all negative training samples. This is the key to make whitening work (interestingly, this averaging idea came up independently to two of the team members). To further reduce the boundary effect from discontinuity and allow ourselves to use a window function that has a larger decay area (for example, Tukey window with a larger alpha), we first **extend the wave while keeping the first derivative continuous at boundaries**. Finally, we apply the window function to each data and **use the average PSD to normalize the wave for different frequencies**.\n\n**For 1D model:**\n- Extend the signal to (3,8192)\n- Tukey window with alpha 0.5\n- Whitening with PSD based on the average FFT of all negative training examples\n\n**For 2D model:**\n- Using the whiten data and apply CQT transform with the following parameters:\n- CQT1992v2(sr=2048, fmin=20, fmax=1000, window='flattop', bins_per_octave=48, filter_scale=0.25, hop_length=8)\n- Resize (128x128, 256x256 …. ) \n\n![](https://i.ibb.co/CtL2nk6/whitening.png)\n### GW Simulation\nThis idea is coming from curriculum learning, and in [this paper](https://arxiv.org/abs/2106.03741), it mentioned that “We find that the deep learning algorithms can generalize **low signal-to-noise ratio (SNR) signals to high SNR ones but not vice versa**”, so we follow it and try to generate a signal and inject into the noise with low SNR. Even though there are around 15 parameters, we found that the most important one is the total mass and mass ratio (maybe we are wrong) because it affects the shape of GW the most through eyeballing.  So we adjust the total mass and mass ratio using different distributions and inject the signal into the noise with a given SNR following max(random.gauss(3.6,1),1) distribution. This SNR distribution is determined by checking the training loss trend: we want it to follow the trend of original data (not too hard, not too simple). This way gives us a **2~8bps increase** depending on the model we use.\nWe also tried to follow this idea by giving Hard Positive samples from the train data more weight but due to the time constraints, we didn’t make it work. It could be a potential win here.\n### Ensemble\nFirst, to confirm that train and test data are similar and do not have any hidden peculiarities we used **adversarial validation** that gave 0.5 AUC = train and test data are indistinguishable.\n\nWe tried many different methods to ensemble the models and saw the following relative performance trend: [CME-ES](https://github.com/CMA-ES/pycma) with Logit > CME-ES with rank> Scipy Optimization > Neural Network > other methods.  We also tried to use `sklearn.preprocessing.PolynomialFeatures` with probability prediction to do the CME-ES optimization. It brings the highest CV and LB but with a little chance of overfitting. We are glad that it turns out to be our best 0.8829 submission.\n\nOur second submission is based on a simulation of the private LB by excluding 16% of our OOF data out of CV optimization. So the produced model weights are more robust to the data noise and potentially can lead to a better performance at the private LB. We do this because we found the CV and LB correlation is very high and we also used adversarial validation to check that indeed they are similar. So We bootstrapped 16% OOF data which has a similar CV to public LB score and **the optimized CV for the remaining data matched the same as the private LB (0.8828 for one of our submissions)**.  \n### Acknowledgment\nSome of the team members used JarvisLabs.ai for their GPU workstations. The cloud platform is easy to use, stable, and very affordable. The pause functionality is great. Recommend this platform to Kagglers.",
      "votes": null
    },
    {
      "id": "1530200",
      "postDate": "10/01/2021 02:55:56",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> on becoming competition GM!</p>",
      "rawMarkdown": "Congratulations @iafoss on becoming competition GM!",
      "votes": null
    },
    {
      "id": "1530205",
      "postDate": "10/01/2021 02:59:21",
      "content": "<p>Thank you everyone! We will post our github repository here once it get cleaned and organized :D </p>",
      "rawMarkdown": "Thank you everyone! We will post our github repository here once it get cleaned and organized :D",
      "votes": null
    },
    {
      "id": "1530217",
      "postDate": "10/01/2021 03:17:13",
      "content": "<p>Thanks for the writeup. Congrats on becoming GM ! Looks like we missed out on the 1D models, should have tried that at the end</p>",
      "rawMarkdown": "Thanks for the writeup. Congrats on becoming GM ! Looks like we missed out on the 1D models, should have tried that at the end",
      "votes": null
    },
    {
      "id": "1530218",
      "postDate": "10/01/2021 03:17:25",
      "content": "<p>Thanks so much <a href=\"https://www.kaggle.com/richx86\" target=\"_blank\">@richx86</a> , it was a long journey for me but finally I reached this point. The time to post some datasets/random posts… just joking.<br>\nI was happy to work with you on this competition. It was really great teamwork, which at the end got rewarded.</p>",
      "rawMarkdown": "Thanks so much @richx86 , it was a long journey for me but finally I reached this point. The time to post some datasets/random posts... just joking.\nI was happy to work with you on this competition. It was really great teamwork, which at the end got rewarded.",
      "votes": null
    },
    {
      "id": "1530222",
      "postDate": "10/01/2021 03:21:27",
      "content": "<p>I was happy and honored to work with you! Great teamwork indeed! 1+1+1+1+1&gt;10  (not binary😊)</p>",
      "rawMarkdown": "I was happy and honored to work with you! Great teamwork indeed! 1+1+1+1+1>10  (not binary😊)",
      "votes": null
    },
    {
      "id": "1530228",
      "postDate": "10/01/2021 03:31:36",
      "content": "<p>Thanks so much. 1D models were really super effective here. Fortunately, ~1.5 weeks before the end of the competition we realized that proper architecture may easily give a quite high performance.<br>\nIt was quite a strange feeling when after ~12 hours of 1D architecture turning I got better performance than for 2D models after a month of experimenting with them… Architecture turning is really a cool thing. It was a way how I also got my first gold medal.</p>",
      "rawMarkdown": "Thanks so much. 1D models were really super effective here. Fortunately, ~1.5 weeks before the end of the competition we realized that proper architecture may easily give a quite high performance.\nIt was quite a strange feeling when after ~12 hours of 1D architecture turning I got better performance than for 2D models after a month of experimenting with them... Architecture turning is really a cool thing. It was a way how I also got my first gold medal.",
      "votes": null
    },
    {
      "id": "1530231",
      "postDate": "10/01/2021 03:35:38",
      "content": "<p>Thank you for the tips 😁 I actually started looking into 1D architectures in roughly the same time, in last week, but my final exams were going on and couldn't make time for the training :( Need to try this on some other competition.</p>",
      "rawMarkdown": "Thank you for the tips 😁 I actually started looking into 1D architectures in roughly the same time, in last week, but my final exams were going on and couldn't make time for the training :( Need to try this on some other competition.",
      "votes": null
    },
    {
      "id": "1530304",
      "postDate": "10/01/2021 05:01:58",
      "content": "<p>Congrats on becoming GM <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>. </p>",
      "rawMarkdown": "Congrats on becoming GM @iafoss.",
      "votes": null
    },
    {
      "id": "1530333",
      "postDate": "10/01/2021 05:28:17",
      "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/urvishp80\" target=\"_blank\">@urvishp80</a>. Unfortunately, we missed gold in the previous competition we worked togather, but eventually I received the last gold required for GM.</p>",
      "rawMarkdown": "Thank you so much @urvishp80. Unfortunately, we missed gold in the previous competition we worked togather, but eventually I received the last gold required for GM.",
      "votes": null
    },
    {
      "id": "1530335",
      "postDate": "10/01/2021 05:28:57",
      "content": "<p>Definitely!</p>",
      "rawMarkdown": "Definitely!",
      "votes": null
    },
    {
      "id": "1530338",
      "postDate": "10/01/2021 05:32:54",
      "content": "<p>yeah <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> but I am very happy that we worked together and now you are GM. There was a time when I used to fork your kernels to start the competition like Bengali.Ai one. So, seeing you becoming GM is really inspiring for me. </p>",
      "rawMarkdown": "yeah @iafoss but I am very happy that we worked together and now you are GM. There was a time when I used to fork your kernels to start the competition like Bengali.Ai one. So, seeing you becoming GM is really inspiring for me.",
      "votes": null
    },
    {
      "id": "1530499",
      "postDate": "10/01/2021 07:42:43",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>, <a href=\"https://www.kaggle.com/yamsam\" target=\"_blank\">@yamsam</a>, <a href=\"https://www.kaggle.com/vincentwang25\" target=\"_blank\">@vincentwang25</a> &amp; <a href=\"https://www.kaggle.com/richx86\" target=\"_blank\">@richx86</a>! This was one of the most fun competitions I've taken part in and I had a great time shooting for the top with you guys!</p>",
      "rawMarkdown": "Thank you @iafoss, @yamsam, @vincentwang25 & @richx86! This was one of the most fun competitions I've taken part in and I had a great time shooting for the top with you guys!",
      "votes": null
    },
    {
      "id": "1530575",
      "postDate": "10/01/2021 08:48:48",
      "content": "<p>Congratz <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> for finally reaching competitions GM, and to your entire team for the impressive finish !</p>",
      "rawMarkdown": "Congratz @iafoss for finally reaching competitions GM, and to your entire team for the impressive finish !",
      "votes": null
    },
    {
      "id": "1530916",
      "postDate": "10/01/2021 13:35:03",
      "content": "<p>Thanks you <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a>, likewise, it was really great working with you. </p>",
      "rawMarkdown": "Thanks you @anjum48, likewise, it was really great working with you.",
      "votes": null
    },
    {
      "id": "1530922",
      "postDate": "10/01/2021 13:39:55",
      "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>. Even if we got our first gold medals together, for me it took quite longer than for you to become a competition GM. But I finally made it.</p>",
      "rawMarkdown": "Thank you so much @theoviel. Even if we got our first gold medals together, for me it took quite longer than for you to become a competition GM. But I finally made it.",
      "votes": null
    },
    {
      "id": "1531209",
      "postDate": "10/01/2021 18:15:18",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> ! I had a great time working with you and I learned a lot from you!</p>",
      "rawMarkdown": "Thank you @anjum48 ! I had a great time working with you and I learned a lot from you!",
      "votes": null
    },
    {
      "id": "1531244",
      "postDate": "10/01/2021 18:57:49",
      "content": "<p>What  a great write-up with great diagrams and insights, thanks for sharing! It seems again that 1D models where key in this competition.</p>\n<p>Congratulations again on your 3rd place.</p>",
      "rawMarkdown": "What  a great write-up with great diagrams and insights, thanks for sharing! It seems again that 1D models where key in this competition.\n\nCongratulations again on your 3rd place.",
      "votes": null
    },
    {
      "id": "1531305",
      "postDate": "10/01/2021 20:14:52",
      "content": "<p>Thanks so much. 1D models were really super stars in this competitions.</p>",
      "rawMarkdown": "Thanks so much. 1D models were really super stars in this competitions.",
      "votes": null
    },
    {
      "id": "1531527",
      "postDate": "10/02/2021 05:18:54",
      "content": "<p>Wow. Great solution. Congratulations 👏👏 and a second congratulations for the well deserved grand status <a href=\"https://www.kaggle.com/lafoss\" target=\"_blank\">@lafoss</a> 👍🎉👏</p>",
      "rawMarkdown": "Wow. Great solution. Congratulations 👏👏 and a second congratulations for the well deserved grand status @lafoss 👍🎉👏",
      "votes": null
    },
    {
      "id": "1531540",
      "postDate": "10/02/2021 05:31:25",
      "content": "<p>Thank you so much!</p>",
      "rawMarkdown": "Thank you so much!",
      "votes": null
    },
    {
      "id": "1531866",
      "postDate": "10/02/2021 12:58:36",
      "content": "<p>Congratulations. I believe that proper use of domain knowledge and careful verification led to your results. Thanks for sharing your team's solution. And <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>, congratulations again on becoming competition GM!!</p>",
      "rawMarkdown": "Congratulations. I believe that proper use of domain knowledge and careful verification led to your results. Thanks for sharing your team's solution. And @iafoss, congratulations again on becoming competition GM!!",
      "votes": null
    },
    {
      "id": "1531875",
      "postDate": "10/02/2021 13:06:11",
      "content": "<p>Indeed, I wish I've given them more attention. 😄</p>",
      "rawMarkdown": "Indeed, I wish I've given them more attention. 😄",
      "votes": null
    },
    {
      "id": "1532376",
      "postDate": "10/02/2021 22:21:52",
      "content": "<p><a href=\"https://www.kaggle.com/mahmoudhelmy957\" target=\"_blank\">@mahmoudhelmy957</a>  What a great solution check it </p>",
      "rawMarkdown": "mahmoudhelmy957  What a great solution check it",
      "votes": null
    },
    {
      "id": "1532486",
      "postDate": "10/03/2021 03:29:47",
      "content": "<p><a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a>,  Thank you so much. Congratulations to u too on getting Competition Master, it's an important milestone. I hope you can get your GM title soon!</p>",
      "rawMarkdown": "naoism,  Thank you so much. Congratulations to u too on getting Competition Master, it's an important milestone. I hope you can get your GM title soon!",
      "votes": null
    },
    {
      "id": "1532770",
      "postDate": "10/03/2021 11:25:11",
      "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> congrats yo your team and thank you for the write up. Congrats to you for becoming a GM.  And congrats to <a href=\"https://www.kaggle.com/vincentwang25\" target=\"_blank\">@vincentwang25</a> and <a href=\"https://www.kaggle.com/richx86\" target=\"_blank\">@richx86</a> for their first gold metals.  </p>",
      "rawMarkdown": "iafoss congrats yo your team and thank you for the write up. Congrats to you for becoming a GM.  And congrats to @vincentwang25 and @richx86 for their first gold metals.",
      "votes": null
    },
    {
      "id": "1533056",
      "postDate": "10/03/2021 16:10:14",
      "content": "<p><a href=\"https://www.kaggle.com/yvonnef\" target=\"_blank\">@yvonnef</a>, thanks so much</p>",
      "rawMarkdown": "yvonnef, thanks so much",
      "votes": null
    },
    {
      "id": "1533065",
      "postDate": "10/03/2021 16:18:29",
      "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> out of curiosity what tool did you to create the diagram?  Thank you</p>",
      "rawMarkdown": "iafoss out of curiosity what tool did you to create the diagram?  Thank you",
      "votes": null
    },
    {
      "id": "1533143",
      "postDate": "10/03/2021 17:39:29",
      "content": "<p>Just MS PowerPoint… <br>\nDuring my PhD I realized that it is nearly the best tool to create images for papers/proposals.</p>",
      "rawMarkdown": "Just MS PowerPoint... \nDuring my PhD I realized that it is nearly the best tool to create images for papers/proposals.",
      "votes": null
    },
    {
      "id": "1533276",
      "postDate": "10/03/2021 20:47:09",
      "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> ty.  Wow !I never released power point could do that.  I will have to give that a whirl. </p>",
      "rawMarkdown": "iafoss ty.  Wow !I never released power point could do that.  I will have to give that a whirl.",
      "votes": null
    },
    {
      "id": "1533403",
      "postDate": "10/04/2021 01:55:09",
      "content": "<p><a href=\"https://www.kaggle.com/yvonnef\" target=\"_blank\">@yvonnef</a> Thank you very much!</p>",
      "rawMarkdown": "yvonnef Thank you very much!",
      "votes": null
    },
    {
      "id": "1535922",
      "postDate": "10/06/2021 10:19:36",
      "content": "<p>Congrats GM <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>  ! You definitely deserve it💪<br>\nand also thanks to all teammates! It was such a fun experience for me!</p>",
      "rawMarkdown": "Congrats GM @iafoss  ! You definitely deserve it💪\nand also thanks to all teammates! It was such a fun experience for me!",
      "votes": null
    },
    {
      "id": "1535927",
      "postDate": "10/06/2021 10:22:35",
      "content": "<p>Thank you, <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a>. It was an honor for me to work with you 👍</p>",
      "rawMarkdown": "Thank you, @anjum48. It was an honor for me to work with you 👍",
      "votes": null
    },
    {
      "id": "1536116",
      "postDate": "10/06/2021 13:28:31",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations",
      "votes": null
    },
    {
      "id": "1536657",
      "postDate": "10/06/2021 22:41:26",
      "content": "<p>Thanks      </p>",
      "rawMarkdown": "Thanks",
      "votes": null
    },
    {
      "id": "1536658",
      "postDate": "10/06/2021 22:42:53",
      "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/yamsam\" target=\"_blank\">@yamsam</a>, it was my pleasure to work with u on this challenge!</p>",
      "rawMarkdown": "Thank you so much @yamsam, it was my pleasure to work with u on this challenge!",
      "votes": null
    },
    {
      "id": "1543959",
      "postDate": "10/14/2021 01:31:23",
      "content": "<p>Our code is cleaned and gathered here: <a href=\"https://github.com/VincentWang25/G2Net_GoGoGo\" target=\"_blank\">https://github.com/VincentWang25/G2Net_GoGoGo</a> . Hope you find it helpful and please let us know if you have any question :) </p>",
      "rawMarkdown": "Our code is cleaned and gathered here: https://github.com/VincentWang25/G2Net_GoGoGo . Hope you find it helpful and please let us know if you have any question :)",
      "votes": null
    },
    {
      "id": "1544175",
      "postDate": "10/14/2021 05:58:26",
      "content": "<p>Thanks for sharing the source code!</p>",
      "rawMarkdown": "Thanks for sharing the source code!",
      "votes": null
    },
    {
      "id": "1548207",
      "postDate": "10/18/2021 04:38:40",
      "content": "<p>you are welcome!</p>",
      "rawMarkdown": "you are welcome!",
      "votes": null
    },
    {
      "id": "1559926",
      "postDate": "10/27/2021 08:23:51",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1530200,
      "author_name": "richx86",
      "author_url": "",
      "post_date": "10/01/2021 02:55:56",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> on becoming competition GM!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1530218,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/01/2021 03:17:25",
          "content": "<p>Thanks so much <a href=\"https://www.kaggle.com/richx86\" target=\"_blank\">@richx86</a> , it was a long journey for me but finally I reached this point. The time to post some datasets/random posts… just joking.<br>\nI was happy to work with you on this competition. It was really great teamwork, which at the end got rewarded.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1530222,
          "author_name": "richx86",
          "author_url": "",
          "post_date": "10/01/2021 03:21:27",
          "content": "<p>I was happy and honored to work with you! Great teamwork indeed! 1+1+1+1+1&gt;10  (not binary😊)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1530335,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/01/2021 05:28:57",
          "content": "<p>Definitely!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1530205,
      "author_name": "vincentwang25",
      "author_url": "",
      "post_date": "10/01/2021 02:59:21",
      "content": "<p>Thank you everyone! We will post our github repository here once it get cleaned and organized :D </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1530217,
      "author_name": "zarif98sjs",
      "author_url": "",
      "post_date": "10/01/2021 03:17:13",
      "content": "<p>Thanks for the writeup. Congrats on becoming GM ! Looks like we missed out on the 1D models, should have tried that at the end</p>",
      "votes": null,
      "replies": [
        {
          "id": 1530228,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/01/2021 03:31:36",
          "content": "<p>Thanks so much. 1D models were really super effective here. Fortunately, ~1.5 weeks before the end of the competition we realized that proper architecture may easily give a quite high performance.<br>\nIt was quite a strange feeling when after ~12 hours of 1D architecture turning I got better performance than for 2D models after a month of experimenting with them… Architecture turning is really a cool thing. It was a way how I also got my first gold medal.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1530231,
          "author_name": "zarif98sjs",
          "author_url": "",
          "post_date": "10/01/2021 03:35:38",
          "content": "<p>Thank you for the tips 😁 I actually started looking into 1D architectures in roughly the same time, in last week, but my final exams were going on and couldn't make time for the training :( Need to try this on some other competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1530304,
      "author_name": "urvishp80",
      "author_url": "",
      "post_date": "10/01/2021 05:01:58",
      "content": "<p>Congrats on becoming GM <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1530333,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/01/2021 05:28:17",
          "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/urvishp80\" target=\"_blank\">@urvishp80</a>. Unfortunately, we missed gold in the previous competition we worked togather, but eventually I received the last gold required for GM.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1530338,
          "author_name": "urvishp80",
          "author_url": "",
          "post_date": "10/01/2021 05:32:54",
          "content": "<p>yeah <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> but I am very happy that we worked together and now you are GM. There was a time when I used to fork your kernels to start the competition like Bengali.Ai one. So, seeing you becoming GM is really inspiring for me. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1530499,
      "author_name": "anjum48",
      "author_url": "",
      "post_date": "10/01/2021 07:42:43",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>, <a href=\"https://www.kaggle.com/yamsam\" target=\"_blank\">@yamsam</a>, <a href=\"https://www.kaggle.com/vincentwang25\" target=\"_blank\">@vincentwang25</a> &amp; <a href=\"https://www.kaggle.com/richx86\" target=\"_blank\">@richx86</a>! This was one of the most fun competitions I've taken part in and I had a great time shooting for the top with you guys!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1530916,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/01/2021 13:35:03",
          "content": "<p>Thanks you <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a>, likewise, it was really great working with you. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1531209,
          "author_name": "richx86",
          "author_url": "",
          "post_date": "10/01/2021 18:15:18",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> ! I had a great time working with you and I learned a lot from you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1535927,
          "author_name": "yamsam",
          "author_url": "",
          "post_date": "10/06/2021 10:22:35",
          "content": "<p>Thank you, <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a>. It was an honor for me to work with you 👍</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1530575,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "10/01/2021 08:48:48",
      "content": "<p>Congratz <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> for finally reaching competitions GM, and to your entire team for the impressive finish !</p>",
      "votes": null,
      "replies": [
        {
          "id": 1530922,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/01/2021 13:39:55",
          "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>. Even if we got our first gold medals together, for me it took quite longer than for you to become a competition GM. But I finally made it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1531244,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "10/01/2021 18:57:49",
      "content": "<p>What  a great write-up with great diagrams and insights, thanks for sharing! It seems again that 1D models where key in this competition.</p>\n<p>Congratulations again on your 3rd place.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1531305,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/01/2021 20:14:52",
          "content": "<p>Thanks so much. 1D models were really super stars in this competitions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1531875,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "10/02/2021 13:06:11",
          "content": "<p>Indeed, I wish I've given them more attention. 😄</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1531527,
      "author_name": "kalilurrahman",
      "author_url": "",
      "post_date": "10/02/2021 05:18:54",
      "content": "<p>Wow. Great solution. Congratulations 👏👏 and a second congratulations for the well deserved grand status <a href=\"https://www.kaggle.com/lafoss\" target=\"_blank\">@lafoss</a> 👍🎉👏</p>",
      "votes": null,
      "replies": [
        {
          "id": 1531540,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/02/2021 05:31:25",
          "content": "<p>Thank you so much!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1531866,
      "author_name": "naoism",
      "author_url": "",
      "post_date": "10/02/2021 12:58:36",
      "content": "<p>Congratulations. I believe that proper use of domain knowledge and careful verification led to your results. Thanks for sharing your team's solution. And <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>, congratulations again on becoming competition GM!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1532486,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/03/2021 03:29:47",
          "content": "<p><a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a>,  Thank you so much. Congratulations to u too on getting Competition Master, it's an important milestone. I hope you can get your GM title soon!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1532376,
      "author_name": "sharafcode",
      "author_url": "",
      "post_date": "10/02/2021 22:21:52",
      "content": "<p><a href=\"https://www.kaggle.com/mahmoudhelmy957\" target=\"_blank\">@mahmoudhelmy957</a>  What a great solution check it </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1532770,
      "author_name": "yvonnef",
      "author_url": "",
      "post_date": "10/03/2021 11:25:11",
      "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> congrats yo your team and thank you for the write up. Congrats to you for becoming a GM.  And congrats to <a href=\"https://www.kaggle.com/vincentwang25\" target=\"_blank\">@vincentwang25</a> and <a href=\"https://www.kaggle.com/richx86\" target=\"_blank\">@richx86</a> for their first gold metals.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1533056,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/03/2021 16:10:14",
          "content": "<p><a href=\"https://www.kaggle.com/yvonnef\" target=\"_blank\">@yvonnef</a>, thanks so much</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1533065,
          "author_name": "yvonnef",
          "author_url": "",
          "post_date": "10/03/2021 16:18:29",
          "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> out of curiosity what tool did you to create the diagram?  Thank you</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1533143,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/03/2021 17:39:29",
          "content": "<p>Just MS PowerPoint… <br>\nDuring my PhD I realized that it is nearly the best tool to create images for papers/proposals.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1533276,
          "author_name": "yvonnef",
          "author_url": "",
          "post_date": "10/03/2021 20:47:09",
          "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> ty.  Wow !I never released power point could do that.  I will have to give that a whirl. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1533403,
          "author_name": "richx86",
          "author_url": "",
          "post_date": "10/04/2021 01:55:09",
          "content": "<p><a href=\"https://www.kaggle.com/yvonnef\" target=\"_blank\">@yvonnef</a> Thank you very much!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1535922,
      "author_name": "yamsam",
      "author_url": "",
      "post_date": "10/06/2021 10:19:36",
      "content": "<p>Congrats GM <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>  ! You definitely deserve it💪<br>\nand also thanks to all teammates! It was such a fun experience for me!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1536658,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/06/2021 22:42:53",
          "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/yamsam\" target=\"_blank\">@yamsam</a>, it was my pleasure to work with u on this challenge!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1536116,
      "author_name": "adityakamal7",
      "author_url": "",
      "post_date": "10/06/2021 13:28:31",
      "content": "<p>Congratulations</p>",
      "votes": null,
      "replies": [
        {
          "id": 1536657,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/06/2021 22:41:26",
          "content": "<p>Thanks      </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1543959,
      "author_name": "vincentwang25",
      "author_url": "",
      "post_date": "10/14/2021 01:31:23",
      "content": "<p>Our code is cleaned and gathered here: <a href=\"https://github.com/VincentWang25/G2Net_GoGoGo\" target=\"_blank\">https://github.com/VincentWang25/G2Net_GoGoGo</a> . Hope you find it helpful and please let us know if you have any question :) </p>",
      "votes": null,
      "replies": [
        {
          "id": 1544175,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "10/14/2021 05:58:26",
          "content": "<p>Thanks for sharing the source code!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1548207,
          "author_name": "vincentwang25",
          "author_url": "",
          "post_date": "10/18/2021 04:38:40",
          "content": "<p>you are welcome!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1559926,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 08:23:51",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1530199": "### TL;DR\n##### Models we used in our best submission: 15 1D Model  + 8 2D Model\n**2D Models** [best single model:  **0.87875/0.8805/0.8787** at CV, public, and private LB]\n- CQT, CWT Transformation\n- Efficientnet, Resnext, InceptionV3...\n- TTA: shuffle LIGO channel\n\n**1D Models** [best single model: **0.8819/0.8827/0.8820** at CV, public, and private LB]\n- Customized architecture targeted at GW detection\n- TTA: vflip, shuffle LIGO channels, Gaussian noise, time shift, time mask, MC dropout\n\n**Training:**\n- Pretraining with GW\n- Training with Pseudo Label or Soft Label\n- BCE, Rank Loss\n- MixUp\n- AdamW,  RangerLars+Lookahead optimizer \n\n**Preprocessing:**\n- Avg PSD of target 0 (Design Curves)\n- Extending waves\n- Whitening with Tukey window\n\n**GW Simulation**\n- Distribution of Total Mass\n- SNR injection ratio max(N(3.6,1),1)\n\n**Ensemble**\n- CMA-ES Optimization\n- Hacking Private LB\n\n### Introduction\nOur team would like to thank organizers and kaggle for making this competition possible. Also, I want to express my gratitude to my outstanding teammates @vincentwang25, @richx86, @anjum48, and @yamsam for their incredible contribution towards our final result. In this competition I got the last gold medal required for getting kaggle GM, and also it is the first gold medal for @vincentwang25 and @richx86 . The provided write-up is the result of the cumulative efforts of all our team members, and I just got the privilege to post it.\n\n### Details \n#### 2D Models\nLike many of the participants, our team initially focused on 2D models. Among the various approaches we tried, we found that whitened signals, which we will discuss later, was the most effective preprocessing for signals. We created the following 2D models using the CQT and CWT images generated from the whitened signal as input.\n\nWe also tried various augmentations for 2D images (mixup, cutout, ...), but most of them did not work here, however mixup on the input waveform prior to CQT/CWT did help counteract overfitting. The most effective one was the augmentation of swapping LIGO signals. This worked both for training and inference (TTA, Test Time Augmentation) and we found soft pseudo labeling also improved the score. \n\nThe performance of our best single 2D model is **0.87875/0.8805/0.8787** at CV, public, and private LB. \n- EfficientNet(B3, B4, B5, B7), EfficientNetV2(M), ResNet200D, Inception-V3 (also we performed a number of initial experiments with ResNeXt models)\n- CQT and CWT images generated based on the whitened signal\n- Image size 128 x 128 〜 512 x 512\n- Soft leak-free pseudo labeling from ensemble results\n- LIGO channel swap argumentation (randomly swapping LIGO channels) for both training and TTA\n- 1D mixup prior to CQT/CWT\n- Adding a 4th channel to the spectrogram/scalogram input which is just a linear gradient (-1, 1) along the frequency dimension used for frequency encoding (similar to positional encoding in transformers)\n#### 1D Models\n1D models appeared to be the key component of our solution, even if we didn't make great improvement until the last 1.5 weeks of the competition. The reason why these models were not widely explored by most of the participants may be the need of using signal whitening to reach a performance comparable with 2D models (at least easily), and whitening is not straightforward for 2s long signals (see discussion below). However, 1D models are much faster to train, and they also outperform our 2D setup. The performance of our best single 1D model is **0.8819/0.8827/0.8820** at CV, public, and private LB. So **it can reach top-7 LB after ~8 hours of training**.\n\nOne of the main contributions towards this result is the proper choice of the model architecture for GW detection task. Specifically, **GW is not just a signal of the specific shape, but rather a correlation in signal between multiple detectors**. Meanwhile, signals may be shifted by up to ~10-20 ms because of the time needed for the signal to cross the distance between detectors. So direct concatenation of signals into (3,4096) stack and then applying a joined convolution is not a good idea (our baseline V0 architecture with CV of 0.8768). Random shift between components prohibits the generation of joined features. Thus, we asked the question, why not split the network into branches for different detectors, like proposed in [this paper](https://www.sciencedirect.com/science/article/pii/S0370269320308327)? So the first layers, extractor and following Res blocks, learn how to extract features from a signal, a kind of learnable FFT or wavelet transform. So before merging the signals the network already has a basic understanding of what is going on. We also share weights between LIGO1 and LIGO2 branches because the signals are very similar.\n![](https://i.ibb.co/4PgdmLY/1Dmodel.png)\n\nMerge of the extracted features instead of the original signal mitigates the effect of the relative shift (like in Short Time Fourier Transform correlation turns into a product of two aligned cells). So simple concatenation at this stage (instead of early concatenation) followed by several Res blocks (V1 architecture) gives an improvement from **0.8768 to 0.8789**. However, the model after combining the signal and getting a better idea about GW, may still want to look into individual branches as a reference. So we extend our individual branches and perform the second concatenation at the next convolutional block (V2 architecture). It results in a further improvement of CV to **0.8804**. \n\nAfter the basic structure of the V2 model was defined, we performed a number of additional experiments for further model optimization. The model architecture tricks giving further improvement include the use of SiLU instead of ReLU, use of GeM instead of strided convolution, use of concatenation pooling at the end of the convolutional part. In one of our final runs, we used ResNeSt blocks (Split Attention convolution) having a comparable performance in the preliminary experiments, but it also performed slightly worse at the end of full training. Using CBAM modules and Stochastic Depth modules gave a slight boost and made the 1D models more diversified. The network width, n, is equal to 16 and 32 for our final models. One of our experiments is also performed for a combined 1D+2D models, pretrained separately and then finetuned jointly, which gave 0.8815/0.8831/0.8817 score.\n\n**Things that didn’t work:** use of larger network depth (extra Res blocks), use of smaller/larger conv size in the extractor/first blocks, use of multi-head self-attention blocks before pooling or in the last stages of ResBlocks (i.e. Bottleneck Transformers), use of densely connected or ResNeXt blocks instead of ReBlocks, learnable CWT like extractors (FFT->multiplication by learnable weights->iFFT), WaveNet like network structures, pretrained ViT following the extractors (instead of customized resnet).\n\n##### Training (1D models)\n- Pretraining with Simulated GW for 2~4 epochs (**2-8bps** boost)\n- Training with Pseudo-labeling data for 4~6 epochs depending on the model\n- Training with rank loss and low lr for 2 epoch (**~1bps** performance boost)\n\nTTA based on multiplying the signal by -1 and first/second channel swap gave **~2bps** boost. 64 fold MC dropout gave an additional **~1bps** boost. Also during training in some experiments we used 0.05 spectral dropout: drop the given percentage of FFT spectrum during whitening.\n\n### Preprocessing\nRegarding whitening, direct use of packages, such as pycbc, doesn’t work mainly because of the short duration of the provided signals: only 2 seconds chunks in contrast to the almost unlimited length of data from LIGO/VIRGO detectors. To make the estimated PSD smooth, pycbc package uses an algorithm that corrupts the boundary of data, which is too costly for our dataset, whose duration of signals is only 2 seconds. We reduce the variance of estimated PSD by taking the average PSD for all negative training samples. This is the key to make whitening work (interestingly, this averaging idea came up independently to two of the team members). To further reduce the boundary effect from discontinuity and allow ourselves to use a window function that has a larger decay area (for example, Tukey window with a larger alpha), we first **extend the wave while keeping the first derivative continuous at boundaries**. Finally, we apply the window function to each data and **use the average PSD to normalize the wave for different frequencies**.\n\n**For 1D model:**\n- Extend the signal to (3,8192)\n- Tukey window with alpha 0.5\n- Whitening with PSD based on the average FFT of all negative training examples\n\n**For 2D model:**\n- Using the whiten data and apply CQT transform with the following parameters:\n- CQT1992v2(sr=2048, fmin=20, fmax=1000, window='flattop', bins_per_octave=48, filter_scale=0.25, hop_length=8)\n- Resize (128x128, 256x256 …. ) \n\n![](https://i.ibb.co/CtL2nk6/whitening.png)\n### GW Simulation\nThis idea is coming from curriculum learning, and in [this paper](https://arxiv.org/abs/2106.03741), it mentioned that “We find that the deep learning algorithms can generalize **low signal-to-noise ratio (SNR) signals to high SNR ones but not vice versa**”, so we follow it and try to generate a signal and inject into the noise with low SNR. Even though there are around 15 parameters, we found that the most important one is the total mass and mass ratio (maybe we are wrong) because it affects the shape of GW the most through eyeballing.  So we adjust the total mass and mass ratio using different distributions and inject the signal into the noise with a given SNR following max(random.gauss(3.6,1),1) distribution. This SNR distribution is determined by checking the training loss trend: we want it to follow the trend of original data (not too hard, not too simple). This way gives us a **2~8bps increase** depending on the model we use.\nWe also tried to follow this idea by giving Hard Positive samples from the train data more weight but due to the time constraints, we didn’t make it work. It could be a potential win here.\n### Ensemble\nFirst, to confirm that train and test data are similar and do not have any hidden peculiarities we used **adversarial validation** that gave 0.5 AUC = train and test data are indistinguishable.\n\nWe tried many different methods to ensemble the models and saw the following relative performance trend: [CME-ES](https://github.com/CMA-ES/pycma) with Logit > CME-ES with rank> Scipy Optimization > Neural Network > other methods.  We also tried to use `sklearn.preprocessing.PolynomialFeatures` with probability prediction to do the CME-ES optimization. It brings the highest CV and LB but with a little chance of overfitting. We are glad that it turns out to be our best 0.8829 submission.\n\nOur second submission is based on a simulation of the private LB by excluding 16% of our OOF data out of CV optimization. So the produced model weights are more robust to the data noise and potentially can lead to a better performance at the private LB. We do this because we found the CV and LB correlation is very high and we also used adversarial validation to check that indeed they are similar. So We bootstrapped 16% OOF data which has a similar CV to public LB score and **the optimized CV for the remaining data matched the same as the private LB (0.8828 for one of our submissions)**.  \n### Acknowledgment\nSome of the team members used JarvisLabs.ai for their GPU workstations. The cloud platform is easy to use, stable, and very affordable. The pause functionality is great. Recommend this platform to Kagglers.",
    "1530200": "Congratulations @iafoss on becoming competition GM!",
    "1530205": "Thank you everyone! We will post our github repository here once it get cleaned and organized :D",
    "1530217": "Thanks for the writeup. Congrats on becoming GM ! Looks like we missed out on the 1D models, should have tried that at the end",
    "1530218": "Thanks so much @richx86 , it was a long journey for me but finally I reached this point. The time to post some datasets/random posts... just joking.\nI was happy to work with you on this competition. It was really great teamwork, which at the end got rewarded.",
    "1530222": "I was happy and honored to work with you! Great teamwork indeed! 1+1+1+1+1>10  (not binary😊)",
    "1530228": "Thanks so much. 1D models were really super effective here. Fortunately, ~1.5 weeks before the end of the competition we realized that proper architecture may easily give a quite high performance.\nIt was quite a strange feeling when after ~12 hours of 1D architecture turning I got better performance than for 2D models after a month of experimenting with them... Architecture turning is really a cool thing. It was a way how I also got my first gold medal.",
    "1530231": "Thank you for the tips 😁 I actually started looking into 1D architectures in roughly the same time, in last week, but my final exams were going on and couldn't make time for the training :( Need to try this on some other competition.",
    "1530304": "Congrats on becoming GM @iafoss.",
    "1530333": "Thank you so much @urvishp80. Unfortunately, we missed gold in the previous competition we worked togather, but eventually I received the last gold required for GM.",
    "1530335": "Definitely!",
    "1530338": "yeah @iafoss but I am very happy that we worked together and now you are GM. There was a time when I used to fork your kernels to start the competition like Bengali.Ai one. So, seeing you becoming GM is really inspiring for me.",
    "1530499": "Thank you @iafoss, @yamsam, @vincentwang25 & @richx86! This was one of the most fun competitions I've taken part in and I had a great time shooting for the top with you guys!",
    "1530575": "Congratz @iafoss for finally reaching competitions GM, and to your entire team for the impressive finish !",
    "1530916": "Thanks you @anjum48, likewise, it was really great working with you.",
    "1530922": "Thank you so much @theoviel. Even if we got our first gold medals together, for me it took quite longer than for you to become a competition GM. But I finally made it.",
    "1531209": "Thank you @anjum48 ! I had a great time working with you and I learned a lot from you!",
    "1531244": "What  a great write-up with great diagrams and insights, thanks for sharing! It seems again that 1D models where key in this competition.\n\nCongratulations again on your 3rd place.",
    "1531305": "Thanks so much. 1D models were really super stars in this competitions.",
    "1531527": "Wow. Great solution. Congratulations 👏👏 and a second congratulations for the well deserved grand status @lafoss 👍🎉👏",
    "1531540": "Thank you so much!",
    "1531866": "Congratulations. I believe that proper use of domain knowledge and careful verification led to your results. Thanks for sharing your team's solution. And @iafoss, congratulations again on becoming competition GM!!",
    "1531875": "Indeed, I wish I've given them more attention. 😄",
    "1532376": "mahmoudhelmy957  What a great solution check it",
    "1532486": "naoism,  Thank you so much. Congratulations to u too on getting Competition Master, it's an important milestone. I hope you can get your GM title soon!",
    "1532770": "iafoss congrats yo your team and thank you for the write up. Congrats to you for becoming a GM.  And congrats to @vincentwang25 and @richx86 for their first gold metals.",
    "1533056": "yvonnef, thanks so much",
    "1533065": "iafoss out of curiosity what tool did you to create the diagram?  Thank you",
    "1533143": "Just MS PowerPoint... \nDuring my PhD I realized that it is nearly the best tool to create images for papers/proposals.",
    "1533276": "iafoss ty.  Wow !I never released power point could do that.  I will have to give that a whirl.",
    "1533403": "yvonnef Thank you very much!",
    "1535922": "Congrats GM @iafoss  ! You definitely deserve it💪\nand also thanks to all teammates! It was such a fun experience for me!",
    "1535927": "Thank you, @anjum48. It was an honor for me to work with you 👍",
    "1536116": "Congratulations",
    "1536657": "Thanks",
    "1536658": "Thank you so much @yamsam, it was my pleasure to work with u on this challenge!",
    "1543959": "Our code is cleaned and gathered here: https://github.com/VincentWang25/G2Net_GoGoGo . Hope you find it helpful and please let us know if you have any question :)",
    "1544175": "Thanks for sharing the source code!",
    "1548207": "you are welcome!",
    "1559926": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}