{
  "id": 275356,
  "title": "12th place solution [SiN Nakaism924]",
  "url": "/competitions/g2net-gravitational-wave-detection/writeups/sin-nakaism924-12th-place-solution-sin-nakaism924",
  "author_name": "",
  "post_date": "2021-09-30T23:46:10.457Z",
  "votes": 27,
  "comment_count": 11,
  "views": 0,
  "content": "<p>First of all, I would like to thank my teammates( <a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a>, <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a>, <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>, <a href=\"https://www.kaggle.com/sinpcw\" target=\"_blank\">@sinpcw</a> ) for competing with me. I would also like to thank kaggle and EGO for organizing an interesting competition.<br>\nI explain my approach. My teammates will add their own approaches in the comments!</p>\n<ul>\n<li>Y.Nakama part ( <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529297\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529297</a> )</li>\n<li>SiNpcw part ( <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529331\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529331</a> )</li>\n<li>Naoism part ( <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529412\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529412</a> )</li>\n<li>Hidehisa part ( <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529639\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529639</a> )</li>\n</ul>\n<h2>hirune part intro</h2>\n<p>I joined this competition after the SETI competition was over. This task is very similar to the SETI competition task. This task is very similar to the SETI competition task, so I reused most of the SETI code.</p>\n<h3>preprocessing</h3>\n<ol>\n<li>Divide the all wave form by 4.6152116213830774e-20 (max(abs) of the entire train and test data).</li>\n<li>Use nnAudio to run CQT. Randomly select one of flattop, blackmanharris, or nuttall to the window. This can be used in TTA. Finally, I also added CWT to this.</li>\n<li>The spectrograms are combined in the frequency direction and input to the model as a single-channel image</li>\n<li>Resize to 384x512 before inputting into the model, and normalize the spectrogram of the entire data set in mean, std</li>\n</ol>\n<h3>Augmentation</h3>\n<ul>\n<li>Before converting to a spectrogram, mixup as follows.</li>\n</ul>\n<pre><code>x = x1 + x2\ny = y1 + y2 -(y1*y2)\n</code></pre>\n<ul>\n<li>Randomly roll shift in time direction</li>\n</ul>\n<h3>train setup</h3>\n<ul>\n<li>model: timm tf_efficientnet_b4_ap</li>\n<li>optimizer: Adam</li>\n<li>Lr: 0.001</li>\n<li>scheduler: CosineAnnealingLR</li>\n</ul>\n<h3>pseudo label</h3>\n<p>Using pseudo-labels showed some improvement, but not a lot.</p>\n<h3>stacking</h3>\n<p>Using other team members models, and finally stacking 137 models with NN or XGB, there was a significant score increase. Other models include 1dcnn, swin transformer and efficientnet b5-8.<br>\nTeam members will explain these details later.</p>",
  "messages": [
    {
      "id": "1528973",
      "postDate": "09/30/2021 04:01:50",
      "content": "<p>First of all, I would like to thank my teammates( <a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a>, <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a>, <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>, <a href=\"https://www.kaggle.com/sinpcw\" target=\"_blank\">@sinpcw</a> ) for competing with me. I would also like to thank kaggle and EGO for organizing an interesting competition.<br>\nI explain my approach. My teammates will add their own approaches in the comments!</p>\n<ul>\n<li>Y.Nakama part ( <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529297\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529297</a> )</li>\n<li>SiNpcw part ( <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529331\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529331</a> )</li>\n<li>Naoism part ( <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529412\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529412</a> )</li>\n<li>Hidehisa part ( <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529639\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529639</a> )</li>\n</ul>\n<h2>hirune part intro</h2>\n<p>I joined this competition after the SETI competition was over. This task is very similar to the SETI competition task. This task is very similar to the SETI competition task, so I reused most of the SETI code.</p>\n<h3>preprocessing</h3>\n<ol>\n<li>Divide the all wave form by 4.6152116213830774e-20 (max(abs) of the entire train and test data).</li>\n<li>Use nnAudio to run CQT. Randomly select one of flattop, blackmanharris, or nuttall to the window. This can be used in TTA. Finally, I also added CWT to this.</li>\n<li>The spectrograms are combined in the frequency direction and input to the model as a single-channel image</li>\n<li>Resize to 384x512 before inputting into the model, and normalize the spectrogram of the entire data set in mean, std</li>\n</ol>\n<h3>Augmentation</h3>\n<ul>\n<li>Before converting to a spectrogram, mixup as follows.</li>\n</ul>\n<pre><code>x = x1 + x2\ny = y1 + y2 -(y1*y2)\n</code></pre>\n<ul>\n<li>Randomly roll shift in time direction</li>\n</ul>\n<h3>train setup</h3>\n<ul>\n<li>model: timm tf_efficientnet_b4_ap</li>\n<li>optimizer: Adam</li>\n<li>Lr: 0.001</li>\n<li>scheduler: CosineAnnealingLR</li>\n</ul>\n<h3>pseudo label</h3>\n<p>Using pseudo-labels showed some improvement, but not a lot.</p>\n<h3>stacking</h3>\n<p>Using other team members models, and finally stacking 137 models with NN or XGB, there was a significant score increase. Other models include 1dcnn, swin transformer and efficientnet b5-8.<br>\nTeam members will explain these details later.</p>",
      "rawMarkdown": "First of all, I would like to thank my teammates( @naoism, @hidehisaarai1213, @yasufuminakama, @sinpcw ) for competing with me. I would also like to thank kaggle and EGO for organizing an interesting competition.\nI explain my approach. My teammates will add their own approaches in the comments!\n* Y.Nakama part ( https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529297 )\n* SiNpcw part ( https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529331 )\n* Naoism part ( https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529412 )\n* Hidehisa part ( https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529639 )\n\n## hirune part intro\nI joined this competition after the SETI competition was over. This task is very similar to the SETI competition task. This task is very similar to the SETI competition task, so I reused most of the SETI code.\n\n### preprocessing\n1. Divide the all wave form by 4.6152116213830774e-20 (max(abs) of the entire train and test data).\n2. Use nnAudio to run CQT. Randomly select one of flattop, blackmanharris, or nuttall to the window. This can be used in TTA. Finally, I also added CWT to this.\n3. The spectrograms are combined in the frequency direction and input to the model as a single-channel image\n4. Resize to 384x512 before inputting into the model, and normalize the spectrogram of the entire data set in mean, std\n\n### Augmentation\n* Before converting to a spectrogram, mixup as follows.\n```\nx = x1 + x2\ny = y1 + y2 -(y1*y2)\n```\n* Randomly roll shift in time direction\n\n### train setup\n* model: timm tf_efficientnet_b4_ap\n* optimizer: Adam\n* Lr: 0.001\n* scheduler: CosineAnnealingLR\n\n### pseudo label\nUsing pseudo-labels showed some improvement, but not a lot.\n\n### stacking\nUsing other team members models, and finally stacking 137 models with NN or XGB, there was a significant score increase. Other models include 1dcnn, swin transformer and efficientnet b5-8.\nTeam members will explain these details later.",
      "votes": null
    },
    {
      "id": "1528994",
      "postDate": "09/30/2021 04:39:19",
      "content": "<p><a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a>, I'm glad that you eventually made it! after we missed gold in the previous competition. Big congratulations to u and your team.</p>",
      "rawMarkdown": "naoism, I'm glad that you eventually made it! after we missed gold in the previous competition. Big congratulations to u and your team.",
      "votes": null
    },
    {
      "id": "1529051",
      "postDate": "09/30/2021 05:48:04",
      "content": "<p>congratulation！  137 models!!!!</p>",
      "rawMarkdown": "congratulation！  137 models!!!!",
      "votes": null
    },
    {
      "id": "1529060",
      "postDate": "09/30/2021 05:59:04",
      "content": "<p>Nice to see that it required you 137 models to beat us 🤣<br>\nGood job for the gold 💪</p>",
      "rawMarkdown": "Nice to see that it required you 137 models to beat us 🤣\nGood job for the gold 💪",
      "votes": null
    },
    {
      "id": "1529071",
      "postDate": "09/30/2021 06:17:57",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hirune924\" target=\"_blank\">@hirune924</a> , thanks for sharing the solution , looks like a very elegant one</p>\n<p><code>x = x1 + x2\ny = y1 + y2 -(y1*y2)</code></p>\n<p>Can you report the score increase due to this , we also tried something on similar terms but didn't get a lot out of it</p>",
      "rawMarkdown": "Hi @hirune924 , thanks for sharing the solution , looks like a very elegant one\n\n`x = x1 + x2\ny = y1 + y2 -(y1*y2)`\n\nCan you report the score increase due to this , we also tried something on similar terms but didn't get a lot out of it",
      "votes": null
    },
    {
      "id": "1529095",
      "postDate": "09/30/2021 06:41:08",
      "content": "<p>I don't know the exact difference in scores from regular mixup because I used that trick from the beginning, but using mixup prevented overfitting during long training.<br>\nAlso, considering the possibility that this trick was negatively affecting the training, I reduced the frequency of its use from 50% to 15%, and the cv score decreased by about 0.0002.</p>",
      "rawMarkdown": "I don't know the exact difference in scores from regular mixup because I used that trick from the beginning, but using mixup prevented overfitting during long training.\nAlso, considering the possibility that this trick was negatively affecting the training, I reduced the frequency of its use from 50% to 15%, and the cv score decreased by about 0.0002.",
      "votes": null
    },
    {
      "id": "1529297",
      "postDate": "09/30/2021 09:25:25",
      "content": "<p>First of all, I would like to thank my teammates for competing with me.<br>\nAnd I would like to thank kaggle &amp; host for the interesting competition and to all the participants for giving us a lot of ideas.</p>\n<h1>Y.Nakama part</h1>\n<p>Here is my approach.</p>\n<h2>train setup</h2>\n<ul>\n<li>main model: CWT + tf_efficientnet_b5</li>\n<li>lr: 1e-3</li>\n<li>scheduler: CosineAnnealingLR</li>\n<li>optimizer: Adam (better than AdamW in my experiments)</li>\n</ul>\n<h2>preprocessing</h2>\n<ul>\n<li>z-score normalization</li>\n</ul>\n<pre><code>MEAN_ALL = [6.901084819126393e-26, 5.117726793550612e-26, -1.3831247922595496e-26]\nSTD_ALL = [7.420282939705547e-21, 7.41993950281645e-21, 1.838329278493792e-21]\n\ndef z_score_normalization_wave(_wave, site):\n    x_mean = MEAN_ALL[site]\n    x_std = STD_ALL[site]\n    wave = (_wave - x_mean) / (x_std)\n    return wave\n</code></pre>\n<p>No gain compared to dividing the all wave form by 4.6152116213830774e-20 (max(abs) of the entire train and test data).</p>\n<ul>\n<li>What didn't work for me: highpass filter + lowpassfilter + tukey window</li>\n</ul>\n<h2>Augmentation</h2>\n<ul>\n<li>mixup</li>\n</ul>\n<pre><code>def mixup(image, label, max_noise=0.5, probability=1.0):\n    if np.random.uniform(0, 1) &lt;= probability:\n        bs, c, h, w = image.size()\n        indices = torch.randperm(bs)\n        shuffled_image = image[indices]\n        shuffled_label = label[indices]\n        lam = np.random.uniform(0, 1) * (1.0 - shuffled_label) * max_noise # don't mix shuffled_label=1\n        lam = torch.flatten(lam)\n        image = (image.reshape(-1, bs) * (1 - lam)).reshape(bs, c, h, w) + (shuffled_image.reshape(-1, bs) * lam).reshape(bs, c, h, w)\n    return image, label\n</code></pre>\n<ul>\n<li>time shift with centercrop</li>\n</ul>\n<pre><code>w_diff = int(torch.randint(low=0, high=5, size=(1,)))\nif torch.rand(1) &lt; 0.5:\n    x = x[:,:,self.h_div_start:self.h_div_end,self.w_div_start+w_diff:self.w_div_end+w_diff]\nelse:\n    x = x[:,:,self.h_div_start:self.h_div_end,self.w_div_start-w_diff:self.w_div_end-w_diff]\n</code></pre>\n<ul>\n<li>What didn't work for me: gaussian noise</li>\n</ul>\n<h2>pseudo label</h2>\n<p>Re-training on soft pseudo-labelled test dataset improved score.</p>\n<h2>TTA</h2>\n<ul>\n<li>time shift with centercrop</li>\n</ul>\n<h2>stacking</h2>\n<p>I implemented stacking models as I did in <a href=\"https://www.kaggle.com/c/stanford-covid-vaccine/discussion/189728\" target=\"_blank\">OpenVaccine competition</a>.<br>\nNN stacking model performed better than GBDT stacking models.</p>",
      "rawMarkdown": "First of all, I would like to thank my teammates for competing with me.\nAnd I would like to thank kaggle & host for the interesting competition and to all the participants for giving us a lot of ideas.\n\n# Y.Nakama part \nHere is my approach.\n\n## train setup\n- main model: CWT + tf_efficientnet_b5\n- lr: 1e-3\n- scheduler: CosineAnnealingLR\n- optimizer: Adam (better than AdamW in my experiments)\n\n## preprocessing\n- z-score normalization\n```\nMEAN_ALL = [6.901084819126393e-26, 5.117726793550612e-26, -1.3831247922595496e-26]\nSTD_ALL = [7.420282939705547e-21, 7.41993950281645e-21, 1.838329278493792e-21]\n\ndef z_score_normalization_wave(_wave, site):\n    x_mean = MEAN_ALL[site]\n    x_std = STD_ALL[site]\n    wave = (_wave - x_mean) / (x_std)\n    return wave\n```\nNo gain compared to dividing the all wave form by 4.6152116213830774e-20 (max(abs) of the entire train and test data).\n- What didn't work for me: highpass filter + lowpassfilter + tukey window\n\n## Augmentation\n- mixup\n```\ndef mixup(image, label, max_noise=0.5, probability=1.0):\n    if np.random.uniform(0, 1) <= probability:\n        bs, c, h, w = image.size()\n        indices = torch.randperm(bs)\n        shuffled_image = image[indices]\n        shuffled_label = label[indices]\n        lam = np.random.uniform(0, 1) * (1.0 - shuffled_label) * max_noise # don't mix shuffled_label=1\n        lam = torch.flatten(lam)\n        image = (image.reshape(-1, bs) * (1 - lam)).reshape(bs, c, h, w) + (shuffled_image.reshape(-1, bs) * lam).reshape(bs, c, h, w)\n    return image, label\n```\n- time shift with centercrop\n```\nw_diff = int(torch.randint(low=0, high=5, size=(1,)))\nif torch.rand(1) < 0.5:\n    x = x[:,:,self.h_div_start:self.h_div_end,self.w_div_start+w_diff:self.w_div_end+w_diff]\nelse:\n    x = x[:,:,self.h_div_start:self.h_div_end,self.w_div_start-w_diff:self.w_div_end-w_diff]\n```\n- What didn't work for me: gaussian noise\n\n## pseudo label\nRe-training on soft pseudo-labelled test dataset improved score.\n\n## TTA\n- time shift with centercrop\n\n## stacking\nI implemented stacking models as I did in [OpenVaccine competition](https://www.kaggle.com/c/stanford-covid-vaccine/discussion/189728).\nNN stacking model performed better than GBDT stacking models.",
      "votes": null
    },
    {
      "id": "1529317",
      "postDate": "09/30/2021 09:41:22",
      "content": "<p>Thank you very much. I was finally able to achieve my goal of becoming a Master. I'm really happy.<br>\nAlso, as a team member of the previous competition, I am happy that you finally become a GM. Congratulations!</p>",
      "rawMarkdown": "Thank you very much. I was finally able to achieve my goal of becoming a Master. I'm really happy.\nAlso, as a team member of the previous competition, I am happy that you finally become a GM. Congratulations!",
      "votes": null
    },
    {
      "id": "1529331",
      "postDate": "09/30/2021 09:57:39",
      "content": "<p>Thank you to all the participants for their hard work. <br>\nThank you to the organizers, hosts and teammates.</p>\n<h2>SiNpcw part:</h2>\n<p>Up to a week before the competition, we were working on 2d CNN, but in the final stage, I was looking for improvements to 1d CNN.</p>\n<h2>2d CNN:</h2>\n<ul>\n<li>model: efficentnet_v2_m</li>\n<li>optimizer: Adam</li>\n<li>lr: 1e-3</li>\n<li>scheduler: CosineAnnealingLR</li>\n<li>image size: 138x513</li>\n</ul>\n<p>It is almost the same preprocessing as <a href=\"https://www.kaggle.com/hirune924\" target=\"_blank\">@hirune924</a> but differs in a few details.<br>\nI applied 1dGridDistortion (only in the frequency direction) as a unique element. It has slightly improved the CV.<br>\nIn training the model, we used only nuttall window among the CQTs. This is because nuttall was the best window for CV.<br>\nI tried deleting data that was prone to miss-prediction, but not work.</p>\n<h2>1d CNN:</h2>\n<p>Improvements were made based on the (<a href=\"https://www.kaggle.com/scaomath/g2net-1d-cnn-gem-pool-pytorch-train-inference\" target=\"_blank\">published notebook</a>).<br>\nI used a strategy that of a noisy student, creating pseudo-labels in my own model and learning them again.  <br>\nDue to deadlines, I was not able to try enough of them, but the following ones worked well.</p>\n<ul>\n<li>SAM optimizer</li>\n<li>adjusting bandpass parameters</li>\n<li>ensemble: augmented and non-augmented models</li>\n</ul>",
      "rawMarkdown": "Thank you to all the participants for their hard work. \nThank you to the organizers, hosts and teammates.\n\n## SiNpcw part:\nUp to a week before the competition, we were working on 2d CNN, but in the final stage, I was looking for improvements to 1d CNN.\n\n## 2d CNN:\n- model: efficentnet_v2_m\n- optimizer: Adam\n- lr: 1e-3\n- scheduler: CosineAnnealingLR\n- image size: 138x513\n  \nIt is almost the same preprocessing as @hirune924 but differs in a few details.\nI applied 1dGridDistortion (only in the frequency direction) as a unique element. It has slightly improved the CV.\nIn training the model, we used only nuttall window among the CQTs. This is because nuttall was the best window for CV.\nI tried deleting data that was prone to miss-prediction, but not work.\n\n## 1d CNN:\nImprovements were made based on the ([published notebook](https://www.kaggle.com/scaomath/g2net-1d-cnn-gem-pool-pytorch-train-inference)).\nI used a strategy that of a noisy student, creating pseudo-labels in my own model and learning them again.  \nDue to deadlines, I was not able to try enough of them, but the following ones worked well.\n- SAM optimizer\n- adjusting bandpass parameters\n- ensemble: augmented and non-augmented models",
      "votes": null
    },
    {
      "id": "1529412",
      "postDate": "09/30/2021 11:28:53",
      "content": "<p>At first, I would like to express my thanks to all the organizers for hosting this interesting competition and my teammates.<br>\nAnd thanks to all other competitors for sharing valuable information in forums and notebooks.</p>\n<p>I was finally able to achieve my goal of becoming a Master :)</p>\n<h1>Naoism part</h1>\n<p>My solution is very similar to <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>'s one.<br>\nSo I won't explain each method in detail.</p>\n<h2>train setup</h2>\n<ul>\n<li>models<ul>\n<li>2D model: CQT + EfficientNet, DenseNet and ResNet</li>\n<li>1D model: Simple 1D-CNN</li></ul></li>\n<li>Scheduler: CosineAnnealingLR</li>\n<li>Epoch: 8~10</li>\n<li>Optimizer = AdamW (better than Adam and RAdam) for 2Dmodel and SAM(better than AdamW) for 1D-CNN</li>\n<li>image size = (128, 768)</li>\n</ul>\n<h2>preprocessing</h2>\n<ul>\n<li>pycbc package's highpass filter(32 &lt; f) and notchfilter(30 &lt; f &lt; 80)</li>\n<li>Tukey window</li>\n<li>Z-score normalization</li>\n</ul>\n<pre><code>def z_score_normalization_wave(_wave, site):\n    MEAN_ALL = [-7.687203989766856e-31, 1.5238509999790067e-30, -1.7078072503529966e-30]\n    STD_ALL = [1.618816429314542e-22, 1.6188170892087954e-22, 2.621679130415917e-22]\n    x_mean = MEAN_ALL[site]\n    x_std = STD_ALL[site]\n    wave = (_wave-x_mean) / (x_std)\n    return wave\n\ndef load_wave(path):\n    waves = np.load(path)\n    new_waves = []\n    window = signal.tukey(4096)\n    for i in [0,1,2]:\n        _wave = pycbc.filter.resample.highpass_fir(pycbc.types.TimeSeries(waves[i], epoch=0, delta_t=1.0/2048), frequency=32, order=100)\n        _wave = pycbc.filter.resample.notch_fir(_wave, f1=30, f2=80, order=10, beta=5)\n        _wave = np.array(_wave)\n        _wave = _wave * window\n        wave = z_score_normalization_wave(_wave, site=i)\n        new_waves.append(wave)\n    return np.array(new_waves)\n</code></pre>\n<ul>\n<li>Centercrop(0.4s ~ 1.8s)</li>\n</ul>\n<h2>Augmentation</h2>\n<ul>\n<li>Negative label mixup for image(p=1.0) (Almost the same as the one introduced by <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>)</li>\n<li>Time shift with centercrop</li>\n<li>GaussianNoise(p=0.3, max_SNR = 0.2)</li>\n<li>SpecAugmentation(Frequency mask and Time mask)</li>\n</ul>\n<h2>Other ideas</h2>\n<ul>\n<li>Using the other model's oof as the target improved the CV considerably.</li>\n<li>Also, repeating the process of getting the predictions of a model-&gt;stacking with multiple models-&gt;retraining with oof as target slightly improved the CV.</li>\n</ul>",
      "rawMarkdown": "At first, I would like to express my thanks to all the organizers for hosting this interesting competition and my teammates.\nAnd thanks to all other competitors for sharing valuable information in forums and notebooks.\n\nI was finally able to achieve my goal of becoming a Master :)\n\n# Naoism part\nMy solution is very similar to @yasufuminakama's one.\nSo I won't explain each method in detail.\n\n## train setup\n- models\n  - 2D model: CQT + EfficientNet, DenseNet and ResNet\n  - 1D model: Simple 1D-CNN\n- Scheduler: CosineAnnealingLR\n- Epoch: 8~10\n- Optimizer = AdamW (better than Adam and RAdam) for 2Dmodel and SAM(better than AdamW) for 1D-CNN\n- image size = (128, 768)\n\n## preprocessing\n- pycbc package's highpass filter(32 < f) and notchfilter(30 < f < 80)\n- Tukey window\n- Z-score normalization\n\n```\ndef z_score_normalization_wave(_wave, site):\n    MEAN_ALL = [-7.687203989766856e-31, 1.5238509999790067e-30, -1.7078072503529966e-30]\n    STD_ALL = [1.618816429314542e-22, 1.6188170892087954e-22, 2.621679130415917e-22]\n    x_mean = MEAN_ALL[site]\n    x_std = STD_ALL[site]\n    wave = (_wave-x_mean) / (x_std)\n    return wave\n\ndef load_wave(path):\n    waves = np.load(path)\n    new_waves = []\n    window = signal.tukey(4096)\n    for i in [0,1,2]:\n        _wave = pycbc.filter.resample.highpass_fir(pycbc.types.TimeSeries(waves[i], epoch=0, delta_t=1.0/2048), frequency=32, order=100)\n        _wave = pycbc.filter.resample.notch_fir(_wave, f1=30, f2=80, order=10, beta=5)\n        _wave = np.array(_wave)\n        _wave = _wave * window\n        wave = z_score_normalization_wave(_wave, site=i)\n        new_waves.append(wave)\n    return np.array(new_waves)\n```\n- Centercrop(0.4s ~ 1.8s)\n\n## Augmentation\n\n- Negative label mixup for image(p=1.0) (Almost the same as the one introduced by @yasufuminakama)\n- Time shift with centercrop\n- GaussianNoise(p=0.3, max_SNR = 0.2)\n- SpecAugmentation(Frequency mask and Time mask)\n\n## Other ideas\n- Using the other model's oof as the target improved the CV considerably.\n- Also, repeating the process of getting the predictions of a model->stacking with multiple models->retraining with oof as target slightly improved the CV.",
      "votes": null
    },
    {
      "id": "1529639",
      "postDate": "09/30/2021 14:46:22",
      "content": "<h2>Hidehisa's Part</h2>\n<p>I and <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> and <a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a> have formed a team from comparatively early stage of the competition, so our approaches were quite similar each other, so you may find some duplicates between what I've done and what others have done.</p>\n<p>Here's the list that I've done and those worked.</p>\n<ul>\n<li>TensorFlow pipeline (maybe good in terms of diversity)<ul>\n<li>I was the only one in the team who used TF pipeline until the end</li></ul></li>\n<li>Use statistics of the whole dataset to normalize each samples (not the statistics of each samples!)<ul>\n<li>Separately reported as effective by Naoism and hirune's team (and indeed quite effective)</li></ul></li>\n<li>CWT was almost as good as CQT <ul>\n<li>I used the implementation of <a href=\"https://github.com/Kevin-McIsaac/cmorlet-tensorflow\" target=\"_blank\">https://github.com/Kevin-McIsaac/cmorlet-tensorflow</a></li>\n<li>I've also  implemented PyTorch version of CWT. This implementation was used through the team (code below).</li></ul></li>\n<li>Found Negative label mixup for image(p=1.0) effective (which is described in Y.Nakama's part and Naoism's part)</li>\n<li>Found Center Crop along with the time axis effective (also described in Y.Nakama's part and Naoism's part)</li>\n<li>Found training with OOF somewhat effective (also described in Naoism's part)<ul>\n<li>We thought this was because low SNR signals can be thought as samples with noisy labels. Does someone tried the same thing? (approaches for training with noisy labels settings)</li>\n<li>Our experiment was a bit leaky, so I'd like to check how much it was effective</li></ul></li>\n<li>SwinTransformer Model (used the implementation of <a href=\"https://github.com/rishigami/Swin-Transformer-TF\" target=\"_blank\">https://github.com/rishigami/Swin-Transformer-TF</a>)<ul>\n<li>Not very good, but I think it contributed a bit in terms of diversity</li></ul></li>\n</ul>\n<h3>CWT implementation in pytorch</h3>\n<pre><code>class CWT(nn.Module):\n    def __init__(\n        self,\n        wavelet_width,\n        fs,\n        lower_freq,\n        upper_freq,\n        n_scales,\n        size_factor=1.0,\n        border_crop=0,\n        stride=1\n    ):\n        super().__init__()\n\n        self.initial_wavelet_width = wavelet_width\n        self.fs = fs\n        self.lower_freq = lower_freq\n        self.upper_freq = upper_freq\n        self.size_factor = size_factor\n        self.n_scales = n_scales\n        self.wavelet_width = wavelet_width\n        self.border_crop = border_crop\n        self.stride = stride\n        wavelet_bank_real, wavelet_bank_imag = self._build_wavelet_kernel()\n        self.wavelet_bank_real = nn.Parameter(wavelet_bank_real, requires_grad=False)\n        self.wavelet_bank_imag = nn.Parameter(wavelet_bank_imag, requires_grad=False)\n\n        self.kernel_size = self.wavelet_bank_real.size(3)\n\n    def _build_wavelet_kernel(self):\n        s_0 = 1 / self.upper_freq\n        s_n = 1 / self.lower_freq\n\n        base = np.power(s_n / s_0, 1 / (self.n_scales - 1))\n        scales = s_0 * np.power(base, np.arange(self.n_scales))\n\n        frequencies = 1 / scales\n        truncation_size = scales.max() * np.sqrt(4.5 * self.initial_wavelet_width) * self.fs\n        one_side = int(self.size_factor * truncation_size)\n        kernel_size = 2 * one_side + 1\n\n        k_array = np.arange(kernel_size, dtype=np.float32) - one_side\n        t_array = k_array / self.fs\n\n        wavelet_bank_real = []\n        wavelet_bank_imag = []\n\n        for scale in scales:\n            norm_constant = np.sqrt(np.pi * self.wavelet_width) * scale * self.fs / 2.0\n            scaled_t = t_array / scale\n            exp_term = np.exp(-(scaled_t ** 2) / self.wavelet_width)\n            kernel_base = exp_term / norm_constant\n            kernel_real = kernel_base * np.cos(2 * np.pi * scaled_t)\n            kernel_imag = kernel_base * np.sin(2 * np.pi * scaled_t)\n            wavelet_bank_real.append(kernel_real)\n            wavelet_bank_imag.append(kernel_imag)\n\n        wavelet_bank_real = np.stack(wavelet_bank_real, axis=0)\n        wavelet_bank_imag = np.stack(wavelet_bank_imag, axis=0)\n\n        wavelet_bank_real = torch.from_numpy(wavelet_bank_real).unsqueeze(1).unsqueeze(2)\n        wavelet_bank_imag = torch.from_numpy(wavelet_bank_imag).unsqueeze(1).unsqueeze(2)\n        return wavelet_bank_real, wavelet_bank_imag\n\n    def forward(self, x):\n        border_crop = self.border_crop // self.stride\n        start = border_crop\n        end = (-border_crop) if border_crop &gt; 0 else None\n\n        # x [n_batch, n_channels, time_len]\n        out_reals = []\n        out_imags = []\n\n        in_width = x.size(2)\n        out_width = int(np.ceil(in_width / self.stride))\n        pad_along_width = np.max((out_width - 1) * self.stride + self.kernel_size - in_width, 0)\n        padding = pad_along_width // 2 + 1\n\n        for i in range(3):\n            # [n_batch, 1, 1, time_len]\n            x_ = x[:, i, :].unsqueeze(1).unsqueeze(2)\n            out_real = nn.functional.conv2d(x_, self.wavelet_bank_real, stride=(1, self.stride), padding=(0, padding))\n            out_imag = nn.functional.conv2d(x_, self.wavelet_bank_imag, stride=(1, self.stride), padding=(0, padding))\n            out_real = out_real.transpose(2, 1)\n            out_imag = out_imag.transpose(2, 1)\n            out_reals.append(out_real)\n            out_imags.append(out_imag)\n\n        out_real = torch.cat(out_reals, axis=1)\n        out_imag = torch.cat(out_imags, axis=1)\n\n        out_real = out_real[:, :, :, start:end]\n        out_imag = out_imag[:, :, :, start:end]\n\n        scalograms = torch.sqrt(out_real ** 2 + out_imag ** 2)\n        return scalograms\n</code></pre>",
      "rawMarkdown": "## Hidehisa's Part\n\nI and @yasufuminakama and @naoism have formed a team from comparatively early stage of the competition, so our approaches were quite similar each other, so you may find some duplicates between what I've done and what others have done.\n\nHere's the list that I've done and those worked.\n\n* TensorFlow pipeline (maybe good in terms of diversity)\n  * I was the only one in the team who used TF pipeline until the end\n* Use statistics of the whole dataset to normalize each samples (not the statistics of each samples!)\n  * Separately reported as effective by Naoism and hirune's team (and indeed quite effective)\n* CWT was almost as good as CQT \n  * I used the implementation of https://github.com/Kevin-McIsaac/cmorlet-tensorflow\n  * I've also  implemented PyTorch version of CWT. This implementation was used through the team (code below).\n* Found Negative label mixup for image(p=1.0) effective (which is described in Y.Nakama's part and Naoism's part)\n* Found Center Crop along with the time axis effective (also described in Y.Nakama's part and Naoism's part)\n* Found training with OOF somewhat effective (also described in Naoism's part)\n  * We thought this was because low SNR signals can be thought as samples with noisy labels. Does someone tried the same thing? (approaches for training with noisy labels settings)\n  * Our experiment was a bit leaky, so I'd like to check how much it was effective\n* SwinTransformer Model (used the implementation of https://github.com/rishigami/Swin-Transformer-TF)\n  * Not very good, but I think it contributed a bit in terms of diversity\n\n### CWT implementation in pytorch\n\n```python\nclass CWT(nn.Module):\n    def __init__(\n        self,\n        wavelet_width,\n        fs,\n        lower_freq,\n        upper_freq,\n        n_scales,\n        size_factor=1.0,\n        border_crop=0,\n        stride=1\n    ):\n        super().__init__()\n        \n        self.initial_wavelet_width = wavelet_width\n        self.fs = fs\n        self.lower_freq = lower_freq\n        self.upper_freq = upper_freq\n        self.size_factor = size_factor\n        self.n_scales = n_scales\n        self.wavelet_width = wavelet_width\n        self.border_crop = border_crop\n        self.stride = stride\n        wavelet_bank_real, wavelet_bank_imag = self._build_wavelet_kernel()\n        self.wavelet_bank_real = nn.Parameter(wavelet_bank_real, requires_grad=False)\n        self.wavelet_bank_imag = nn.Parameter(wavelet_bank_imag, requires_grad=False)\n        \n        self.kernel_size = self.wavelet_bank_real.size(3)\n        \n    def _build_wavelet_kernel(self):\n        s_0 = 1 / self.upper_freq\n        s_n = 1 / self.lower_freq\n        \n        base = np.power(s_n / s_0, 1 / (self.n_scales - 1))\n        scales = s_0 * np.power(base, np.arange(self.n_scales))\n        \n        frequencies = 1 / scales\n        truncation_size = scales.max() * np.sqrt(4.5 * self.initial_wavelet_width) * self.fs\n        one_side = int(self.size_factor * truncation_size)\n        kernel_size = 2 * one_side + 1\n        \n        k_array = np.arange(kernel_size, dtype=np.float32) - one_side\n        t_array = k_array / self.fs\n        \n        wavelet_bank_real = []\n        wavelet_bank_imag = []\n        \n        for scale in scales:\n            norm_constant = np.sqrt(np.pi * self.wavelet_width) * scale * self.fs / 2.0\n            scaled_t = t_array / scale\n            exp_term = np.exp(-(scaled_t ** 2) / self.wavelet_width)\n            kernel_base = exp_term / norm_constant\n            kernel_real = kernel_base * np.cos(2 * np.pi * scaled_t)\n            kernel_imag = kernel_base * np.sin(2 * np.pi * scaled_t)\n            wavelet_bank_real.append(kernel_real)\n            wavelet_bank_imag.append(kernel_imag)\n            \n        wavelet_bank_real = np.stack(wavelet_bank_real, axis=0)\n        wavelet_bank_imag = np.stack(wavelet_bank_imag, axis=0)\n        \n        wavelet_bank_real = torch.from_numpy(wavelet_bank_real).unsqueeze(1).unsqueeze(2)\n        wavelet_bank_imag = torch.from_numpy(wavelet_bank_imag).unsqueeze(1).unsqueeze(2)\n        return wavelet_bank_real, wavelet_bank_imag\n    \n    def forward(self, x):\n        border_crop = self.border_crop // self.stride\n        start = border_crop\n        end = (-border_crop) if border_crop > 0 else None\n        \n        # x [n_batch, n_channels, time_len]\n        out_reals = []\n        out_imags = []\n        \n        in_width = x.size(2)\n        out_width = int(np.ceil(in_width / self.stride))\n        pad_along_width = np.max((out_width - 1) * self.stride + self.kernel_size - in_width, 0)\n        padding = pad_along_width // 2 + 1\n        \n        for i in range(3):\n            # [n_batch, 1, 1, time_len]\n            x_ = x[:, i, :].unsqueeze(1).unsqueeze(2)\n            out_real = nn.functional.conv2d(x_, self.wavelet_bank_real, stride=(1, self.stride), padding=(0, padding))\n            out_imag = nn.functional.conv2d(x_, self.wavelet_bank_imag, stride=(1, self.stride), padding=(0, padding))\n            out_real = out_real.transpose(2, 1)\n            out_imag = out_imag.transpose(2, 1)\n            out_reals.append(out_real)\n            out_imags.append(out_imag)\n            \n        out_real = torch.cat(out_reals, axis=1)\n        out_imag = torch.cat(out_imags, axis=1)\n        \n        out_real = out_real[:, :, :, start:end]\n        out_imag = out_imag[:, :, :, start:end]\n        \n        scalograms = torch.sqrt(out_real ** 2 + out_imag ** 2)\n        return scalograms\n```",
      "votes": null
    },
    {
      "id": "1559887",
      "postDate": "10/27/2021 08:06:06",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1528994,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "09/30/2021 04:39:19",
      "content": "<p><a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a>, I'm glad that you eventually made it! after we missed gold in the previous competition. Big congratulations to u and your team.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1529317,
          "author_name": "naoism",
          "author_url": "",
          "post_date": "09/30/2021 09:41:22",
          "content": "<p>Thank you very much. I was finally able to achieve my goal of becoming a Master. I'm really happy.<br>\nAlso, as a team member of the previous competition, I am happy that you finally become a GM. Congratulations!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1529051,
      "author_name": "jdxyw2004",
      "author_url": "",
      "post_date": "09/30/2021 05:48:04",
      "content": "<p>congratulation！  137 models!!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1529060,
      "author_name": "callmeb",
      "author_url": "",
      "post_date": "09/30/2021 05:59:04",
      "content": "<p>Nice to see that it required you 137 models to beat us 🤣<br>\nGood job for the gold 💪</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1529071,
      "author_name": "tanulsingh077",
      "author_url": "",
      "post_date": "09/30/2021 06:17:57",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hirune924\" target=\"_blank\">@hirune924</a> , thanks for sharing the solution , looks like a very elegant one</p>\n<p><code>x = x1 + x2\ny = y1 + y2 -(y1*y2)</code></p>\n<p>Can you report the score increase due to this , we also tried something on similar terms but didn't get a lot out of it</p>",
      "votes": null,
      "replies": [
        {
          "id": 1529095,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "09/30/2021 06:41:08",
          "content": "<p>I don't know the exact difference in scores from regular mixup because I used that trick from the beginning, but using mixup prevented overfitting during long training.<br>\nAlso, considering the possibility that this trick was negatively affecting the training, I reduced the frequency of its use from 50% to 15%, and the cv score decreased by about 0.0002.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1529297,
      "author_name": "yasufuminakama",
      "author_url": "",
      "post_date": "09/30/2021 09:25:25",
      "content": "<p>First of all, I would like to thank my teammates for competing with me.<br>\nAnd I would like to thank kaggle &amp; host for the interesting competition and to all the participants for giving us a lot of ideas.</p>\n<h1>Y.Nakama part</h1>\n<p>Here is my approach.</p>\n<h2>train setup</h2>\n<ul>\n<li>main model: CWT + tf_efficientnet_b5</li>\n<li>lr: 1e-3</li>\n<li>scheduler: CosineAnnealingLR</li>\n<li>optimizer: Adam (better than AdamW in my experiments)</li>\n</ul>\n<h2>preprocessing</h2>\n<ul>\n<li>z-score normalization</li>\n</ul>\n<pre><code>MEAN_ALL = [6.901084819126393e-26, 5.117726793550612e-26, -1.3831247922595496e-26]\nSTD_ALL = [7.420282939705547e-21, 7.41993950281645e-21, 1.838329278493792e-21]\n\ndef z_score_normalization_wave(_wave, site):\n    x_mean = MEAN_ALL[site]\n    x_std = STD_ALL[site]\n    wave = (_wave - x_mean) / (x_std)\n    return wave\n</code></pre>\n<p>No gain compared to dividing the all wave form by 4.6152116213830774e-20 (max(abs) of the entire train and test data).</p>\n<ul>\n<li>What didn't work for me: highpass filter + lowpassfilter + tukey window</li>\n</ul>\n<h2>Augmentation</h2>\n<ul>\n<li>mixup</li>\n</ul>\n<pre><code>def mixup(image, label, max_noise=0.5, probability=1.0):\n    if np.random.uniform(0, 1) &lt;= probability:\n        bs, c, h, w = image.size()\n        indices = torch.randperm(bs)\n        shuffled_image = image[indices]\n        shuffled_label = label[indices]\n        lam = np.random.uniform(0, 1) * (1.0 - shuffled_label) * max_noise # don't mix shuffled_label=1\n        lam = torch.flatten(lam)\n        image = (image.reshape(-1, bs) * (1 - lam)).reshape(bs, c, h, w) + (shuffled_image.reshape(-1, bs) * lam).reshape(bs, c, h, w)\n    return image, label\n</code></pre>\n<ul>\n<li>time shift with centercrop</li>\n</ul>\n<pre><code>w_diff = int(torch.randint(low=0, high=5, size=(1,)))\nif torch.rand(1) &lt; 0.5:\n    x = x[:,:,self.h_div_start:self.h_div_end,self.w_div_start+w_diff:self.w_div_end+w_diff]\nelse:\n    x = x[:,:,self.h_div_start:self.h_div_end,self.w_div_start-w_diff:self.w_div_end-w_diff]\n</code></pre>\n<ul>\n<li>What didn't work for me: gaussian noise</li>\n</ul>\n<h2>pseudo label</h2>\n<p>Re-training on soft pseudo-labelled test dataset improved score.</p>\n<h2>TTA</h2>\n<ul>\n<li>time shift with centercrop</li>\n</ul>\n<h2>stacking</h2>\n<p>I implemented stacking models as I did in <a href=\"https://www.kaggle.com/c/stanford-covid-vaccine/discussion/189728\" target=\"_blank\">OpenVaccine competition</a>.<br>\nNN stacking model performed better than GBDT stacking models.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1529331,
      "author_name": "sinpcw",
      "author_url": "",
      "post_date": "09/30/2021 09:57:39",
      "content": "<p>Thank you to all the participants for their hard work. <br>\nThank you to the organizers, hosts and teammates.</p>\n<h2>SiNpcw part:</h2>\n<p>Up to a week before the competition, we were working on 2d CNN, but in the final stage, I was looking for improvements to 1d CNN.</p>\n<h2>2d CNN:</h2>\n<ul>\n<li>model: efficentnet_v2_m</li>\n<li>optimizer: Adam</li>\n<li>lr: 1e-3</li>\n<li>scheduler: CosineAnnealingLR</li>\n<li>image size: 138x513</li>\n</ul>\n<p>It is almost the same preprocessing as <a href=\"https://www.kaggle.com/hirune924\" target=\"_blank\">@hirune924</a> but differs in a few details.<br>\nI applied 1dGridDistortion (only in the frequency direction) as a unique element. It has slightly improved the CV.<br>\nIn training the model, we used only nuttall window among the CQTs. This is because nuttall was the best window for CV.<br>\nI tried deleting data that was prone to miss-prediction, but not work.</p>\n<h2>1d CNN:</h2>\n<p>Improvements were made based on the (<a href=\"https://www.kaggle.com/scaomath/g2net-1d-cnn-gem-pool-pytorch-train-inference\" target=\"_blank\">published notebook</a>).<br>\nI used a strategy that of a noisy student, creating pseudo-labels in my own model and learning them again.  <br>\nDue to deadlines, I was not able to try enough of them, but the following ones worked well.</p>\n<ul>\n<li>SAM optimizer</li>\n<li>adjusting bandpass parameters</li>\n<li>ensemble: augmented and non-augmented models</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1529412,
      "author_name": "naoism",
      "author_url": "",
      "post_date": "09/30/2021 11:28:53",
      "content": "<p>At first, I would like to express my thanks to all the organizers for hosting this interesting competition and my teammates.<br>\nAnd thanks to all other competitors for sharing valuable information in forums and notebooks.</p>\n<p>I was finally able to achieve my goal of becoming a Master :)</p>\n<h1>Naoism part</h1>\n<p>My solution is very similar to <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>'s one.<br>\nSo I won't explain each method in detail.</p>\n<h2>train setup</h2>\n<ul>\n<li>models<ul>\n<li>2D model: CQT + EfficientNet, DenseNet and ResNet</li>\n<li>1D model: Simple 1D-CNN</li></ul></li>\n<li>Scheduler: CosineAnnealingLR</li>\n<li>Epoch: 8~10</li>\n<li>Optimizer = AdamW (better than Adam and RAdam) for 2Dmodel and SAM(better than AdamW) for 1D-CNN</li>\n<li>image size = (128, 768)</li>\n</ul>\n<h2>preprocessing</h2>\n<ul>\n<li>pycbc package's highpass filter(32 &lt; f) and notchfilter(30 &lt; f &lt; 80)</li>\n<li>Tukey window</li>\n<li>Z-score normalization</li>\n</ul>\n<pre><code>def z_score_normalization_wave(_wave, site):\n    MEAN_ALL = [-7.687203989766856e-31, 1.5238509999790067e-30, -1.7078072503529966e-30]\n    STD_ALL = [1.618816429314542e-22, 1.6188170892087954e-22, 2.621679130415917e-22]\n    x_mean = MEAN_ALL[site]\n    x_std = STD_ALL[site]\n    wave = (_wave-x_mean) / (x_std)\n    return wave\n\ndef load_wave(path):\n    waves = np.load(path)\n    new_waves = []\n    window = signal.tukey(4096)\n    for i in [0,1,2]:\n        _wave = pycbc.filter.resample.highpass_fir(pycbc.types.TimeSeries(waves[i], epoch=0, delta_t=1.0/2048), frequency=32, order=100)\n        _wave = pycbc.filter.resample.notch_fir(_wave, f1=30, f2=80, order=10, beta=5)\n        _wave = np.array(_wave)\n        _wave = _wave * window\n        wave = z_score_normalization_wave(_wave, site=i)\n        new_waves.append(wave)\n    return np.array(new_waves)\n</code></pre>\n<ul>\n<li>Centercrop(0.4s ~ 1.8s)</li>\n</ul>\n<h2>Augmentation</h2>\n<ul>\n<li>Negative label mixup for image(p=1.0) (Almost the same as the one introduced by <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>)</li>\n<li>Time shift with centercrop</li>\n<li>GaussianNoise(p=0.3, max_SNR = 0.2)</li>\n<li>SpecAugmentation(Frequency mask and Time mask)</li>\n</ul>\n<h2>Other ideas</h2>\n<ul>\n<li>Using the other model's oof as the target improved the CV considerably.</li>\n<li>Also, repeating the process of getting the predictions of a model-&gt;stacking with multiple models-&gt;retraining with oof as target slightly improved the CV.</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1529639,
      "author_name": "hidehisaarai1213",
      "author_url": "",
      "post_date": "09/30/2021 14:46:22",
      "content": "<h2>Hidehisa's Part</h2>\n<p>I and <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> and <a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a> have formed a team from comparatively early stage of the competition, so our approaches were quite similar each other, so you may find some duplicates between what I've done and what others have done.</p>\n<p>Here's the list that I've done and those worked.</p>\n<ul>\n<li>TensorFlow pipeline (maybe good in terms of diversity)<ul>\n<li>I was the only one in the team who used TF pipeline until the end</li></ul></li>\n<li>Use statistics of the whole dataset to normalize each samples (not the statistics of each samples!)<ul>\n<li>Separately reported as effective by Naoism and hirune's team (and indeed quite effective)</li></ul></li>\n<li>CWT was almost as good as CQT <ul>\n<li>I used the implementation of <a href=\"https://github.com/Kevin-McIsaac/cmorlet-tensorflow\" target=\"_blank\">https://github.com/Kevin-McIsaac/cmorlet-tensorflow</a></li>\n<li>I've also  implemented PyTorch version of CWT. This implementation was used through the team (code below).</li></ul></li>\n<li>Found Negative label mixup for image(p=1.0) effective (which is described in Y.Nakama's part and Naoism's part)</li>\n<li>Found Center Crop along with the time axis effective (also described in Y.Nakama's part and Naoism's part)</li>\n<li>Found training with OOF somewhat effective (also described in Naoism's part)<ul>\n<li>We thought this was because low SNR signals can be thought as samples with noisy labels. Does someone tried the same thing? (approaches for training with noisy labels settings)</li>\n<li>Our experiment was a bit leaky, so I'd like to check how much it was effective</li></ul></li>\n<li>SwinTransformer Model (used the implementation of <a href=\"https://github.com/rishigami/Swin-Transformer-TF\" target=\"_blank\">https://github.com/rishigami/Swin-Transformer-TF</a>)<ul>\n<li>Not very good, but I think it contributed a bit in terms of diversity</li></ul></li>\n</ul>\n<h3>CWT implementation in pytorch</h3>\n<pre><code>class CWT(nn.Module):\n    def __init__(\n        self,\n        wavelet_width,\n        fs,\n        lower_freq,\n        upper_freq,\n        n_scales,\n        size_factor=1.0,\n        border_crop=0,\n        stride=1\n    ):\n        super().__init__()\n\n        self.initial_wavelet_width = wavelet_width\n        self.fs = fs\n        self.lower_freq = lower_freq\n        self.upper_freq = upper_freq\n        self.size_factor = size_factor\n        self.n_scales = n_scales\n        self.wavelet_width = wavelet_width\n        self.border_crop = border_crop\n        self.stride = stride\n        wavelet_bank_real, wavelet_bank_imag = self._build_wavelet_kernel()\n        self.wavelet_bank_real = nn.Parameter(wavelet_bank_real, requires_grad=False)\n        self.wavelet_bank_imag = nn.Parameter(wavelet_bank_imag, requires_grad=False)\n\n        self.kernel_size = self.wavelet_bank_real.size(3)\n\n    def _build_wavelet_kernel(self):\n        s_0 = 1 / self.upper_freq\n        s_n = 1 / self.lower_freq\n\n        base = np.power(s_n / s_0, 1 / (self.n_scales - 1))\n        scales = s_0 * np.power(base, np.arange(self.n_scales))\n\n        frequencies = 1 / scales\n        truncation_size = scales.max() * np.sqrt(4.5 * self.initial_wavelet_width) * self.fs\n        one_side = int(self.size_factor * truncation_size)\n        kernel_size = 2 * one_side + 1\n\n        k_array = np.arange(kernel_size, dtype=np.float32) - one_side\n        t_array = k_array / self.fs\n\n        wavelet_bank_real = []\n        wavelet_bank_imag = []\n\n        for scale in scales:\n            norm_constant = np.sqrt(np.pi * self.wavelet_width) * scale * self.fs / 2.0\n            scaled_t = t_array / scale\n            exp_term = np.exp(-(scaled_t ** 2) / self.wavelet_width)\n            kernel_base = exp_term / norm_constant\n            kernel_real = kernel_base * np.cos(2 * np.pi * scaled_t)\n            kernel_imag = kernel_base * np.sin(2 * np.pi * scaled_t)\n            wavelet_bank_real.append(kernel_real)\n            wavelet_bank_imag.append(kernel_imag)\n\n        wavelet_bank_real = np.stack(wavelet_bank_real, axis=0)\n        wavelet_bank_imag = np.stack(wavelet_bank_imag, axis=0)\n\n        wavelet_bank_real = torch.from_numpy(wavelet_bank_real).unsqueeze(1).unsqueeze(2)\n        wavelet_bank_imag = torch.from_numpy(wavelet_bank_imag).unsqueeze(1).unsqueeze(2)\n        return wavelet_bank_real, wavelet_bank_imag\n\n    def forward(self, x):\n        border_crop = self.border_crop // self.stride\n        start = border_crop\n        end = (-border_crop) if border_crop &gt; 0 else None\n\n        # x [n_batch, n_channels, time_len]\n        out_reals = []\n        out_imags = []\n\n        in_width = x.size(2)\n        out_width = int(np.ceil(in_width / self.stride))\n        pad_along_width = np.max((out_width - 1) * self.stride + self.kernel_size - in_width, 0)\n        padding = pad_along_width // 2 + 1\n\n        for i in range(3):\n            # [n_batch, 1, 1, time_len]\n            x_ = x[:, i, :].unsqueeze(1).unsqueeze(2)\n            out_real = nn.functional.conv2d(x_, self.wavelet_bank_real, stride=(1, self.stride), padding=(0, padding))\n            out_imag = nn.functional.conv2d(x_, self.wavelet_bank_imag, stride=(1, self.stride), padding=(0, padding))\n            out_real = out_real.transpose(2, 1)\n            out_imag = out_imag.transpose(2, 1)\n            out_reals.append(out_real)\n            out_imags.append(out_imag)\n\n        out_real = torch.cat(out_reals, axis=1)\n        out_imag = torch.cat(out_imags, axis=1)\n\n        out_real = out_real[:, :, :, start:end]\n        out_imag = out_imag[:, :, :, start:end]\n\n        scalograms = torch.sqrt(out_real ** 2 + out_imag ** 2)\n        return scalograms\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1559887,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 08:06:06",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1528973": "First of all, I would like to thank my teammates( @naoism, @hidehisaarai1213, @yasufuminakama, @sinpcw ) for competing with me. I would also like to thank kaggle and EGO for organizing an interesting competition.\nI explain my approach. My teammates will add their own approaches in the comments!\n* Y.Nakama part ( https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529297 )\n* SiNpcw part ( https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529331 )\n* Naoism part ( https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529412 )\n* Hidehisa part ( https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275356#1529639 )\n\n## hirune part intro\nI joined this competition after the SETI competition was over. This task is very similar to the SETI competition task. This task is very similar to the SETI competition task, so I reused most of the SETI code.\n\n### preprocessing\n1. Divide the all wave form by 4.6152116213830774e-20 (max(abs) of the entire train and test data).\n2. Use nnAudio to run CQT. Randomly select one of flattop, blackmanharris, or nuttall to the window. This can be used in TTA. Finally, I also added CWT to this.\n3. The spectrograms are combined in the frequency direction and input to the model as a single-channel image\n4. Resize to 384x512 before inputting into the model, and normalize the spectrogram of the entire data set in mean, std\n\n### Augmentation\n* Before converting to a spectrogram, mixup as follows.\n```\nx = x1 + x2\ny = y1 + y2 -(y1*y2)\n```\n* Randomly roll shift in time direction\n\n### train setup\n* model: timm tf_efficientnet_b4_ap\n* optimizer: Adam\n* Lr: 0.001\n* scheduler: CosineAnnealingLR\n\n### pseudo label\nUsing pseudo-labels showed some improvement, but not a lot.\n\n### stacking\nUsing other team members models, and finally stacking 137 models with NN or XGB, there was a significant score increase. Other models include 1dcnn, swin transformer and efficientnet b5-8.\nTeam members will explain these details later.",
    "1528994": "naoism, I'm glad that you eventually made it! after we missed gold in the previous competition. Big congratulations to u and your team.",
    "1529051": "congratulation！  137 models!!!!",
    "1529060": "Nice to see that it required you 137 models to beat us 🤣\nGood job for the gold 💪",
    "1529071": "Hi @hirune924 , thanks for sharing the solution , looks like a very elegant one\n\n`x = x1 + x2\ny = y1 + y2 -(y1*y2)`\n\nCan you report the score increase due to this , we also tried something on similar terms but didn't get a lot out of it",
    "1529095": "I don't know the exact difference in scores from regular mixup because I used that trick from the beginning, but using mixup prevented overfitting during long training.\nAlso, considering the possibility that this trick was negatively affecting the training, I reduced the frequency of its use from 50% to 15%, and the cv score decreased by about 0.0002.",
    "1529297": "First of all, I would like to thank my teammates for competing with me.\nAnd I would like to thank kaggle & host for the interesting competition and to all the participants for giving us a lot of ideas.\n\n# Y.Nakama part \nHere is my approach.\n\n## train setup\n- main model: CWT + tf_efficientnet_b5\n- lr: 1e-3\n- scheduler: CosineAnnealingLR\n- optimizer: Adam (better than AdamW in my experiments)\n\n## preprocessing\n- z-score normalization\n```\nMEAN_ALL = [6.901084819126393e-26, 5.117726793550612e-26, -1.3831247922595496e-26]\nSTD_ALL = [7.420282939705547e-21, 7.41993950281645e-21, 1.838329278493792e-21]\n\ndef z_score_normalization_wave(_wave, site):\n    x_mean = MEAN_ALL[site]\n    x_std = STD_ALL[site]\n    wave = (_wave - x_mean) / (x_std)\n    return wave\n```\nNo gain compared to dividing the all wave form by 4.6152116213830774e-20 (max(abs) of the entire train and test data).\n- What didn't work for me: highpass filter + lowpassfilter + tukey window\n\n## Augmentation\n- mixup\n```\ndef mixup(image, label, max_noise=0.5, probability=1.0):\n    if np.random.uniform(0, 1) <= probability:\n        bs, c, h, w = image.size()\n        indices = torch.randperm(bs)\n        shuffled_image = image[indices]\n        shuffled_label = label[indices]\n        lam = np.random.uniform(0, 1) * (1.0 - shuffled_label) * max_noise # don't mix shuffled_label=1\n        lam = torch.flatten(lam)\n        image = (image.reshape(-1, bs) * (1 - lam)).reshape(bs, c, h, w) + (shuffled_image.reshape(-1, bs) * lam).reshape(bs, c, h, w)\n    return image, label\n```\n- time shift with centercrop\n```\nw_diff = int(torch.randint(low=0, high=5, size=(1,)))\nif torch.rand(1) < 0.5:\n    x = x[:,:,self.h_div_start:self.h_div_end,self.w_div_start+w_diff:self.w_div_end+w_diff]\nelse:\n    x = x[:,:,self.h_div_start:self.h_div_end,self.w_div_start-w_diff:self.w_div_end-w_diff]\n```\n- What didn't work for me: gaussian noise\n\n## pseudo label\nRe-training on soft pseudo-labelled test dataset improved score.\n\n## TTA\n- time shift with centercrop\n\n## stacking\nI implemented stacking models as I did in [OpenVaccine competition](https://www.kaggle.com/c/stanford-covid-vaccine/discussion/189728).\nNN stacking model performed better than GBDT stacking models.",
    "1529317": "Thank you very much. I was finally able to achieve my goal of becoming a Master. I'm really happy.\nAlso, as a team member of the previous competition, I am happy that you finally become a GM. Congratulations!",
    "1529331": "Thank you to all the participants for their hard work. \nThank you to the organizers, hosts and teammates.\n\n## SiNpcw part:\nUp to a week before the competition, we were working on 2d CNN, but in the final stage, I was looking for improvements to 1d CNN.\n\n## 2d CNN:\n- model: efficentnet_v2_m\n- optimizer: Adam\n- lr: 1e-3\n- scheduler: CosineAnnealingLR\n- image size: 138x513\n  \nIt is almost the same preprocessing as @hirune924 but differs in a few details.\nI applied 1dGridDistortion (only in the frequency direction) as a unique element. It has slightly improved the CV.\nIn training the model, we used only nuttall window among the CQTs. This is because nuttall was the best window for CV.\nI tried deleting data that was prone to miss-prediction, but not work.\n\n## 1d CNN:\nImprovements were made based on the ([published notebook](https://www.kaggle.com/scaomath/g2net-1d-cnn-gem-pool-pytorch-train-inference)).\nI used a strategy that of a noisy student, creating pseudo-labels in my own model and learning them again.  \nDue to deadlines, I was not able to try enough of them, but the following ones worked well.\n- SAM optimizer\n- adjusting bandpass parameters\n- ensemble: augmented and non-augmented models",
    "1529412": "At first, I would like to express my thanks to all the organizers for hosting this interesting competition and my teammates.\nAnd thanks to all other competitors for sharing valuable information in forums and notebooks.\n\nI was finally able to achieve my goal of becoming a Master :)\n\n# Naoism part\nMy solution is very similar to @yasufuminakama's one.\nSo I won't explain each method in detail.\n\n## train setup\n- models\n  - 2D model: CQT + EfficientNet, DenseNet and ResNet\n  - 1D model: Simple 1D-CNN\n- Scheduler: CosineAnnealingLR\n- Epoch: 8~10\n- Optimizer = AdamW (better than Adam and RAdam) for 2Dmodel and SAM(better than AdamW) for 1D-CNN\n- image size = (128, 768)\n\n## preprocessing\n- pycbc package's highpass filter(32 < f) and notchfilter(30 < f < 80)\n- Tukey window\n- Z-score normalization\n\n```\ndef z_score_normalization_wave(_wave, site):\n    MEAN_ALL = [-7.687203989766856e-31, 1.5238509999790067e-30, -1.7078072503529966e-30]\n    STD_ALL = [1.618816429314542e-22, 1.6188170892087954e-22, 2.621679130415917e-22]\n    x_mean = MEAN_ALL[site]\n    x_std = STD_ALL[site]\n    wave = (_wave-x_mean) / (x_std)\n    return wave\n\ndef load_wave(path):\n    waves = np.load(path)\n    new_waves = []\n    window = signal.tukey(4096)\n    for i in [0,1,2]:\n        _wave = pycbc.filter.resample.highpass_fir(pycbc.types.TimeSeries(waves[i], epoch=0, delta_t=1.0/2048), frequency=32, order=100)\n        _wave = pycbc.filter.resample.notch_fir(_wave, f1=30, f2=80, order=10, beta=5)\n        _wave = np.array(_wave)\n        _wave = _wave * window\n        wave = z_score_normalization_wave(_wave, site=i)\n        new_waves.append(wave)\n    return np.array(new_waves)\n```\n- Centercrop(0.4s ~ 1.8s)\n\n## Augmentation\n\n- Negative label mixup for image(p=1.0) (Almost the same as the one introduced by @yasufuminakama)\n- Time shift with centercrop\n- GaussianNoise(p=0.3, max_SNR = 0.2)\n- SpecAugmentation(Frequency mask and Time mask)\n\n## Other ideas\n- Using the other model's oof as the target improved the CV considerably.\n- Also, repeating the process of getting the predictions of a model->stacking with multiple models->retraining with oof as target slightly improved the CV.",
    "1529639": "## Hidehisa's Part\n\nI and @yasufuminakama and @naoism have formed a team from comparatively early stage of the competition, so our approaches were quite similar each other, so you may find some duplicates between what I've done and what others have done.\n\nHere's the list that I've done and those worked.\n\n* TensorFlow pipeline (maybe good in terms of diversity)\n  * I was the only one in the team who used TF pipeline until the end\n* Use statistics of the whole dataset to normalize each samples (not the statistics of each samples!)\n  * Separately reported as effective by Naoism and hirune's team (and indeed quite effective)\n* CWT was almost as good as CQT \n  * I used the implementation of https://github.com/Kevin-McIsaac/cmorlet-tensorflow\n  * I've also  implemented PyTorch version of CWT. This implementation was used through the team (code below).\n* Found Negative label mixup for image(p=1.0) effective (which is described in Y.Nakama's part and Naoism's part)\n* Found Center Crop along with the time axis effective (also described in Y.Nakama's part and Naoism's part)\n* Found training with OOF somewhat effective (also described in Naoism's part)\n  * We thought this was because low SNR signals can be thought as samples with noisy labels. Does someone tried the same thing? (approaches for training with noisy labels settings)\n  * Our experiment was a bit leaky, so I'd like to check how much it was effective\n* SwinTransformer Model (used the implementation of https://github.com/rishigami/Swin-Transformer-TF)\n  * Not very good, but I think it contributed a bit in terms of diversity\n\n### CWT implementation in pytorch\n\n```python\nclass CWT(nn.Module):\n    def __init__(\n        self,\n        wavelet_width,\n        fs,\n        lower_freq,\n        upper_freq,\n        n_scales,\n        size_factor=1.0,\n        border_crop=0,\n        stride=1\n    ):\n        super().__init__()\n        \n        self.initial_wavelet_width = wavelet_width\n        self.fs = fs\n        self.lower_freq = lower_freq\n        self.upper_freq = upper_freq\n        self.size_factor = size_factor\n        self.n_scales = n_scales\n        self.wavelet_width = wavelet_width\n        self.border_crop = border_crop\n        self.stride = stride\n        wavelet_bank_real, wavelet_bank_imag = self._build_wavelet_kernel()\n        self.wavelet_bank_real = nn.Parameter(wavelet_bank_real, requires_grad=False)\n        self.wavelet_bank_imag = nn.Parameter(wavelet_bank_imag, requires_grad=False)\n        \n        self.kernel_size = self.wavelet_bank_real.size(3)\n        \n    def _build_wavelet_kernel(self):\n        s_0 = 1 / self.upper_freq\n        s_n = 1 / self.lower_freq\n        \n        base = np.power(s_n / s_0, 1 / (self.n_scales - 1))\n        scales = s_0 * np.power(base, np.arange(self.n_scales))\n        \n        frequencies = 1 / scales\n        truncation_size = scales.max() * np.sqrt(4.5 * self.initial_wavelet_width) * self.fs\n        one_side = int(self.size_factor * truncation_size)\n        kernel_size = 2 * one_side + 1\n        \n        k_array = np.arange(kernel_size, dtype=np.float32) - one_side\n        t_array = k_array / self.fs\n        \n        wavelet_bank_real = []\n        wavelet_bank_imag = []\n        \n        for scale in scales:\n            norm_constant = np.sqrt(np.pi * self.wavelet_width) * scale * self.fs / 2.0\n            scaled_t = t_array / scale\n            exp_term = np.exp(-(scaled_t ** 2) / self.wavelet_width)\n            kernel_base = exp_term / norm_constant\n            kernel_real = kernel_base * np.cos(2 * np.pi * scaled_t)\n            kernel_imag = kernel_base * np.sin(2 * np.pi * scaled_t)\n            wavelet_bank_real.append(kernel_real)\n            wavelet_bank_imag.append(kernel_imag)\n            \n        wavelet_bank_real = np.stack(wavelet_bank_real, axis=0)\n        wavelet_bank_imag = np.stack(wavelet_bank_imag, axis=0)\n        \n        wavelet_bank_real = torch.from_numpy(wavelet_bank_real).unsqueeze(1).unsqueeze(2)\n        wavelet_bank_imag = torch.from_numpy(wavelet_bank_imag).unsqueeze(1).unsqueeze(2)\n        return wavelet_bank_real, wavelet_bank_imag\n    \n    def forward(self, x):\n        border_crop = self.border_crop // self.stride\n        start = border_crop\n        end = (-border_crop) if border_crop > 0 else None\n        \n        # x [n_batch, n_channels, time_len]\n        out_reals = []\n        out_imags = []\n        \n        in_width = x.size(2)\n        out_width = int(np.ceil(in_width / self.stride))\n        pad_along_width = np.max((out_width - 1) * self.stride + self.kernel_size - in_width, 0)\n        padding = pad_along_width // 2 + 1\n        \n        for i in range(3):\n            # [n_batch, 1, 1, time_len]\n            x_ = x[:, i, :].unsqueeze(1).unsqueeze(2)\n            out_real = nn.functional.conv2d(x_, self.wavelet_bank_real, stride=(1, self.stride), padding=(0, padding))\n            out_imag = nn.functional.conv2d(x_, self.wavelet_bank_imag, stride=(1, self.stride), padding=(0, padding))\n            out_real = out_real.transpose(2, 1)\n            out_imag = out_imag.transpose(2, 1)\n            out_reals.append(out_real)\n            out_imags.append(out_imag)\n            \n        out_real = torch.cat(out_reals, axis=1)\n        out_imag = torch.cat(out_imags, axis=1)\n        \n        out_real = out_real[:, :, :, start:end]\n        out_imag = out_imag[:, :, :, start:end]\n        \n        scalograms = torch.sqrt(out_real ** 2 + out_imag ** 2)\n        return scalograms\n```",
    "1559887": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}