{
  "id": 275433,
  "title": "8th solution ",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/275433",
  "author_name": "",
  "post_date": "2021-09-30T12:34:46.755204500Z",
  "votes": 69,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I want to thank Kaggle and Host for this amazing challenge. Topic was great, and teaming with <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> and <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> was great too.</p>\n<p>We led the competition for a while then all of a sudden teams started to pass us.  In hindsight we know it is thanks to using 1D models and/or resnets. Unfortunately for us we kept working with efficientnet v2 as we were having good results so far.  We have no excuse given both ideas were shared in the forum soon enough to be leveraged.  To be honest we tried resnet34 but it did not match effnetv2 results for us.  I wonder why it is different for other teams.  We also tried 1D models briefly, but, same, they were low quality compared to effnets.  We should have submitted a blend obviously.  </p>\n<p>Here is a snapshot of our solution.</p>\n<p><strong>Data processing</strong></p>\n<p>Data was multiplied by 1e19 (1e21 later) to ensure that CQT and CWT could be performed correctly using FP16.</p>\n<p>CQT and CWT (results are very similar).  Best model  (public 0.8826 private 0.8805)  used these setting for CQT:</p>\n<pre><code>sr = 2048\nhop_length = 5\nfmin = 22\nfmax = 22*16\nbins_per_octave = 8\nn_octaves = 4\nn_bins = n_octaves * bins_per_octave\nfscale = 1\ncqt = CQT1992v2(sr=sr, hop_length=hop_length, fmin=fmin, fmax=fmax,\n                        n_bins=n_bins, bins_per_octave=bins_per_octave * fscale,         \n                window=('kaiser', 14), filter_scale=1/fscale,\n               )\n</code></pre>\n<p>We used fscale = 2 (i.e. filter_scale = 0.5) in few models, it improves a bit but we could not train them as much as our best model.  We also varied hop length to get different image sizes. We then concatenated the 3 CQT images on the frequency dim, getting images of size 96xL where L depends on hop length (L = 820 for the above setting).  </p>\n<p><strong>Signal To Noise EDA</strong></p>\n<p>We found the best frequency range by doing a signal ratio analysis.  We computed the average of CQT images for positive samples (POS) , and divided it by the average of CQT images for negative samples (NEG).  This gives us what we called a mask.  Here is the mask for above settings</p>\n<p><img src=\"https://i.imgur.com/SXWVA5I.png\" alt=\"mask\"></p>\n<p>We see where the waves are on average.</p>\n<p>By taking mask average over detectors and time we see that we can crop images left and right safely (y axis labels should be multiplied by 3, sorry):</p>\n<p><img src=\"https://i.imgur.com/BuBZGJo.png\" alt=\"time\"></p>\n<p>Cropping sides gets rid of the CQT border artifacts.</p>\n<p>By taking the mask mean over detectors and time we get signal to noise ratio per frequency (y axis labels should be multiplied by 3, sorry):</p>\n<p><img src=\"https://i.imgur.com/QTTfca3.png\" alt=\"freq\"></p>\n<p>x is log scale, but doing the math we see that snr is 1 outside 22Hz-352Hz</p>\n<p>We tuned bins per octave so that the image height is a multiple of 16 (for speed), and 32 bins overall was a very good tradeoff.</p>\n<p><strong>Input images</strong></p>\n<p>We used nnAudio CQT (or CWT) as the first model layer.  Then 3 outputs are stacked on frequency axis. We then divided the image by NEG (average of negative sample images) to get rid of noise on average.  This was quite better than whitening. We finally applied a log scaling, a bit similar to amplitude to db in audio signal processing.  An example of resulting image is shown below (ignore frequency axis labels):</p>\n<p><img src=\"https://i.imgur.com/ERThCPY.png\" alt=\"image\"></p>\n<p>Depending on the model we added the mask as an additional channel. We then scaled data to have multiple of 16 dimensions.  For instance resizing 96x820 to 1x96x768</p>\n<p><strong>Model</strong></p>\n<p>We mostly used efficientnet_v2s_s and efficientnet_v2_m from timm package.  Models were trained using pytorch.cuda.amp.  By using batch size and image dimensions that are multiple of 16 we use of RT cores and speed up training significantly on GPU V100.</p>\n<p>We used BCEWithLogitsLoss, Adam and OneCycleLR most of the competition. We later found that retraining models was helping. Our best model went through 4 training cycles.  Using cosine annealing from the start would probably have been a good choice.</p>\n<p><strong>Augmentations</strong></p>\n<p>Kazuki shared most of our augmentations here: <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275335\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275335</a></p>\n<p>We also used a random time shift in addition to these.</p>\n<p><strong>Pseudo Labeling</strong></p>\n<p>Pseudo labeling gave us a 0.0002 boost on LB when we used it.  We used non rounded predictions from teacher model for test dataset and concatenated to training fold.  It improved CV by 0.0008 on average but LB by only 0.0002.  However, our best model was not a pseudo labeling model.</p>\n<p><strong>Ensembling</strong></p>\n<p>We kept oof predictions for all our models and trained second level models on them. We used logistic regression on logits, scipy optimize on ranked predictions, and XGBoost on ranked predictions.  They yield similar results, and averaging them did not help really.  </p>\n<p>Ensembling moved us from 8826 for best single model to 8832 on public LB.</p>\n<p><strong>What did not work</strong></p>\n<ul>\n<li>Giba tried hard to generate additional positive samples.  Issue was to calibrate these to be of the same distribution as training data.  We did not find the right way in time.</li>\n<li>Resnet34.  As written above, we could not meet effnetv2s performance.  And speedup was only 2x anyway.</li>\n<li>We tried many other things that did not work well enough, like triplet models (a backbone on each detector then a common head), or using detectors as RGB channels.</li>\n</ul>\n<p>That's it.  We clearly missed 1D model, but for the rest we did everything we could.  And I learned from my team mates!  Teaming is good!.</p>",
  "messages": [
    {
      "id": "1529466",
      "postDate": "09/30/2021 12:34:46",
      "content": "<p>I want to thank Kaggle and Host for this amazing challenge. Topic was great, and teaming with <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> and <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> was great too.</p>\n<p>We led the competition for a while then all of a sudden teams started to pass us.  In hindsight we know it is thanks to using 1D models and/or resnets. Unfortunately for us we kept working with efficientnet v2 as we were having good results so far.  We have no excuse given both ideas were shared in the forum soon enough to be leveraged.  To be honest we tried resnet34 but it did not match effnetv2 results for us.  I wonder why it is different for other teams.  We also tried 1D models briefly, but, same, they were low quality compared to effnets.  We should have submitted a blend obviously.  </p>\n<p>Here is a snapshot of our solution.</p>\n<p><strong>Data processing</strong></p>\n<p>Data was multiplied by 1e19 (1e21 later) to ensure that CQT and CWT could be performed correctly using FP16.</p>\n<p>CQT and CWT (results are very similar).  Best model  (public 0.8826 private 0.8805)  used these setting for CQT:</p>\n<pre><code>sr = 2048\nhop_length = 5\nfmin = 22\nfmax = 22*16\nbins_per_octave = 8\nn_octaves = 4\nn_bins = n_octaves * bins_per_octave\nfscale = 1\ncqt = CQT1992v2(sr=sr, hop_length=hop_length, fmin=fmin, fmax=fmax,\n                        n_bins=n_bins, bins_per_octave=bins_per_octave * fscale,         \n                window=('kaiser', 14), filter_scale=1/fscale,\n               )\n</code></pre>\n<p>We used fscale = 2 (i.e. filter_scale = 0.5) in few models, it improves a bit but we could not train them as much as our best model.  We also varied hop length to get different image sizes. We then concatenated the 3 CQT images on the frequency dim, getting images of size 96xL where L depends on hop length (L = 820 for the above setting).  </p>\n<p><strong>Signal To Noise EDA</strong></p>\n<p>We found the best frequency range by doing a signal ratio analysis.  We computed the average of CQT images for positive samples (POS) , and divided it by the average of CQT images for negative samples (NEG).  This gives us what we called a mask.  Here is the mask for above settings</p>\n<p><img src=\"https://i.imgur.com/SXWVA5I.png\" alt=\"mask\"></p>\n<p>We see where the waves are on average.</p>\n<p>By taking mask average over detectors and time we see that we can crop images left and right safely (y axis labels should be multiplied by 3, sorry):</p>\n<p><img src=\"https://i.imgur.com/BuBZGJo.png\" alt=\"time\"></p>\n<p>Cropping sides gets rid of the CQT border artifacts.</p>\n<p>By taking the mask mean over detectors and time we get signal to noise ratio per frequency (y axis labels should be multiplied by 3, sorry):</p>\n<p><img src=\"https://i.imgur.com/QTTfca3.png\" alt=\"freq\"></p>\n<p>x is log scale, but doing the math we see that snr is 1 outside 22Hz-352Hz</p>\n<p>We tuned bins per octave so that the image height is a multiple of 16 (for speed), and 32 bins overall was a very good tradeoff.</p>\n<p><strong>Input images</strong></p>\n<p>We used nnAudio CQT (or CWT) as the first model layer.  Then 3 outputs are stacked on frequency axis. We then divided the image by NEG (average of negative sample images) to get rid of noise on average.  This was quite better than whitening. We finally applied a log scaling, a bit similar to amplitude to db in audio signal processing.  An example of resulting image is shown below (ignore frequency axis labels):</p>\n<p><img src=\"https://i.imgur.com/ERThCPY.png\" alt=\"image\"></p>\n<p>Depending on the model we added the mask as an additional channel. We then scaled data to have multiple of 16 dimensions.  For instance resizing 96x820 to 1x96x768</p>\n<p><strong>Model</strong></p>\n<p>We mostly used efficientnet_v2s_s and efficientnet_v2_m from timm package.  Models were trained using pytorch.cuda.amp.  By using batch size and image dimensions that are multiple of 16 we use of RT cores and speed up training significantly on GPU V100.</p>\n<p>We used BCEWithLogitsLoss, Adam and OneCycleLR most of the competition. We later found that retraining models was helping. Our best model went through 4 training cycles.  Using cosine annealing from the start would probably have been a good choice.</p>\n<p><strong>Augmentations</strong></p>\n<p>Kazuki shared most of our augmentations here: <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275335\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275335</a></p>\n<p>We also used a random time shift in addition to these.</p>\n<p><strong>Pseudo Labeling</strong></p>\n<p>Pseudo labeling gave us a 0.0002 boost on LB when we used it.  We used non rounded predictions from teacher model for test dataset and concatenated to training fold.  It improved CV by 0.0008 on average but LB by only 0.0002.  However, our best model was not a pseudo labeling model.</p>\n<p><strong>Ensembling</strong></p>\n<p>We kept oof predictions for all our models and trained second level models on them. We used logistic regression on logits, scipy optimize on ranked predictions, and XGBoost on ranked predictions.  They yield similar results, and averaging them did not help really.  </p>\n<p>Ensembling moved us from 8826 for best single model to 8832 on public LB.</p>\n<p><strong>What did not work</strong></p>\n<ul>\n<li>Giba tried hard to generate additional positive samples.  Issue was to calibrate these to be of the same distribution as training data.  We did not find the right way in time.</li>\n<li>Resnet34.  As written above, we could not meet effnetv2s performance.  And speedup was only 2x anyway.</li>\n<li>We tried many other things that did not work well enough, like triplet models (a backbone on each detector then a common head), or using detectors as RGB channels.</li>\n</ul>\n<p>That's it.  We clearly missed 1D model, but for the rest we did everything we could.  And I learned from my team mates!  Teaming is good!.</p>",
      "rawMarkdown": "I want to thank Kaggle and Host for this amazing challenge. Topic was great, and teaming with @onodera and @titericz was great too.\n\nWe led the competition for a while then all of a sudden teams started to pass us.  In hindsight we know it is thanks to using 1D models and/or resnets. Unfortunately for us we kept working with efficientnet v2 as we were having good results so far.  We have no excuse given both ideas were shared in the forum soon enough to be leveraged.  To be honest we tried resnet34 but it did not match effnetv2 results for us.  I wonder why it is different for other teams.  We also tried 1D models briefly, but, same, they were low quality compared to effnets.  We should have submitted a blend obviously.  \n\nHere is a snapshot of our solution.\n\n**Data processing**\n\nData was multiplied by 1e19 (1e21 later) to ensure that CQT and CWT could be performed correctly using FP16.\n\nCQT and CWT (results are very similar).  Best model  (public 0.8826 private 0.8805)  used these setting for CQT:\n\n```\nsr = 2048\nhop_length = 5\nfmin = 22\nfmax = 22*16\nbins_per_octave = 8\nn_octaves = 4\nn_bins = n_octaves * bins_per_octave\nfscale = 1\ncqt = CQT1992v2(sr=sr, hop_length=hop_length, fmin=fmin, fmax=fmax,\n                        n_bins=n_bins, bins_per_octave=bins_per_octave * fscale,         \n                window=('kaiser', 14), filter_scale=1/fscale,\n               )\n```\n\nWe used fscale = 2 (i.e. filter_scale = 0.5) in few models, it improves a bit but we could not train them as much as our best model.  We also varied hop length to get different image sizes. We then concatenated the 3 CQT images on the frequency dim, getting images of size 96xL where L depends on hop length (L = 820 for the above setting).  \n\n**Signal To Noise EDA**\n\nWe found the best frequency range by doing a signal ratio analysis.  We computed the average of CQT images for positive samples (POS) , and divided it by the average of CQT images for negative samples (NEG).  This gives us what we called a mask.  Here is the mask for above settings\n\n![mask](https://i.imgur.com/SXWVA5I.png)\n\nWe see where the waves are on average.\n\nBy taking mask average over detectors and time we see that we can crop images left and right safely (y axis labels should be multiplied by 3, sorry):\n\n ![time](https://i.imgur.com/BuBZGJo.png)\n\nCropping sides gets rid of the CQT border artifacts.\n\nBy taking the mask mean over detectors and time we get signal to noise ratio per frequency (y axis labels should be multiplied by 3, sorry):\n\n![freq](https://i.imgur.com/QTTfca3.png)\n\nx is log scale, but doing the math we see that snr is 1 outside 22Hz-352Hz\n\nWe tuned bins per octave so that the image height is a multiple of 16 (for speed), and 32 bins overall was a very good tradeoff.\n\n**Input images**\n\nWe used nnAudio CQT (or CWT) as the first model layer.  Then 3 outputs are stacked on frequency axis. We then divided the image by NEG (average of negative sample images) to get rid of noise on average.  This was quite better than whitening. We finally applied a log scaling, a bit similar to amplitude to db in audio signal processing.  An example of resulting image is shown below (ignore frequency axis labels):\n\n![image](https://i.imgur.com/ERThCPY.png)\n\nDepending on the model we added the mask as an additional channel. We then scaled data to have multiple of 16 dimensions.  For instance resizing 96x820 to 1x96x768\n\n**Model**\n\nWe mostly used efficientnet_v2s_s and efficientnet_v2_m from timm package.  Models were trained using pytorch.cuda.amp.  By using batch size and image dimensions that are multiple of 16 we use of RT cores and speed up training significantly on GPU V100.\n\nWe used BCEWithLogitsLoss, Adam and OneCycleLR most of the competition. We later found that retraining models was helping. Our best model went through 4 training cycles.  Using cosine annealing from the start would probably have been a good choice.\n\n**Augmentations**\n\nKazuki shared most of our augmentations here: https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275335\n\nWe also used a random time shift in addition to these.\n\n**Pseudo Labeling**\n\nPseudo labeling gave us a 0.0002 boost on LB when we used it.  We used non rounded predictions from teacher model for test dataset and concatenated to training fold.  It improved CV by 0.0008 on average but LB by only 0.0002.  However, our best model was not a pseudo labeling model.\n\n**Ensembling**\n\nWe kept oof predictions for all our models and trained second level models on them. We used logistic regression on logits, scipy optimize on ranked predictions, and XGBoost on ranked predictions.  They yield similar results, and averaging them did not help really.  \n\nEnsembling moved us from 8826 for best single model to 8832 on public LB.\n\n**What did not work**\n\n- Giba tried hard to generate additional positive samples.  Issue was to calibrate these to be of the same distribution as training data.  We did not find the right way in time.\n- Resnet34.  As written above, we could not meet effnetv2s performance.  And speedup was only 2x anyway.\n- We tried many other things that did not work well enough, like triplet models (a backbone on each detector then a common head), or using detectors as RGB channels.\n\nThat's it.  We clearly missed 1D model, but for the rest we did everything we could.  And I learned from my team mates!  Teaming is good!.",
      "votes": null
    },
    {
      "id": "1529482",
      "postDate": "09/30/2021 12:45:39",
      "content": "<blockquote>\n  <p>We found the best frequency range by doing a signal ratio analysis. We computed the average of CQT images for positive samples (POS) , and divided it by the average of CQT images for negative samples (NEG). This gives us what we called a mask. Here is the mask for above settings</p>\n</blockquote>\n<p>So simple and genius that it's frustrating it did not come to me to attempt this.</p>\n<blockquote>\n  <p>We then divided the image by NEG (average of negative sample images) to get rid of noise on average. This was quite better than whitening.</p>\n</blockquote>\n<p>omg nnnrrrrrrrrggh!</p>\n<p>As always, an elegant solution, well deserved.</p>\n<p>PS:</p>\n<blockquote>\n  <p>We kept oof predictions for all our models and trained second level models on them. We used logistic regression on logits, scipy optimize on ranked predictions, and XGBoost on ranked predictions. They yield similar results, and averaging them did not help really. </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> if you haven't deleted your preds, might I convince you to humor me and attempt to run SGD for your second level stacker? Interested in what the CV / late private LB of that would be…</p>",
      "rawMarkdown": "> We found the best frequency range by doing a signal ratio analysis. We computed the average of CQT images for positive samples (POS) , and divided it by the average of CQT images for negative samples (NEG). This gives us what we called a mask. Here is the mask for above settings\n\nSo simple and genius that it's frustrating it did not come to me to attempt this.\n\n> We then divided the image by NEG (average of negative sample images) to get rid of noise on average. This was quite better than whitening.\n\nomg nnnrrrrrrrrggh!\n\nAs always, an elegant solution, well deserved.\n\nPS:\n\n> We kept oof predictions for all our models and trained second level models on them. We used logistic regression on logits, scipy optimize on ranked predictions, and XGBoost on ranked predictions. They yield similar results, and averaging them did not help really. \n\n@cpmpml if you haven't deleted your preds, might I convince you to humor me and attempt to run SGD for your second level stacker? Interested in what the CV / late private LB of that would be...",
      "votes": null
    },
    {
      "id": "1529490",
      "postDate": "09/30/2021 12:52:50",
      "content": "<p>Thanks!  We thought that signal processing was they, but we totally missed the 1D avenue still.  </p>",
      "rawMarkdown": "Thanks!  We thought that signal processing was they, but we totally missed the 1D avenue still.",
      "votes": null
    },
    {
      "id": "1529515",
      "postDate": "09/30/2021 13:09:25",
      "content": "<p>\"To be honest we tried resnet34 but it did not match effnetv2 results for us. I wonder why it is different for other teams. \"</p>\n<p>Resnet34 is not a strong classifier. the performance actually comes from the learnable front-end for my case. If the backend is not strong, the model is forced to learn a better front end. just some experiment results from my side:</p>\n<p>fixed cqt + resnet34 : cv 0.872   <br>\nsame fixed cqt + effb2 : cv 0.874<br>\nsame fixed cqt + effb7 : cv 0.876</p>\n<p>learnable cqt + resnet34 : cv 0.878<br>\nlearnable cqt + effb2 : cv 0.877</p>\n<p>you can consider fixed cqt  as non optimised cqt in some sense</p>\n<p>I am surprised that the mask cqt images you showed have the GW waves extending beyond 1000Hz</p>",
      "rawMarkdown": "\"To be honest we tried resnet34 but it did not match effnetv2 results for us. I wonder why it is different for other teams. \"\n\nResnet34 is not a strong classifier. the performance actually comes from the learnable front-end for my case. If the backend is not strong, the model is forced to learn a better front end. just some experiment results from my side:\n\nfixed cqt + resnet34 : cv 0.872   \nsame fixed cqt + effb2 : cv 0.874\nsame fixed cqt + effb7 : cv 0.876\n\nlearnable cqt + resnet34 : cv 0.878\nlearnable cqt + effb2 : cv 0.877\n\nyou can consider fixed cqt  as non optimised cqt in some sense\n\nI am surprised that the mask cqt images you showed have the GW waves extending beyond 1000Hz",
      "votes": null
    },
    {
      "id": "1529527",
      "postDate": "09/30/2021 13:15:18",
      "content": "<p>The frequency labels are not right because CQT output is on a log scale.  The mask does not extend above 350 Hz as shown by the snr curve.</p>",
      "rawMarkdown": "The frequency labels are not right because CQT output is on a log scale.  The mask does not extend above 350 Hz as shown by the snr curve.",
      "votes": null
    },
    {
      "id": "1529536",
      "postDate": "09/30/2021 13:20:25",
      "content": "<p>\"We then divided the image by NEG (average of negative sample images) to get rid of noise on average. This was quite better than whitening.\"</p>\n<p>just an idea … use some different sets(random set or clustering …. even from the test set) of  NEG template as TTA (and training)? </p>",
      "rawMarkdown": "\"We then divided the image by NEG (average of negative sample images) to get rid of noise on average. This was quite better than whitening.\"\n\njust an idea ... use some different sets(random set or clustering .... even from the test set) of  NEG template as TTA (and training)?",
      "votes": null
    },
    {
      "id": "1529594",
      "postDate": "09/30/2021 14:03:55",
      "content": "<p>Neg estimate didn't vary much when subsampling, even 10k samples.  Anyway, goal is to improve snr, hence NEG average is the best noise estimate we can get..  But we made NEG and MASK trainable last few days of comp, and this improved models significantly.</p>\n<p>I wish we had tried trainable CQT too as you did.</p>",
      "rawMarkdown": "Neg estimate didn't vary much when subsampling, even 10k samples.  Anyway, goal is to improve snr, hence NEG average is the best noise estimate we can get..  But we made NEG and MASK trainable last few days of comp, and this improved models significantly.\n\n\nI wish we had tried trainable CQT too as you did.",
      "votes": null
    },
    {
      "id": "1529608",
      "postDate": "09/30/2021 14:14:00",
      "content": "<p>Thank you for sharing and congrats! Did you manage to get your STFT Transformer from Birdclef to work for this? I tried the same approach (but replacing STFT with CQT), but couldn't get decent results. Wondering if you tried the same?</p>",
      "rawMarkdown": "Thank you for sharing and congrats! Did you manage to get your STFT Transformer from Birdclef to work for this? I tried the same approach (but replacing STFT with CQT), but couldn't get decent results. Wondering if you tried the same?",
      "votes": null
    },
    {
      "id": "1529644",
      "postDate": "09/30/2021 14:53:47",
      "content": "<p>I tried transformers indeed, slow and not as good as effnets here!</p>",
      "rawMarkdown": "I tried transformers indeed, slow and not as good as effnets here!",
      "votes": null
    },
    {
      "id": "1529654",
      "postDate": "09/30/2021 15:02:00",
      "content": "<p>trainable NEG and MASK, that is a very good idea.</p>\n<p>basically, make an \"attention mask\", then make it trainable … I can see potential use in the future👍👍👍</p>",
      "rawMarkdown": "trainable NEG and MASK, that is a very good idea.\n\nbasically, make an \"attention mask\", then make it trainable ... I can see potential use in the future👍👍👍",
      "votes": null
    },
    {
      "id": "1529705",
      "postDate": "09/30/2021 15:52:03",
      "content": "<p>We've also tried, and had the same result. </p>",
      "rawMarkdown": "We've also tried, and had the same result.",
      "votes": null
    },
    {
      "id": "1531478",
      "postDate": "10/02/2021 03:28:17",
      "content": "<p>I also tried it, following your Birdclef report, but without much success( probably next time… It is really a great idea.</p>",
      "rawMarkdown": "I also tried it, following your Birdclef report, but without much success( probably next time... It is really a great idea.",
      "votes": null
    },
    {
      "id": "1532869",
      "postDate": "10/03/2021 12:58:09",
      "content": "<p>I missed that.  Which SGD you want me to run?</p>",
      "rawMarkdown": "I missed that.  Which SGD you want me to run?",
      "votes": null
    },
    {
      "id": "1549136",
      "postDate": "10/18/2021 20:20:49",
      "content": "<p>Thanks a lot for sharing an for such an understandable explanation. A couple of questions:</p>\n<blockquote>\n  <p>We later found that retraining models was helping. Our best model went through 4 training cycles.</p>\n</blockquote>\n<p>You mean 4 cycles of the scheduler, right? How did you figure out that?</p>\n<blockquote>\n  <p>But we made NEG and MASK trainable last few days of comp, and this improved models significantly.</p>\n</blockquote>\n<p>How did you do that? Considering them as a torch nn.Parameter or something like that?</p>",
      "rawMarkdown": "Thanks a lot for sharing an for such an understandable explanation. A couple of questions:\n\n> We later found that retraining models was helping. Our best model went through 4 training cycles.\n\nYou mean 4 cycles of the scheduler, right? How did you figure out that?\n\n> But we made NEG and MASK trainable last few days of comp, and this improved models significantly.\n\nHow did you do that? Considering them as a torch nn.Parameter or something like that?",
      "votes": null
    },
    {
      "id": "1559928",
      "postDate": "10/27/2021 08:24:04",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1529482,
      "author_name": "authman",
      "author_url": "",
      "post_date": "09/30/2021 12:45:39",
      "content": "<blockquote>\n  <p>We found the best frequency range by doing a signal ratio analysis. We computed the average of CQT images for positive samples (POS) , and divided it by the average of CQT images for negative samples (NEG). This gives us what we called a mask. Here is the mask for above settings</p>\n</blockquote>\n<p>So simple and genius that it's frustrating it did not come to me to attempt this.</p>\n<blockquote>\n  <p>We then divided the image by NEG (average of negative sample images) to get rid of noise on average. This was quite better than whitening.</p>\n</blockquote>\n<p>omg nnnrrrrrrrrggh!</p>\n<p>As always, an elegant solution, well deserved.</p>\n<p>PS:</p>\n<blockquote>\n  <p>We kept oof predictions for all our models and trained second level models on them. We used logistic regression on logits, scipy optimize on ranked predictions, and XGBoost on ranked predictions. They yield similar results, and averaging them did not help really. </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> if you haven't deleted your preds, might I convince you to humor me and attempt to run SGD for your second level stacker? Interested in what the CV / late private LB of that would be…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1529490,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/30/2021 12:52:50",
          "content": "<p>Thanks!  We thought that signal processing was they, but we totally missed the 1D avenue still.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1532869,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "10/03/2021 12:58:09",
          "content": "<p>I missed that.  Which SGD you want me to run?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1529515,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/30/2021 13:09:25",
      "content": "<p>\"To be honest we tried resnet34 but it did not match effnetv2 results for us. I wonder why it is different for other teams. \"</p>\n<p>Resnet34 is not a strong classifier. the performance actually comes from the learnable front-end for my case. If the backend is not strong, the model is forced to learn a better front end. just some experiment results from my side:</p>\n<p>fixed cqt + resnet34 : cv 0.872   <br>\nsame fixed cqt + effb2 : cv 0.874<br>\nsame fixed cqt + effb7 : cv 0.876</p>\n<p>learnable cqt + resnet34 : cv 0.878<br>\nlearnable cqt + effb2 : cv 0.877</p>\n<p>you can consider fixed cqt  as non optimised cqt in some sense</p>\n<p>I am surprised that the mask cqt images you showed have the GW waves extending beyond 1000Hz</p>",
      "votes": null,
      "replies": [
        {
          "id": 1529527,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/30/2021 13:15:18",
          "content": "<p>The frequency labels are not right because CQT output is on a log scale.  The mask does not extend above 350 Hz as shown by the snr curve.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1529536,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/30/2021 13:20:25",
      "content": "<p>\"We then divided the image by NEG (average of negative sample images) to get rid of noise on average. This was quite better than whitening.\"</p>\n<p>just an idea … use some different sets(random set or clustering …. even from the test set) of  NEG template as TTA (and training)? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1529594,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/30/2021 14:03:55",
          "content": "<p>Neg estimate didn't vary much when subsampling, even 10k samples.  Anyway, goal is to improve snr, hence NEG average is the best noise estimate we can get..  But we made NEG and MASK trainable last few days of comp, and this improved models significantly.</p>\n<p>I wish we had tried trainable CQT too as you did.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1529654,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/30/2021 15:02:00",
          "content": "<p>trainable NEG and MASK, that is a very good idea.</p>\n<p>basically, make an \"attention mask\", then make it trainable … I can see potential use in the future👍👍👍</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1529608,
      "author_name": "anjum48",
      "author_url": "",
      "post_date": "09/30/2021 14:14:00",
      "content": "<p>Thank you for sharing and congrats! Did you manage to get your STFT Transformer from Birdclef to work for this? I tried the same approach (but replacing STFT with CQT), but couldn't get decent results. Wondering if you tried the same?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1529644,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/30/2021 14:53:47",
          "content": "<p>I tried transformers indeed, slow and not as good as effnets here!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1529705,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "09/30/2021 15:52:03",
          "content": "<p>We've also tried, and had the same result. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1531478,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/02/2021 03:28:17",
          "content": "<p>I also tried it, following your Birdclef report, but without much success( probably next time… It is really a great idea.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1549136,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "10/18/2021 20:20:49",
      "content": "<p>Thanks a lot for sharing an for such an understandable explanation. A couple of questions:</p>\n<blockquote>\n  <p>We later found that retraining models was helping. Our best model went through 4 training cycles.</p>\n</blockquote>\n<p>You mean 4 cycles of the scheduler, right? How did you figure out that?</p>\n<blockquote>\n  <p>But we made NEG and MASK trainable last few days of comp, and this improved models significantly.</p>\n</blockquote>\n<p>How did you do that? Considering them as a torch nn.Parameter or something like that?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1559928,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 08:24:04",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1529466": "I want to thank Kaggle and Host for this amazing challenge. Topic was great, and teaming with @onodera and @titericz was great too.\n\nWe led the competition for a while then all of a sudden teams started to pass us.  In hindsight we know it is thanks to using 1D models and/or resnets. Unfortunately for us we kept working with efficientnet v2 as we were having good results so far.  We have no excuse given both ideas were shared in the forum soon enough to be leveraged.  To be honest we tried resnet34 but it did not match effnetv2 results for us.  I wonder why it is different for other teams.  We also tried 1D models briefly, but, same, they were low quality compared to effnets.  We should have submitted a blend obviously.  \n\nHere is a snapshot of our solution.\n\n**Data processing**\n\nData was multiplied by 1e19 (1e21 later) to ensure that CQT and CWT could be performed correctly using FP16.\n\nCQT and CWT (results are very similar).  Best model  (public 0.8826 private 0.8805)  used these setting for CQT:\n\n```\nsr = 2048\nhop_length = 5\nfmin = 22\nfmax = 22*16\nbins_per_octave = 8\nn_octaves = 4\nn_bins = n_octaves * bins_per_octave\nfscale = 1\ncqt = CQT1992v2(sr=sr, hop_length=hop_length, fmin=fmin, fmax=fmax,\n                        n_bins=n_bins, bins_per_octave=bins_per_octave * fscale,         \n                window=('kaiser', 14), filter_scale=1/fscale,\n               )\n```\n\nWe used fscale = 2 (i.e. filter_scale = 0.5) in few models, it improves a bit but we could not train them as much as our best model.  We also varied hop length to get different image sizes. We then concatenated the 3 CQT images on the frequency dim, getting images of size 96xL where L depends on hop length (L = 820 for the above setting).  \n\n**Signal To Noise EDA**\n\nWe found the best frequency range by doing a signal ratio analysis.  We computed the average of CQT images for positive samples (POS) , and divided it by the average of CQT images for negative samples (NEG).  This gives us what we called a mask.  Here is the mask for above settings\n\n![mask](https://i.imgur.com/SXWVA5I.png)\n\nWe see where the waves are on average.\n\nBy taking mask average over detectors and time we see that we can crop images left and right safely (y axis labels should be multiplied by 3, sorry):\n\n ![time](https://i.imgur.com/BuBZGJo.png)\n\nCropping sides gets rid of the CQT border artifacts.\n\nBy taking the mask mean over detectors and time we get signal to noise ratio per frequency (y axis labels should be multiplied by 3, sorry):\n\n![freq](https://i.imgur.com/QTTfca3.png)\n\nx is log scale, but doing the math we see that snr is 1 outside 22Hz-352Hz\n\nWe tuned bins per octave so that the image height is a multiple of 16 (for speed), and 32 bins overall was a very good tradeoff.\n\n**Input images**\n\nWe used nnAudio CQT (or CWT) as the first model layer.  Then 3 outputs are stacked on frequency axis. We then divided the image by NEG (average of negative sample images) to get rid of noise on average.  This was quite better than whitening. We finally applied a log scaling, a bit similar to amplitude to db in audio signal processing.  An example of resulting image is shown below (ignore frequency axis labels):\n\n![image](https://i.imgur.com/ERThCPY.png)\n\nDepending on the model we added the mask as an additional channel. We then scaled data to have multiple of 16 dimensions.  For instance resizing 96x820 to 1x96x768\n\n**Model**\n\nWe mostly used efficientnet_v2s_s and efficientnet_v2_m from timm package.  Models were trained using pytorch.cuda.amp.  By using batch size and image dimensions that are multiple of 16 we use of RT cores and speed up training significantly on GPU V100.\n\nWe used BCEWithLogitsLoss, Adam and OneCycleLR most of the competition. We later found that retraining models was helping. Our best model went through 4 training cycles.  Using cosine annealing from the start would probably have been a good choice.\n\n**Augmentations**\n\nKazuki shared most of our augmentations here: https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275335\n\nWe also used a random time shift in addition to these.\n\n**Pseudo Labeling**\n\nPseudo labeling gave us a 0.0002 boost on LB when we used it.  We used non rounded predictions from teacher model for test dataset and concatenated to training fold.  It improved CV by 0.0008 on average but LB by only 0.0002.  However, our best model was not a pseudo labeling model.\n\n**Ensembling**\n\nWe kept oof predictions for all our models and trained second level models on them. We used logistic regression on logits, scipy optimize on ranked predictions, and XGBoost on ranked predictions.  They yield similar results, and averaging them did not help really.  \n\nEnsembling moved us from 8826 for best single model to 8832 on public LB.\n\n**What did not work**\n\n- Giba tried hard to generate additional positive samples.  Issue was to calibrate these to be of the same distribution as training data.  We did not find the right way in time.\n- Resnet34.  As written above, we could not meet effnetv2s performance.  And speedup was only 2x anyway.\n- We tried many other things that did not work well enough, like triplet models (a backbone on each detector then a common head), or using detectors as RGB channels.\n\nThat's it.  We clearly missed 1D model, but for the rest we did everything we could.  And I learned from my team mates!  Teaming is good!.",
    "1529482": "> We found the best frequency range by doing a signal ratio analysis. We computed the average of CQT images for positive samples (POS) , and divided it by the average of CQT images for negative samples (NEG). This gives us what we called a mask. Here is the mask for above settings\n\nSo simple and genius that it's frustrating it did not come to me to attempt this.\n\n> We then divided the image by NEG (average of negative sample images) to get rid of noise on average. This was quite better than whitening.\n\nomg nnnrrrrrrrrggh!\n\nAs always, an elegant solution, well deserved.\n\nPS:\n\n> We kept oof predictions for all our models and trained second level models on them. We used logistic regression on logits, scipy optimize on ranked predictions, and XGBoost on ranked predictions. They yield similar results, and averaging them did not help really. \n\n@cpmpml if you haven't deleted your preds, might I convince you to humor me and attempt to run SGD for your second level stacker? Interested in what the CV / late private LB of that would be...",
    "1529490": "Thanks!  We thought that signal processing was they, but we totally missed the 1D avenue still.",
    "1529515": "\"To be honest we tried resnet34 but it did not match effnetv2 results for us. I wonder why it is different for other teams. \"\n\nResnet34 is not a strong classifier. the performance actually comes from the learnable front-end for my case. If the backend is not strong, the model is forced to learn a better front end. just some experiment results from my side:\n\nfixed cqt + resnet34 : cv 0.872   \nsame fixed cqt + effb2 : cv 0.874\nsame fixed cqt + effb7 : cv 0.876\n\nlearnable cqt + resnet34 : cv 0.878\nlearnable cqt + effb2 : cv 0.877\n\nyou can consider fixed cqt  as non optimised cqt in some sense\n\nI am surprised that the mask cqt images you showed have the GW waves extending beyond 1000Hz",
    "1529527": "The frequency labels are not right because CQT output is on a log scale.  The mask does not extend above 350 Hz as shown by the snr curve.",
    "1529536": "\"We then divided the image by NEG (average of negative sample images) to get rid of noise on average. This was quite better than whitening.\"\n\njust an idea ... use some different sets(random set or clustering .... even from the test set) of  NEG template as TTA (and training)?",
    "1529594": "Neg estimate didn't vary much when subsampling, even 10k samples.  Anyway, goal is to improve snr, hence NEG average is the best noise estimate we can get..  But we made NEG and MASK trainable last few days of comp, and this improved models significantly.\n\n\nI wish we had tried trainable CQT too as you did.",
    "1529608": "Thank you for sharing and congrats! Did you manage to get your STFT Transformer from Birdclef to work for this? I tried the same approach (but replacing STFT with CQT), but couldn't get decent results. Wondering if you tried the same?",
    "1529644": "I tried transformers indeed, slow and not as good as effnets here!",
    "1529654": "trainable NEG and MASK, that is a very good idea.\n\nbasically, make an \"attention mask\", then make it trainable ... I can see potential use in the future👍👍👍",
    "1529705": "We've also tried, and had the same result.",
    "1531478": "I also tried it, following your Birdclef report, but without much success( probably next time... It is really a great idea.",
    "1532869": "I missed that.  Which SGD you want me to run?",
    "1549136": "Thanks a lot for sharing an for such an understandable explanation. A couple of questions:\n\n> We later found that retraining models was helping. Our best model went through 4 training cycles.\n\nYou mean 4 cycles of the scheduler, right? How did you figure out that?\n\n> But we made NEG and MASK trainable last few days of comp, and this improved models significantly.\n\nHow did you do that? Considering them as a torch nn.Parameter or something like that?",
    "1559928": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}