{
  "id": 275318,
  "title": "rank18th solution : private/public LB(0.8806/0.8822) learnable CQT + resnet34",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/275318",
  "author_name": "hengck23",
  "post_date": "2021-09-29T22:46:00.750000",
  "votes": 60,
  "comment_count": 18,
  "views": 0,
  "content": "<p>the secret of resnet34? here are some initial writeup:<br>\n<img src=\"https://i.ibb.co/6WYFtKN/Selection-933.png\" alt=\"https://i.ibb.co/6WYFtKN/Selection-933.png\"><br>\n<img src=\"https://i.ibb.co/vXb6zqX/Selection-932.png\" alt=\"https://i.ibb.co/vXb6zqX/Selection-932.png\"><br>\nlearnable QT code: <a href=\"https://gist.github.com/hengck23/0dd338bed8aad0e98b7477a6f17f9de7\" target=\"_blank\">https://gist.github.com/hengck23/0dd338bed8aad0e98b7477a6f17f9de7</a></p>\n<pre><code>      #----\n     self.register_buffer('freq', freq)\n     self.register_buffer('length', length)\n     self.register_buffer('cqt_window', cqt_window) #this is not trainable !!!!\n      if trainable:\n         self.cqt_kernel_real = nn.Parameter(cqt_kernel_real)\n         self.cqt_kernel_imag = nn.Parameter(cqt_kernel_imag)\n     else:\n         self.register_buffer('cqt_kernel_real', cqt_kernel_real)\n         self.register_buffer('cqt_kernel_imag', cqt_kernel_imag)\n</code></pre>\n<p>just by plugging a learnable CQT to resnet34, you will easily get CV in the range of 0.880 and LB in the range of 0.881.</p>\n<ul>\n<li>some important points when training:<br>\n(1)  pre-trained with initial cqt kernels first. set trainable=False<br>\n(2)  after one epoch, you can set trainable=True<br>\n(3)  for TTA to work, use with shifted-versions of the sample wave as train augmentation<br>\n(4)  bandpass is a necessity. Tapering is good to have.</li>\n<li>I cannot get whitening to work, so my solution don't use whitening. But I think it may lead to better solution?</li>\n<li>Lastly, I am grateful to Z by HP &amp; NVIDIA for providing a powerful workstation HP Z8 with two Quadro RTX 8000 (48 GB VRAM each) for this competition. For resnet34 and learnable CQT of size 128x256, it takes only** 10 min for one epoch**. This lets me perform many experiments in very short time. With large memory, I can try efficient net b7 with large image size. And for training transformers, I can disable torch.amp (to avoid NAN in attention computation) too.</li>\n</ul>",
  "messages": [
    {
      "id": 1528710,
      "postDate": "2021-09-29T22:46:00.750Z",
      "content": "<p>the secret of resnet34? here are some initial writeup:<br>\n<img src=\"https://i.ibb.co/6WYFtKN/Selection-933.png\" alt=\"https://i.ibb.co/6WYFtKN/Selection-933.png\"><br>\n<img src=\"https://i.ibb.co/vXb6zqX/Selection-932.png\" alt=\"https://i.ibb.co/vXb6zqX/Selection-932.png\"><br>\nlearnable QT code: <a href=\"https://gist.github.com/hengck23/0dd338bed8aad0e98b7477a6f17f9de7\" target=\"_blank\">https://gist.github.com/hengck23/0dd338bed8aad0e98b7477a6f17f9de7</a></p>\n<pre><code>      #----\n     self.register_buffer('freq', freq)\n     self.register_buffer('length', length)\n     self.register_buffer('cqt_window', cqt_window) #this is not trainable !!!!\n      if trainable:\n         self.cqt_kernel_real = nn.Parameter(cqt_kernel_real)\n         self.cqt_kernel_imag = nn.Parameter(cqt_kernel_imag)\n     else:\n         self.register_buffer('cqt_kernel_real', cqt_kernel_real)\n         self.register_buffer('cqt_kernel_imag', cqt_kernel_imag)\n</code></pre>\n<p>just by plugging a learnable CQT to resnet34, you will easily get CV in the range of 0.880 and LB in the range of 0.881.</p>\n<ul>\n<li>some important points when training:<br>\n(1)  pre-trained with initial cqt kernels first. set trainable=False<br>\n(2)  after one epoch, you can set trainable=True<br>\n(3)  for TTA to work, use with shifted-versions of the sample wave as train augmentation<br>\n(4)  bandpass is a necessity. Tapering is good to have.</li>\n<li>I cannot get whitening to work, so my solution don't use whitening. But I think it may lead to better solution?</li>\n<li>Lastly, I am grateful to Z by HP &amp; NVIDIA for providing a powerful workstation HP Z8 with two Quadro RTX 8000 (48 GB VRAM each) for this competition. For resnet34 and learnable CQT of size 128x256, it takes only** 10 min for one epoch**. This lets me perform many experiments in very short time. With large memory, I can try efficient net b7 with large image size. And for training transformers, I can disable torch.amp (to avoid NAN in attention computation) too.</li>\n</ul>",
      "rawMarkdown": " \n\nthe secret of resnet34? here are some initial writeup:\n\n![https://i.ibb.co/6WYFtKN/Selection-933.png](https://i.ibb.co/6WYFtKN/Selection-933.png)\n![https://i.ibb.co/vXb6zqX/Selection-932.png](https://i.ibb.co/vXb6zqX/Selection-932.png)\n\nlearnable QT code: https://gist.github.com/hengck23/0dd338bed8aad0e98b7477a6f17f9de7\n\n\n```\n\n        #----\n        self.register_buffer('freq', freq)\n        self.register_buffer('length', length)\n        self.register_buffer('cqt_window', cqt_window) #this is not trainable !!!!\n\n        if trainable:\n            self.cqt_kernel_real = nn.Parameter(cqt_kernel_real)\n            self.cqt_kernel_imag = nn.Parameter(cqt_kernel_imag)\n        else:\n            self.register_buffer('cqt_kernel_real', cqt_kernel_real)\n            self.register_buffer('cqt_kernel_imag', cqt_kernel_imag)\n\n```\n\n\n\n\njust by plugging a learnable CQT to resnet34, you will easily get CV in the range of 0.880 and LB in the range of 0.881.\n\n- some important points when training:\n\n(1)  pre-trained with initial cqt kernels first. set trainable=False\n(2)  after one epoch, you can set trainable=True\n(3)  for TTA to work, use with shifted-versions of the sample wave as train augmentation\n(4)  bandpass is a necessity. Tapering is good to have.\n\n\n\n\n- I cannot get whitening to work, so my solution don't use whitening. But I think it may lead to better solution?\n\n \n- Lastly, I am grateful to Z by HP & NVIDIA for providing a powerful workstation HP Z8 with two Quadro RTX 8000 (48 GB VRAM each) for this competition. For resnet34 and learnable CQT of size 128x256, it takes only** 10 min for one epoch**. This lets me perform many experiments in very short time. With large memory, I can try efficient net b7 with large image size. And for training transformers, I can disable torch.amp (to avoid NAN in attention computation) too.\n",
      "votes": 58
    },
    {
      "id": 1528723,
      "postDate": "2021-09-29T23:02:09.960Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> - Just don't be the second person to accidentally share a top solution too early!</p>",
      "rawMarkdown": "@hengck23 - Just don't be the second person to accidentally share a top solution too early!",
      "votes": 3,
      "replies": [
        {
          "id": 1528730,
          "postDate": "2021-09-29T23:14:51.543Z",
          "content": "<p>i don't have solution in kaggle notebook, so it won't happen</p>",
          "rawMarkdown": "i don't have solution in kaggle notebook, so it won't happen",
          "votes": 2
        }
      ]
    },
    {
      "id": 1531141,
      "postDate": "2021-10-01T16:47:31.303Z",
      "content": "<p>visualisation of what is learned:</p>\n<p><a href=\"https://ibb.co/j4z0TVk\"><img src=\"https://i.ibb.co/GP9Y0sQ/Selection-985.png\" alt=\"Selection-985\"></a><br>\n<a href=\"https://ibb.co/zmVYSLg\"><img src=\"https://i.ibb.co/tHptbN6/Selection-986.png\" alt=\"Selection-986\"></a><br>\n<a href=\"https://ibb.co/xFQJRgF\"><img src=\"https://i.ibb.co/cYdD78Y/Selection-987.png\" alt=\"Selection-987\"></a></p>",
      "rawMarkdown": "visualisation of what is learned:\n\n<a href=\"https://ibb.co/j4z0TVk\"><img src=\"https://i.ibb.co/GP9Y0sQ/Selection-985.png\" alt=\"Selection-985\" border=\"0\"></a>\n<a href=\"https://ibb.co/zmVYSLg\"><img src=\"https://i.ibb.co/tHptbN6/Selection-986.png\" alt=\"Selection-986\" border=\"0\"></a>\n<a href=\"https://ibb.co/xFQJRgF\"><img src=\"https://i.ibb.co/cYdD78Y/Selection-987.png\" alt=\"Selection-987\" border=\"0\"></a>\n",
      "votes": 1
    },
    {
      "id": 1528785,
      "postDate": "2021-09-30T00:51:12.803Z",
      "content": "<p>for us, using Radam (+lookahead) was enough to get trainable cqt to work. no need to freeze it/pre train first</p>",
      "rawMarkdown": "for us, using Radam (+lookahead) was enough to get trainable cqt to work. no need to freeze it/pre train first",
      "votes": 2,
      "replies": [
        {
          "id": 1528890,
          "postDate": "2021-09-30T02:46:52.290Z",
          "content": "<p><a href=\"https://www.kaggle.com/proletheus\" target=\"_blank\">@proletheus</a> <br>\nthanks for the feedback.</p>\n<p>I had compared results with and without pretraining. For me, pre-training works better. I think it also depends on my augmentation and regularisation (e.g. dropout).</p>\n<p>I use Radam (+lookahead) as well.</p>",
          "rawMarkdown": "@proletheus \nthanks for the feedback.\n\nI had compared results with and without pretraining. For me, pre-training works better. I think it also depends on my augmentation and regularisation (e.g. dropout).\n\nI use Radam (+lookahead) as well.",
          "votes": 1
        },
        {
          "id": 1528997,
          "postDate": "2021-09-30T04:42:02.277Z",
          "content": "<p>i used it because of your bms kernel;) i noticed that the loss would first decrease (0.7-&gt; .5) then increase again with trainabble cqts, so i also tried freezing it first but it would go back to 0.7 loss within a few batches when using just adam</p>",
          "rawMarkdown": "i used it because of your bms kernel;) i noticed that the loss would first decrease (0.7-> .5) then increase again with trainabble cqts, so i also tried freezing it first but it would go back to 0.7 loss within a few batches when using just adam"
        }
      ]
    },
    {
      "id": 1528715,
      "postDate": "2021-09-29T22:49:27.610Z",
      "content": "<p>I was really expecting it to go all the way to the zero at first 😀</p>",
      "rawMarkdown": "I was really expecting it to go all the way to the zero at first 😀",
      "votes": 2
    },
    {
      "id": 1559899,
      "postDate": "2021-10-27T08:07:07.940Z",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
    },
    {
      "id": 1529641,
      "postDate": "2021-09-30T14:47:41.267Z",
      "content": "<p>I used the similar approach as yours.<br>\nCWT/CQT is a kind of 1D Conv in nature. But it is only one layer and the kernel is too big. So  I just replaced the CWT/CQT with a trainable multiple layers 1D Conv. <br>\nThe detail is here.<br>\n<a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275353\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275353</a></p>",
      "rawMarkdown": "I used the similar approach as yours.\nCWT/CQT is a kind of 1D Conv in nature. But it is only one layer and the kernel is too big. So  I just replaced the CWT/CQT with a trainable multiple layers 1D Conv. \nThe detail is here.\nhttps://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275353",
      "replies": [
        {
          "id": 1529661,
          "postDate": "2021-09-30T15:06:05.277Z",
          "content": "<p>thanks for the link.</p>\n<p>I tried rectangle window. In some cases, hann or gaussian-shaped windows work better. It all has to depend on the value of Q (cycles in your kernel size) and whether you need to localize/restrict frequency response to large or small bandwidth</p>",
          "rawMarkdown": "thanks for the link.\n\nI tried rectangle window. In some cases, hann or gaussian-shaped windows work better. It all has to depend on the value of Q (cycles in your kernel size) and whether you need to localize/restrict frequency response to large or small bandwidth"
        }
      ]
    },
    {
      "id": 1529044,
      "postDate": "2021-09-30T05:39:40.733Z",
      "content": "<p>whitening is only useful if you explore the cross correlation between the 3 signals. I tried many different correlation methods, but got no luck to make it work with this low SNR dataset</p>",
      "rawMarkdown": "whitening is only useful if you explore the cross correlation between the 3 signals. I tried many different correlation methods, but got no luck to make it work with this low SNR dataset"
    },
    {
      "id": 1528880,
      "postDate": "2021-09-30T02:38:14.483Z",
      "content": "<p>Ha somehow my ResNet34 is trained the same way…but is no where as good. I guess the devil is really in the details.</p>\n<p><img src=\"https://i.imgur.com/AnD86Ji.png\" alt=\"Imgur\"></p>",
      "rawMarkdown": "Ha somehow my ResNet34 is trained the same way...but is no where as good. I guess the devil is really in the details.\n\n![Imgur](https://i.imgur.com/AnD86Ji.png)",
      "replies": [
        {
          "id": 1529247,
          "postDate": "2021-09-30T08:48:56.747Z",
          "content": "<p><a href=\"https://www.kaggle.com/scaomath\" target=\"_blank\">@scaomath</a> if you are making post-competition submission, how about starting a slack team and send me an invite email.</p>\n<p>i will send you my code to repeat results and take a look at your results. Maybe I can offer advice on how to debug your algorithm.</p>\n<p>after working on this for many weeks (and falling many attempts), maybe I can pinpoint what goes wrong.</p>",
          "rawMarkdown": "@scaomath if you are making post-competition submission, how about starting a slack team and send me an invite email.\n\ni will send you my code to repeat results and take a look at your results. Maybe I can offer advice on how to debug your algorithm.\n\nafter working on this for many weeks (and falling many attempts), maybe I can pinpoint what goes wrong.\n",
          "votes": 2
        },
        {
          "id": 1530006,
          "postDate": "2021-09-30T20:58:45.403Z",
          "content": "<p>I just created a post-competition Slack channel and have invited you. Greatly appreciated.</p>",
          "rawMarkdown": "I just created a post-competition Slack channel and have invited you. Greatly appreciated."
        }
      ]
    },
    {
      "id": 1528806,
      "postDate": "2021-09-30T01:13:32.210Z",
      "content": "<p>I may see what you mean 1dcnn_resnet34?</p>",
      "rawMarkdown": "I may see what you mean 1dcnn_resnet34?"
    },
    {
      "id": 1528796,
      "postDate": "2021-09-30T01:06:55.690Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1528947,
      "postDate": "2021-09-30T03:50:23.723Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 1528783,
      "postDate": "2021-09-30T00:49:22.063Z",
      "content": "<p>Great Points! Thank you for sharing !!!</p>",
      "rawMarkdown": "Great Points! Thank you for sharing !!!"
    }
  ],
  "comments": [
    {
      "id": 1528723,
      "author_name": "John Mitchell",
      "author_url": "",
      "post_date": "2021-09-29T23:02:09.960000",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> - Just don't be the second person to accidentally share a top solution too early!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1528730,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-09-29T23:14:51.543000",
          "content": "<p>i don't have solution in kaggle notebook, so it won't happen</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1531141,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-10-01T16:47:31.303000",
      "content": "<p>visualisation of what is learned:</p>\n<p><a href=\"https://ibb.co/j4z0TVk\"><img src=\"https://i.ibb.co/GP9Y0sQ/Selection-985.png\" alt=\"Selection-985\"></a><br>\n<a href=\"https://ibb.co/zmVYSLg\"><img src=\"https://i.ibb.co/tHptbN6/Selection-986.png\" alt=\"Selection-986\"></a><br>\n<a href=\"https://ibb.co/xFQJRgF\"><img src=\"https://i.ibb.co/cYdD78Y/Selection-987.png\" alt=\"Selection-987\"></a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1528785,
      "author_name": "Proletheus",
      "author_url": "",
      "post_date": "2021-09-30T00:51:12.803000",
      "content": "<p>for us, using Radam (+lookahead) was enough to get trainable cqt to work. no need to freeze it/pre train first</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1528890,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-09-30T02:46:52.290000",
          "content": "<p><a href=\"https://www.kaggle.com/proletheus\" target=\"_blank\">@proletheus</a> <br>\nthanks for the feedback.</p>\n<p>I had compared results with and without pretraining. For me, pre-training works better. I think it also depends on my augmentation and regularisation (e.g. dropout).</p>\n<p>I use Radam (+lookahead) as well.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1528997,
          "author_name": "Proletheus",
          "author_url": "",
          "post_date": "2021-09-30T04:42:02.277000",
          "content": "<p>i used it because of your bms kernel;) i noticed that the loss would first decrease (0.7-&gt; .5) then increase again with trainabble cqts, so i also tried freezing it first but it would go back to 0.7 loss within a few batches when using just adam</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1528715,
      "author_name": "Sinan Calisir",
      "author_url": "",
      "post_date": "2021-09-29T22:49:27.610000",
      "content": "<p>I was really expecting it to go all the way to the zero at first 😀</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1559899,
      "author_name": "ChristopherZerafa",
      "author_url": "",
      "post_date": "2021-10-27T08:07:07.940000",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1529641,
      "author_name": "Johnny Lee",
      "author_url": "",
      "post_date": "2021-09-30T14:47:41.267000",
      "content": "<p>I used the similar approach as yours.<br>\nCWT/CQT is a kind of 1D Conv in nature. But it is only one layer and the kernel is too big. So  I just replaced the CWT/CQT with a trainable multiple layers 1D Conv. <br>\nThe detail is here.<br>\n<a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275353\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275353</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 1529661,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-09-30T15:06:05.277000",
          "content": "<p>thanks for the link.</p>\n<p>I tried rectangle window. In some cases, hann or gaussian-shaped windows work better. It all has to depend on the value of Q (cycles in your kernel size) and whether you need to localize/restrict frequency response to large or small bandwidth</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1529044,
      "author_name": "jobless physicist",
      "author_url": "",
      "post_date": "2021-09-30T05:39:40.733000",
      "content": "<p>whitening is only useful if you explore the cross correlation between the 3 signals. I tried many different correlation methods, but got no luck to make it work with this low SNR dataset</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1528880,
      "author_name": "Shuhao Cao",
      "author_url": "",
      "post_date": "2021-09-30T02:38:14.483000",
      "content": "<p>Ha somehow my ResNet34 is trained the same way…but is no where as good. I guess the devil is really in the details.</p>\n<p><img src=\"https://i.imgur.com/AnD86Ji.png\" alt=\"Imgur\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 1529247,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-09-30T08:48:56.747000",
          "content": "<p><a href=\"https://www.kaggle.com/scaomath\" target=\"_blank\">@scaomath</a> if you are making post-competition submission, how about starting a slack team and send me an invite email.</p>\n<p>i will send you my code to repeat results and take a look at your results. Maybe I can offer advice on how to debug your algorithm.</p>\n<p>after working on this for many weeks (and falling many attempts), maybe I can pinpoint what goes wrong.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1530006,
          "author_name": "Shuhao Cao",
          "author_url": "",
          "post_date": "2021-09-30T20:58:45.403000",
          "content": "<p>I just created a post-competition Slack channel and have invited you. Greatly appreciated.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1528806,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "2021-09-30T01:13:32.210000",
      "content": "<p>I may see what you mean 1dcnn_resnet34?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1528796,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-09-30T01:06:55.690000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1528947,
      "author_name": "YOg19",
      "author_url": "",
      "post_date": "2021-09-30T03:50:23.723000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1528783,
      "author_name": "Chansey Zheng",
      "author_url": "",
      "post_date": "2021-09-30T00:49:22.063000",
      "content": "<p>Great Points! Thank you for sharing !!!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1528710": " \n\nthe secret of resnet34? here are some initial writeup:\n\n![https://i.ibb.co/6WYFtKN/Selection-933.png](https://i.ibb.co/6WYFtKN/Selection-933.png)\n![https://i.ibb.co/vXb6zqX/Selection-932.png](https://i.ibb.co/vXb6zqX/Selection-932.png)\n\nlearnable QT code: https://gist.github.com/hengck23/0dd338bed8aad0e98b7477a6f17f9de7\n\n\n```\n\n        #----\n        self.register_buffer('freq', freq)\n        self.register_buffer('length', length)\n        self.register_buffer('cqt_window', cqt_window) #this is not trainable !!!!\n\n        if trainable:\n            self.cqt_kernel_real = nn.Parameter(cqt_kernel_real)\n            self.cqt_kernel_imag = nn.Parameter(cqt_kernel_imag)\n        else:\n            self.register_buffer('cqt_kernel_real', cqt_kernel_real)\n            self.register_buffer('cqt_kernel_imag', cqt_kernel_imag)\n\n```\n\n\n\n\njust by plugging a learnable CQT to resnet34, you will easily get CV in the range of 0.880 and LB in the range of 0.881.\n\n- some important points when training:\n\n(1)  pre-trained with initial cqt kernels first. set trainable=False\n(2)  after one epoch, you can set trainable=True\n(3)  for TTA to work, use with shifted-versions of the sample wave as train augmentation\n(4)  bandpass is a necessity. Tapering is good to have.\n\n\n\n\n- I cannot get whitening to work, so my solution don't use whitening. But I think it may lead to better solution?\n\n \n- Lastly, I am grateful to Z by HP & NVIDIA for providing a powerful workstation HP Z8 with two Quadro RTX 8000 (48 GB VRAM each) for this competition. For resnet34 and learnable CQT of size 128x256, it takes only** 10 min for one epoch**. This lets me perform many experiments in very short time. With large memory, I can try efficient net b7 with large image size. And for training transformers, I can disable torch.amp (to avoid NAN in attention computation) too.\n",
    "1528723": "@hengck23 - Just don't be the second person to accidentally share a top solution too early!",
    "1531141": "visualisation of what is learned:\n\n<a href=\"https://ibb.co/j4z0TVk\"><img src=\"https://i.ibb.co/GP9Y0sQ/Selection-985.png\" alt=\"Selection-985\" border=\"0\"></a>\n<a href=\"https://ibb.co/zmVYSLg\"><img src=\"https://i.ibb.co/tHptbN6/Selection-986.png\" alt=\"Selection-986\" border=\"0\"></a>\n<a href=\"https://ibb.co/xFQJRgF\"><img src=\"https://i.ibb.co/cYdD78Y/Selection-987.png\" alt=\"Selection-987\" border=\"0\"></a>\n",
    "1528785": "for us, using Radam (+lookahead) was enough to get trainable cqt to work. no need to freeze it/pre train first",
    "1528715": "I was really expecting it to go all the way to the zero at first 😀",
    "1559899": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
    "1529641": "I used the similar approach as yours.\nCWT/CQT is a kind of 1D Conv in nature. But it is only one layer and the kernel is too big. So  I just replaced the CWT/CQT with a trainable multiple layers 1D Conv. \nThe detail is here.\nhttps://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/275353",
    "1529044": "whitening is only useful if you explore the cross correlation between the 3 signals. I tried many different correlation methods, but got no luck to make it work with this low SNR dataset",
    "1528880": "Ha somehow my ResNet34 is trained the same way...but is no where as good. I guess the devil is really in the details.\n\n![Imgur](https://i.imgur.com/AnD86Ji.png)",
    "1528806": "I may see what you mean 1dcnn_resnet34?",
    "1528796": "",
    "1528947": "Thanks for sharing!",
    "1528783": "Great Points! Thank you for sharing !!!"
  }
}