{
  "id": 275341,
  "title": "2nd Place Solution: trainable custom frontend [EventHorizon] ",
  "url": "/competitions/g2net-gravitational-wave-detection/writeups/aillis-jp-eventhorizon-2nd-place-solution-trainabl",
  "author_name": "",
  "post_date": "2021-10-07T07:27:27.367Z",
  "votes": 127,
  "comment_count": 17,
  "views": 0,
  "content": "<p>First of all, I would like to express deep gratitude to organizers and all the teams for making this competition so interesting and exciting. Also, I want to express a big congratulations to the first place, who dominated this competition with a single ResNet34 model :)</p>\n<p>Code is available at: <a href=\"https://github.com/analokmaus/kaggle-g2net-public\" target=\"_blank\">https://github.com/analokmaus/kaggle-g2net-public</a></p>\n<h1>Common settings</h1>\n<p>num epochs = 8<br>\noptimizer = Adam<br>\nscheduler = CosineAnnealingWarmRestarts(8)<br>\nloss function = BCE<br>\ncross validation = target stratified 5 fold cross validation</p>\n<h1>Frontend architectures</h1>\n<p>Neural network architecture played the most important role for improving the performance.<br>\nHere are the frontend architectures I used. <br>\nTrainable frontend in general outperformed fixed frontend.<br>\n<img src=\"https://pbs.twimg.com/media/FAfu6LJVEAA8u53?format=jpg&amp;name=medium\" alt=\"\"><br>\n<img src=\"https://pbs.twimg.com/media/FAfu7VvUYAYjT1y?format=jpg&amp;name=medium\" alt=\"\"><br>\n<img src=\"https://pbs.twimg.com/media/FAfu7_RUUAQM2x4?format=jpg&amp;name=medium\" alt=\"\"><br>\n<img src=\"https://pbs.twimg.com/media/FAfu8t2VgAgWiUX?format=jpg&amp;name=medium\" alt=\"\"></p>\n<h1>Preprocessing</h1>\n<p>I applied bandpass filter to all networks. [16, 512] for CWT-CNN and Trainable frontend CNN, [30, 300] for 1d-CNN. Whitening did not work.</p>\n<h1>Augmentations</h1>\n<p>I tested several types of augmentations on wave and spectrogram, and only a few of wave augmentations worked. I added gaussian noise for 2d-CNN networks, and flipped wave amplitude for 1d-CNN.</p>\n<h1>Pseudo-labeling</h1>\n<p>Re-training on soft(continuous) pseudo-labelled test dataset improved AUC by ~0.001. Label smoothing during pseudo-label also helped a bit. </p>\n<h1>Stacking</h1>\n<p>I kept oofs and predictions from all my experiments. <br>\nCross validated Ridge regression model was used to combine the outputs from neural network models. <br>\nA constant improvement in both CV and LB was observed as I add more model into the stacking model.<br>\nFinally, I run greedy model selection and chose 20(/10/5) models to maximize CV. </p>\n<p><strong>Stacking 20 models: CV 0.88283 / Public LB 0.8845 / Private LB 0.8829</strong><br>\nStacking 10 models: CV 0.88270 / Public LB 0.8842 / Private LB 0.8827<br>\nStacking 5 models: CV 0.88242 / Public LB 0.8839 / Private LB 0.8825</p>\n<h2>Appendix: all networks</h2>\n<p><img src=\"https://pbs.twimg.com/media/FAf6m7SUUAM2Plk?format=jpg&amp;name=medium\" alt=\"\"><br>\n<img src=\"https://pbs.twimg.com/media/FAf6oBSVEAAFwAa?format=jpg&amp;name=medium\" alt=\"\"></p>",
  "messages": [
    {
      "id": "1528856",
      "postDate": "09/30/2021 02:16:05",
      "content": "<p>First of all, I would like to express deep gratitude to organizers and all the teams for making this competition so interesting and exciting. Also, I want to express a big congratulations to the first place, who dominated this competition with a single ResNet34 model :)</p>\n<p>Code is available at: <a href=\"https://github.com/analokmaus/kaggle-g2net-public\" target=\"_blank\">https://github.com/analokmaus/kaggle-g2net-public</a></p>\n<h1>Common settings</h1>\n<p>num epochs = 8<br>\noptimizer = Adam<br>\nscheduler = CosineAnnealingWarmRestarts(8)<br>\nloss function = BCE<br>\ncross validation = target stratified 5 fold cross validation</p>\n<h1>Frontend architectures</h1>\n<p>Neural network architecture played the most important role for improving the performance.<br>\nHere are the frontend architectures I used. <br>\nTrainable frontend in general outperformed fixed frontend.<br>\n<img src=\"https://pbs.twimg.com/media/FAfu6LJVEAA8u53?format=jpg&amp;name=medium\" alt=\"\"><br>\n<img src=\"https://pbs.twimg.com/media/FAfu7VvUYAYjT1y?format=jpg&amp;name=medium\" alt=\"\"><br>\n<img src=\"https://pbs.twimg.com/media/FAfu7_RUUAQM2x4?format=jpg&amp;name=medium\" alt=\"\"><br>\n<img src=\"https://pbs.twimg.com/media/FAfu8t2VgAgWiUX?format=jpg&amp;name=medium\" alt=\"\"></p>\n<h1>Preprocessing</h1>\n<p>I applied bandpass filter to all networks. [16, 512] for CWT-CNN and Trainable frontend CNN, [30, 300] for 1d-CNN. Whitening did not work.</p>\n<h1>Augmentations</h1>\n<p>I tested several types of augmentations on wave and spectrogram, and only a few of wave augmentations worked. I added gaussian noise for 2d-CNN networks, and flipped wave amplitude for 1d-CNN.</p>\n<h1>Pseudo-labeling</h1>\n<p>Re-training on soft(continuous) pseudo-labelled test dataset improved AUC by ~0.001. Label smoothing during pseudo-label also helped a bit. </p>\n<h1>Stacking</h1>\n<p>I kept oofs and predictions from all my experiments. <br>\nCross validated Ridge regression model was used to combine the outputs from neural network models. <br>\nA constant improvement in both CV and LB was observed as I add more model into the stacking model.<br>\nFinally, I run greedy model selection and chose 20(/10/5) models to maximize CV. </p>\n<p><strong>Stacking 20 models: CV 0.88283 / Public LB 0.8845 / Private LB 0.8829</strong><br>\nStacking 10 models: CV 0.88270 / Public LB 0.8842 / Private LB 0.8827<br>\nStacking 5 models: CV 0.88242 / Public LB 0.8839 / Private LB 0.8825</p>\n<h2>Appendix: all networks</h2>\n<p><img src=\"https://pbs.twimg.com/media/FAf6m7SUUAM2Plk?format=jpg&amp;name=medium\" alt=\"\"><br>\n<img src=\"https://pbs.twimg.com/media/FAf6oBSVEAAFwAa?format=jpg&amp;name=medium\" alt=\"\"></p>",
      "rawMarkdown": "First of all, I would like to express deep gratitude to organizers and all the teams for making this competition so interesting and exciting. Also, I want to express a big congratulations to the first place, who dominated this competition with a single ResNet34 model :)\n\nCode is available at: https://github.com/analokmaus/kaggle-g2net-public\n\n# Common settings\nnum epochs = 8\noptimizer = Adam\nscheduler = CosineAnnealingWarmRestarts(8)\nloss function = BCE\ncross validation = target stratified 5 fold cross validation\n\n# Frontend architectures\nNeural network architecture played the most important role for improving the performance.\nHere are the frontend architectures I used. \nTrainable frontend in general outperformed fixed frontend.\n<img src=\"https://pbs.twimg.com/media/FAfu6LJVEAA8u53?format=jpg&name=medium\" alt=\"\" width=\"500\"/>\n<img src=\"https://pbs.twimg.com/media/FAfu7VvUYAYjT1y?format=jpg&name=medium\" alt=\"\" width=\"500\"/>\n<img src=\"https://pbs.twimg.com/media/FAfu7_RUUAQM2x4?format=jpg&name=medium\" alt=\"\" width=\"500\"/>\n<img src=\"https://pbs.twimg.com/media/FAfu8t2VgAgWiUX?format=jpg&name=medium\" alt=\"\" width=\"500\"/>\n\n# Preprocessing\nI applied bandpass filter to all networks. [16, 512] for CWT-CNN and Trainable frontend CNN, [30, 300] for 1d-CNN. Whitening did not work.\n\n# Augmentations\nI tested several types of augmentations on wave and spectrogram, and only a few of wave augmentations worked. I added gaussian noise for 2d-CNN networks, and flipped wave amplitude for 1d-CNN.\n\n# Pseudo-labeling\nRe-training on soft(continuous) pseudo-labelled test dataset improved AUC by ~0.001. Label smoothing during pseudo-label also helped a bit. \n\n# Stacking\nI kept oofs and predictions from all my experiments. \nCross validated Ridge regression model was used to combine the outputs from neural network models. \nA constant improvement in both CV and LB was observed as I add more model into the stacking model.\nFinally, I run greedy model selection and chose 20(/10/5) models to maximize CV. \n\n**Stacking 20 models: CV 0.88283 / Public LB 0.8845 / Private LB 0.8829**\nStacking 10 models: CV 0.88270 / Public LB 0.8842 / Private LB 0.8827\nStacking 5 models: CV 0.88242 / Public LB 0.8839 / Private LB 0.8825\n\n## Appendix: all networks \n<img src=\"https://pbs.twimg.com/media/FAf6m7SUUAM2Plk?format=jpg&name=medium\" alt=\"\" width=600>\n<img src=\"https://pbs.twimg.com/media/FAf6oBSVEAAFwAa?format=jpg&name=medium\" alt=\"\" width=600>",
      "votes": null
    },
    {
      "id": "1528875",
      "postDate": "09/30/2021 02:34:40",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/analokamus\" target=\"_blank\">@analokamus</a> on 2nd place it's good to see wavenet working so well, some of my simple tests it's not so good</p>",
      "rawMarkdown": "Congrats @analokamus on 2nd place it's good to see wavenet working so well, some of my simple tests it's not so good",
      "votes": null
    },
    {
      "id": "1528885",
      "postDate": "09/30/2021 02:43:29",
      "content": "<p>congrats for being 2nd in ranking and thanks for the write-up. I have a question on the table of results:</p>\n<ul>\n<li>is that for a single fold or k-fold?</li>\n<li>does it include TTA, etc</li>\n<li>for example in model 11, it is frontend=CNN, backend=efficient b6. CNN refers to  1d CNN wavegram?</li>\n</ul>",
      "rawMarkdown": "congrats for being 2nd in ranking and thanks for the write-up. I have a question on the table of results:\n- is that for a single fold or k-fold?\n- does it include TTA, etc\n- for example in model 11, it is frontend=CNN, backend=efficient b6. CNN refers to  1d CNN wavegram?",
      "votes": null
    },
    {
      "id": "1528886",
      "postDate": "09/30/2021 02:44:17",
      "content": "<p>Original wavenet with tanh and sigmoid activation converged very slow, that's why I used a simplified version of it (in the figure: \"1d-CNN\").</p>",
      "rawMarkdown": "Original wavenet with tanh and sigmoid activation converged very slow, that's why I used a simplified version of it (in the figure: \"1d-CNN\").",
      "votes": null
    },
    {
      "id": "1528889",
      "postDate": "09/30/2021 02:46:42",
      "content": "<p>Thank you for the questions. I forgot to add some key information 😫</p>\n<ul>\n<li>is that for a single fold or k-fold?<br>\nYes, all results above are from 5-fold CV.</li>\n<li>does it include TTA, etc<br>\nYes, the same augmentation used in training was used in TTA.</li>\n<li>CNN refers to 1d CNN wavegram?<br>\nYes, CNN = 1d-CNN wavegram.</li>\n</ul>",
      "rawMarkdown": "Thank you for the questions. I forgot to add some key information 😫\n- is that for a single fold or k-fold?\nYes, all results above are from 5-fold CV.\n- does it include TTA, etc\nYes, the same augmentation used in training was used in TTA.\n- CNN refers to 1d CNN wavegram?\nYes, CNN = 1d-CNN wavegram.",
      "votes": null
    },
    {
      "id": "1529239",
      "postDate": "09/30/2021 08:41:27",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/analokamus\" target=\"_blank\">@analokamus</a> for a vigorous finish. And thank you for sharing your approach 🤘 </p>\n<p>On the pseudo labeling phase, which model did you use to label the test dataset? And did you add this data to your full pipeline later on together with the train data?</p>",
      "rawMarkdown": "Congrats @analokamus for a vigorous finish. And thank you for sharing your approach 🤘 \n\nOn the pseudo labeling phase, which model did you use to label the test dataset? And did you add this data to your full pipeline later on together with the train data?",
      "votes": null
    },
    {
      "id": "1529262",
      "postDate": "09/30/2021 08:57:42",
      "content": "<p>Congratulations on the solo gold! Really impressive that you managed to dig those big claws of yours into the top of the leaderboard and stay there for so long! 🐻</p>",
      "rawMarkdown": "Congratulations on the solo gold! Really impressive that you managed to dig those big claws of yours into the top of the leaderboard and stay there for so long! 🐻",
      "votes": null
    },
    {
      "id": "1529385",
      "postDate": "09/30/2021 10:53:04",
      "content": "<p>Congratulations and thanks for sharing! </p>",
      "rawMarkdown": "Congratulations and thanks for sharing!",
      "votes": null
    },
    {
      "id": "1529457",
      "postDate": "09/30/2021 12:22:38",
      "content": "<p>This is just beautiful.<br>\nCongratulations, well deserved!</p>",
      "rawMarkdown": "This is just beautiful.\nCongratulations, well deserved!",
      "votes": null
    },
    {
      "id": "1529459",
      "postDate": "09/30/2021 12:26:32",
      "content": "<p>In order to avoid information leak, the pseudo labels are from the same model. Pseudo labeled test set was added to the training set in each fold.</p>",
      "rawMarkdown": "In order to avoid information leak, the pseudo labels are from the same model. Pseudo labeled test set was added to the training set in each fold.",
      "votes": null
    },
    {
      "id": "1530305",
      "postDate": "10/01/2021 05:05:04",
      "content": "<p>Congrats on the solo gold!  Thanks for the nice write-up!</p>",
      "rawMarkdown": "Congrats on the solo gold!  Thanks for the nice write-up!",
      "votes": null
    },
    {
      "id": "1530319",
      "postDate": "10/01/2021 05:16:40",
      "content": "<p>Congratulations and thanks for sharing</p>",
      "rawMarkdown": "Congratulations and thanks for sharing",
      "votes": null
    },
    {
      "id": "1530350",
      "postDate": "10/01/2021 05:41:53",
      "content": "<p>Code is now available: <a href=\"https://github.com/analokmaus/kaggle-g2net-public\" target=\"_blank\">https://github.com/analokmaus/kaggle-g2net-public</a></p>",
      "rawMarkdown": "Code is now available: https://github.com/analokmaus/kaggle-g2net-public",
      "votes": null
    },
    {
      "id": "1534870",
      "postDate": "10/05/2021 09:48:31",
      "content": "<p>Congrats and thanks a lot for sharing the code!<br>\nCould you give a basic intuition about how the CNN and WaveNet frontends work? </p>",
      "rawMarkdown": "Congrats and thanks a lot for sharing the code!\nCould you give a basic intuition about how the CNN and WaveNet frontends work?",
      "votes": null
    },
    {
      "id": "1535166",
      "postDate": "10/05/2021 14:30:06",
      "content": "<p>While a spectrogram shows the intensity of a specific range of frequencies,  a CNN frontend can extract information from waveforms more dynamically.</p>",
      "rawMarkdown": "While a spectrogram shows the intensity of a specific range of frequencies,  a CNN frontend can extract information from waveforms more dynamically.",
      "votes": null
    },
    {
      "id": "1537815",
      "postDate": "10/07/2021 19:37:25",
      "content": "<p>Congratulations and thanks so much for sharing your code. I have a question about the <code>cwt</code> code. Is the combined options of <code>trainable_width=True</code> and <code>trainable_filter=False</code> working? I guess the variable <code>self.wavelet_width</code> is only used once at initialization step and wouldn't be included in each gradient update?</p>",
      "rawMarkdown": "Congratulations and thanks so much for sharing your code. I have a question about the `cwt` code. Is the combined options of `trainable_width=True` and `trainable_filter=False` working? I guess the variable `self.wavelet_width` is only used once at initialization step and wouldn't be included in each gradient update?",
      "votes": null
    },
    {
      "id": "1548689",
      "postDate": "10/18/2021 12:50:05",
      "content": "<p>Thanks for sharing. Btw, what is a 'trainable frontend'? I am a noob. </p>",
      "rawMarkdown": "Thanks for sharing. Btw, what is a 'trainable frontend'? I am a noob.",
      "votes": null
    },
    {
      "id": "1559927",
      "postDate": "10/27/2021 08:23:58",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1528875,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "09/30/2021 02:34:40",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/analokamus\" target=\"_blank\">@analokamus</a> on 2nd place it's good to see wavenet working so well, some of my simple tests it's not so good</p>",
      "votes": null,
      "replies": [
        {
          "id": 1528886,
          "author_name": "analokamus",
          "author_url": "",
          "post_date": "09/30/2021 02:44:17",
          "content": "<p>Original wavenet with tanh and sigmoid activation converged very slow, that's why I used a simplified version of it (in the figure: \"1d-CNN\").</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1528885,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/30/2021 02:43:29",
      "content": "<p>congrats for being 2nd in ranking and thanks for the write-up. I have a question on the table of results:</p>\n<ul>\n<li>is that for a single fold or k-fold?</li>\n<li>does it include TTA, etc</li>\n<li>for example in model 11, it is frontend=CNN, backend=efficient b6. CNN refers to  1d CNN wavegram?</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1528889,
          "author_name": "analokamus",
          "author_url": "",
          "post_date": "09/30/2021 02:46:42",
          "content": "<p>Thank you for the questions. I forgot to add some key information 😫</p>\n<ul>\n<li>is that for a single fold or k-fold?<br>\nYes, all results above are from 5-fold CV.</li>\n<li>does it include TTA, etc<br>\nYes, the same augmentation used in training was used in TTA.</li>\n<li>CNN refers to 1d CNN wavegram?<br>\nYes, CNN = 1d-CNN wavegram.</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1529239,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "09/30/2021 08:41:27",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/analokamus\" target=\"_blank\">@analokamus</a> for a vigorous finish. And thank you for sharing your approach 🤘 </p>\n<p>On the pseudo labeling phase, which model did you use to label the test dataset? And did you add this data to your full pipeline later on together with the train data?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1529459,
          "author_name": "analokamus",
          "author_url": "",
          "post_date": "09/30/2021 12:26:32",
          "content": "<p>In order to avoid information leak, the pseudo labels are from the same model. Pseudo labeled test set was added to the training set in each fold.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1529262,
      "author_name": "anjum48",
      "author_url": "",
      "post_date": "09/30/2021 08:57:42",
      "content": "<p>Congratulations on the solo gold! Really impressive that you managed to dig those big claws of yours into the top of the leaderboard and stay there for so long! 🐻</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1529385,
      "author_name": "danskas",
      "author_url": "",
      "post_date": "09/30/2021 10:53:04",
      "content": "<p>Congratulations and thanks for sharing! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1529457,
      "author_name": "rolandw0w",
      "author_url": "",
      "post_date": "09/30/2021 12:22:38",
      "content": "<p>This is just beautiful.<br>\nCongratulations, well deserved!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1530305,
      "author_name": "richx86",
      "author_url": "",
      "post_date": "10/01/2021 05:05:04",
      "content": "<p>Congrats on the solo gold!  Thanks for the nice write-up!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1530319,
      "author_name": "sidharkal",
      "author_url": "",
      "post_date": "10/01/2021 05:16:40",
      "content": "<p>Congratulations and thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1530350,
      "author_name": "analokamus",
      "author_url": "",
      "post_date": "10/01/2021 05:41:53",
      "content": "<p>Code is now available: <a href=\"https://github.com/analokmaus/kaggle-g2net-public\" target=\"_blank\">https://github.com/analokmaus/kaggle-g2net-public</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1537815,
          "author_name": "ggrizzly",
          "author_url": "",
          "post_date": "10/07/2021 19:37:25",
          "content": "<p>Congratulations and thanks so much for sharing your code. I have a question about the <code>cwt</code> code. Is the combined options of <code>trainable_width=True</code> and <code>trainable_filter=False</code> working? I guess the variable <code>self.wavelet_width</code> is only used once at initialization step and wouldn't be included in each gradient update?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1534870,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "10/05/2021 09:48:31",
      "content": "<p>Congrats and thanks a lot for sharing the code!<br>\nCould you give a basic intuition about how the CNN and WaveNet frontends work? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1535166,
          "author_name": "analokamus",
          "author_url": "",
          "post_date": "10/05/2021 14:30:06",
          "content": "<p>While a spectrogram shows the intensity of a specific range of frequencies,  a CNN frontend can extract information from waveforms more dynamically.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1548689,
      "author_name": "smoochy",
      "author_url": "",
      "post_date": "10/18/2021 12:50:05",
      "content": "<p>Thanks for sharing. Btw, what is a 'trainable frontend'? I am a noob. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1559927,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 08:23:58",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1528856": "First of all, I would like to express deep gratitude to organizers and all the teams for making this competition so interesting and exciting. Also, I want to express a big congratulations to the first place, who dominated this competition with a single ResNet34 model :)\n\nCode is available at: https://github.com/analokmaus/kaggle-g2net-public\n\n# Common settings\nnum epochs = 8\noptimizer = Adam\nscheduler = CosineAnnealingWarmRestarts(8)\nloss function = BCE\ncross validation = target stratified 5 fold cross validation\n\n# Frontend architectures\nNeural network architecture played the most important role for improving the performance.\nHere are the frontend architectures I used. \nTrainable frontend in general outperformed fixed frontend.\n<img src=\"https://pbs.twimg.com/media/FAfu6LJVEAA8u53?format=jpg&name=medium\" alt=\"\" width=\"500\"/>\n<img src=\"https://pbs.twimg.com/media/FAfu7VvUYAYjT1y?format=jpg&name=medium\" alt=\"\" width=\"500\"/>\n<img src=\"https://pbs.twimg.com/media/FAfu7_RUUAQM2x4?format=jpg&name=medium\" alt=\"\" width=\"500\"/>\n<img src=\"https://pbs.twimg.com/media/FAfu8t2VgAgWiUX?format=jpg&name=medium\" alt=\"\" width=\"500\"/>\n\n# Preprocessing\nI applied bandpass filter to all networks. [16, 512] for CWT-CNN and Trainable frontend CNN, [30, 300] for 1d-CNN. Whitening did not work.\n\n# Augmentations\nI tested several types of augmentations on wave and spectrogram, and only a few of wave augmentations worked. I added gaussian noise for 2d-CNN networks, and flipped wave amplitude for 1d-CNN.\n\n# Pseudo-labeling\nRe-training on soft(continuous) pseudo-labelled test dataset improved AUC by ~0.001. Label smoothing during pseudo-label also helped a bit. \n\n# Stacking\nI kept oofs and predictions from all my experiments. \nCross validated Ridge regression model was used to combine the outputs from neural network models. \nA constant improvement in both CV and LB was observed as I add more model into the stacking model.\nFinally, I run greedy model selection and chose 20(/10/5) models to maximize CV. \n\n**Stacking 20 models: CV 0.88283 / Public LB 0.8845 / Private LB 0.8829**\nStacking 10 models: CV 0.88270 / Public LB 0.8842 / Private LB 0.8827\nStacking 5 models: CV 0.88242 / Public LB 0.8839 / Private LB 0.8825\n\n## Appendix: all networks \n<img src=\"https://pbs.twimg.com/media/FAf6m7SUUAM2Plk?format=jpg&name=medium\" alt=\"\" width=600>\n<img src=\"https://pbs.twimg.com/media/FAf6oBSVEAAFwAa?format=jpg&name=medium\" alt=\"\" width=600>",
    "1528875": "Congrats @analokamus on 2nd place it's good to see wavenet working so well, some of my simple tests it's not so good",
    "1528885": "congrats for being 2nd in ranking and thanks for the write-up. I have a question on the table of results:\n- is that for a single fold or k-fold?\n- does it include TTA, etc\n- for example in model 11, it is frontend=CNN, backend=efficient b6. CNN refers to  1d CNN wavegram?",
    "1528886": "Original wavenet with tanh and sigmoid activation converged very slow, that's why I used a simplified version of it (in the figure: \"1d-CNN\").",
    "1528889": "Thank you for the questions. I forgot to add some key information 😫\n- is that for a single fold or k-fold?\nYes, all results above are from 5-fold CV.\n- does it include TTA, etc\nYes, the same augmentation used in training was used in TTA.\n- CNN refers to 1d CNN wavegram?\nYes, CNN = 1d-CNN wavegram.",
    "1529239": "Congrats @analokamus for a vigorous finish. And thank you for sharing your approach 🤘 \n\nOn the pseudo labeling phase, which model did you use to label the test dataset? And did you add this data to your full pipeline later on together with the train data?",
    "1529262": "Congratulations on the solo gold! Really impressive that you managed to dig those big claws of yours into the top of the leaderboard and stay there for so long! 🐻",
    "1529385": "Congratulations and thanks for sharing!",
    "1529457": "This is just beautiful.\nCongratulations, well deserved!",
    "1529459": "In order to avoid information leak, the pseudo labels are from the same model. Pseudo labeled test set was added to the training set in each fold.",
    "1530305": "Congrats on the solo gold!  Thanks for the nice write-up!",
    "1530319": "Congratulations and thanks for sharing",
    "1530350": "Code is now available: https://github.com/analokmaus/kaggle-g2net-public",
    "1534870": "Congrats and thanks a lot for sharing the code!\nCould you give a basic intuition about how the CNN and WaveNet frontends work?",
    "1535166": "While a spectrogram shows the intensity of a specific range of frequencies,  a CNN frontend can extract information from waveforms more dynamically.",
    "1537815": "Congratulations and thanks so much for sharing your code. I have a question about the `cwt` code. Is the combined options of `trainable_width=True` and `trainable_filter=False` working? I guess the variable `self.wavelet_width` is only used once at initialization step and wouldn't be included in each gradient update?",
    "1548689": "Thanks for sharing. Btw, what is a 'trainable frontend'? I am a noob.",
    "1559927": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}