{
  "id": 266711,
  "title": "12th place solution - experiments with augmentation",
  "url": "/competitions/seti-breakthrough-listen/discussion/266711",
  "author_name": "Mohsin hasan",
  "post_date": "2021-08-20T05:18:28.128000",
  "votes": 26,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Thanks to hosts, all competitors and Kaggle for once again providing a great learning opportunity.</p>\n<h3>Models</h3>\n<p>For most experiments I used <code>efficientnet_b0</code>. During final days, I added <code>efficientnetv2s</code> and <code>efficientnetb1</code>, <code>efficientnetb2</code> to ensemble, which is a geometric mean of different runs.<br>\n(I missed the trick here by not adding bigger models).</p>\n<h3>Input</h3>\n<p>For most of my experiments I used all 6 inputs stacked by time and resized to either 512x512 or 768x512.<br>\nFor most most experiments, original stacked image is used as channel 0, frequency normalized image is used as channel 1, and 3rd channel is 0 and 1 basis whether that part of input is ON target or OFF target.</p>\n<h3>Training setup</h3>\n<p>I extensively use config files and <a href=\"https://hydra.cc/\" target=\"_blank\">hydra</a> to manage experiments. <a href=\"https://wandb.ai/site\" target=\"_blank\">Weights and biases</a> for logging and monitoring. <a href=\"https://github.com/PyTorchLightning/pytorch-lightning\" target=\"_blank\">Pytorch lightning</a> for boiler plate training code. Most of the code is <strong>flake8</strong> linted and auto formatted by <a href=\"https://github.com/psf/black\" target=\"_blank\">black</a>. Great thanks to <code>https://github.com/ashleve/lightning-hydra-template</code> for providing a good starter code base. Thanks to all the developers of those libraries!</p>\n<p>All my codes are available here: <a href=\"https://github.com/mohsinkhn/seti-kaggle-tezdhar\" target=\"_blank\">https://github.com/mohsinkhn/seti-kaggle-tezdhar</a></p>\n<h3>Experiments:</h3>\n<ul>\n<li><p>Including off channels gives small boost on CV but a bigger one (0.1-.2) on LB for some of my early models. It could be explained by narrowband drift signals existing in both on-off channels (more so in test data).</p></li>\n<li><p>Untied mixup ( <a href=\"https://openreview.net/pdf?id=SkgjKR4YwH\" target=\"_blank\">https://openreview.net/pdf?id=SkgjKR4YwH</a> ). The basic idea is to have different distributions for mixing of inputs and targets. For inputs I use beta distribution with alpha=1.0, beta=1.0 while for targets I use beta distribution with alpha=1.0, beta=0.5.</p></li>\n<li><p>Augmentations: To detect unknown classes, I wanted to enforce the condition that signal should be present in only ON channels and NOT OFF inputs. To this end, I roll data in inputs by 1 step ( [A0, B, A1, C, A2, D] --&gt; [D, A0, B, A1, C, A2]) and make target 0. The image in input 0 becomes image in input 1, image in input 1 becomes image in input 2 and so on. <strong>Doing so, along with SpecAug, I am able to successfully predict new class in test data</strong></p></li>\n<li><p>Other than above augmentation, I used hflip, roll, specAug, meanshift, verticalshift(2 steps as 6 input image can't be vertically flipped)</p></li>\n<li><p>I tried many normalisation techniques, mean absolute deviation scaling, sigmoid scaling, random power after min-max scaling etc. but none of them gave significant boost.</p></li>\n<li><p>Tried pseudo labelling towards the end with only confident predictions, but couldn't get much benefit.</p></li>\n</ul>\n<p>Although, feels bad to loose solo gold by 1 rank, it was really good learning experience for me, especially my experimentation stack has improved significantly. .</p>",
  "messages": [
    {
      "id": 1482454,
      "postDate": "2021-08-20T05:18:28.130Z",
      "content": "<p>Thanks to hosts, all competitors and Kaggle for once again providing a great learning opportunity.</p>\n<h3>Models</h3>\n<p>For most experiments I used <code>efficientnet_b0</code>. During final days, I added <code>efficientnetv2s</code> and <code>efficientnetb1</code>, <code>efficientnetb2</code> to ensemble, which is a geometric mean of different runs.<br>\n(I missed the trick here by not adding bigger models).</p>\n<h3>Input</h3>\n<p>For most of my experiments I used all 6 inputs stacked by time and resized to either 512x512 or 768x512.<br>\nFor most most experiments, original stacked image is used as channel 0, frequency normalized image is used as channel 1, and 3rd channel is 0 and 1 basis whether that part of input is ON target or OFF target.</p>\n<h3>Training setup</h3>\n<p>I extensively use config files and <a href=\"https://hydra.cc/\" target=\"_blank\">hydra</a> to manage experiments. <a href=\"https://wandb.ai/site\" target=\"_blank\">Weights and biases</a> for logging and monitoring. <a href=\"https://github.com/PyTorchLightning/pytorch-lightning\" target=\"_blank\">Pytorch lightning</a> for boiler plate training code. Most of the code is <strong>flake8</strong> linted and auto formatted by <a href=\"https://github.com/psf/black\" target=\"_blank\">black</a>. Great thanks to <code>https://github.com/ashleve/lightning-hydra-template</code> for providing a good starter code base. Thanks to all the developers of those libraries!</p>\n<p>All my codes are available here: <a href=\"https://github.com/mohsinkhn/seti-kaggle-tezdhar\" target=\"_blank\">https://github.com/mohsinkhn/seti-kaggle-tezdhar</a></p>\n<h3>Experiments:</h3>\n<ul>\n<li><p>Including off channels gives small boost on CV but a bigger one (0.1-.2) on LB for some of my early models. It could be explained by narrowband drift signals existing in both on-off channels (more so in test data).</p></li>\n<li><p>Untied mixup ( <a href=\"https://openreview.net/pdf?id=SkgjKR4YwH\" target=\"_blank\">https://openreview.net/pdf?id=SkgjKR4YwH</a> ). The basic idea is to have different distributions for mixing of inputs and targets. For inputs I use beta distribution with alpha=1.0, beta=1.0 while for targets I use beta distribution with alpha=1.0, beta=0.5.</p></li>\n<li><p>Augmentations: To detect unknown classes, I wanted to enforce the condition that signal should be present in only ON channels and NOT OFF inputs. To this end, I roll data in inputs by 1 step ( [A0, B, A1, C, A2, D] --&gt; [D, A0, B, A1, C, A2]) and make target 0. The image in input 0 becomes image in input 1, image in input 1 becomes image in input 2 and so on. <strong>Doing so, along with SpecAug, I am able to successfully predict new class in test data</strong></p></li>\n<li><p>Other than above augmentation, I used hflip, roll, specAug, meanshift, verticalshift(2 steps as 6 input image can't be vertically flipped)</p></li>\n<li><p>I tried many normalisation techniques, mean absolute deviation scaling, sigmoid scaling, random power after min-max scaling etc. but none of them gave significant boost.</p></li>\n<li><p>Tried pseudo labelling towards the end with only confident predictions, but couldn't get much benefit.</p></li>\n</ul>\n<p>Although, feels bad to loose solo gold by 1 rank, it was really good learning experience for me, especially my experimentation stack has improved significantly. .</p>",
      "rawMarkdown": "Thanks to hosts, all competitors and Kaggle for once again providing a great learning opportunity.\n\n### Models\nFor most experiments I used `efficientnet_b0`. During final days, I added `efficientnetv2s` and `efficientnetb1`, `efficientnetb2` to ensemble, which is a geometric mean of different runs.\n(I missed the trick here by not adding bigger models).\n\n### Input\nFor most of my experiments I used all 6 inputs stacked by time and resized to either 512x512 or 768x512.\nFor most most experiments, original stacked image is used as channel 0, frequency normalized image is used as channel 1, and 3rd channel is 0 and 1 basis whether that part of input is ON target or OFF target.\n\n### Training setup\nI extensively use config files and [hydra](https://hydra.cc/) to manage experiments. [Weights and biases](https://wandb.ai/site) for logging and monitoring. [Pytorch lightning](https://github.com/PyTorchLightning/pytorch-lightning) for boiler plate training code. Most of the code is **flake8** linted and auto formatted by [black](https://github.com/psf/black). Great thanks to `https://github.com/ashleve/lightning-hydra-template` for providing a good starter code base. Thanks to all the developers of those libraries!\n\nAll my codes are available here: https://github.com/mohsinkhn/seti-kaggle-tezdhar\n\n\n### Experiments:\n* Including off channels gives small boost on CV but a bigger one (0.1-.2) on LB for some of my early models. It could be explained by narrowband drift signals existing in both on-off channels (more so in test data).\n\n* Untied mixup ( https://openreview.net/pdf?id=SkgjKR4YwH ). The basic idea is to have different distributions for mixing of inputs and targets. For inputs I use beta distribution with alpha=1.0, beta=1.0 while for targets I use beta distribution with alpha=1.0, beta=0.5.\n\n* Augmentations: To detect unknown classes, I wanted to enforce the condition that signal should be present in only ON channels and NOT OFF inputs. To this end, I roll data in inputs by 1 step ( [A0, B, A1, C, A2, D] --> [D, A0, B, A1, C, A2]) and make target 0. The image in input 0 becomes image in input 1, image in input 1 becomes image in input 2 and so on. **Doing so, along with SpecAug, I am able to successfully predict new class in test data**\n\n* Other than above augmentation, I used hflip, roll, specAug, meanshift, verticalshift(2 steps as 6 input image can't be vertically flipped)\n\n* I tried many normalisation techniques, mean absolute deviation scaling, sigmoid scaling, random power after min-max scaling etc. but none of them gave significant boost.\n\n* Tried pseudo labelling towards the end with only confident predictions, but couldn't get much benefit.\n\n\nAlthough, feels bad to loose solo gold by 1 rank, it was really good learning experience for me, especially my experimentation stack has improved significantly. .\n\n\n\n\n",
      "votes": 26
    },
    {
      "id": 1482696,
      "postDate": "2021-08-20T07:47:02.013Z",
      "content": "<blockquote>\n  <p>Augmentations: To detect unknown classes, I wanted to enforce the condition that signal should be present in only ON channels and NOT OFF inputs. To this end, I roll data in inputs by 1 step and make target 0. The image in input 0 becomes image in input 1, image in input 1 becomes image in input 2 and so on. Doing so, along with SpecAug, I am able to successfully predict new class in test data</p>\n</blockquote>\n<p>OK I love this. It's like the trick I used of adding an indicator channel <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266460\" target=\"_blank\">(link)</a>, but better, because it FORCES the network to be aware of the location of the needle in terms of on/off, rather than just overfitting to its shape. Damn, wish I'd thought of this!</p>",
      "rawMarkdown": "> Augmentations: To detect unknown classes, I wanted to enforce the condition that signal should be present in only ON channels and NOT OFF inputs. To this end, I roll data in inputs by 1 step and make target 0. The image in input 0 becomes image in input 1, image in input 1 becomes image in input 2 and so on. Doing so, along with SpecAug, I am able to successfully predict new class in test data\n\nOK I love this. It's like the trick I used of adding an indicator channel [(link)](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266460), but better, because it FORCES the network to be aware of the location of the needle in terms of on/off, rather than just overfitting to its shape. Damn, wish I'd thought of this!",
      "votes": 3,
      "replies": [
        {
          "id": 1482706,
          "postDate": "2021-08-20T08:00:10.307Z",
          "content": "<p>It's a great idea, and I feel it's like the idea in <a href=\"https://arxiv.org/abs/1807.03247\" target=\"_blank\">CoordConv</a>, to let Conv network to learn some location information instead of only shapes and patterns. <br>\n<a href=\"https://www.kaggle.com/tezdhar\" target=\"_blank\">@tezdhar</a> , may I ask when you add the indicatior channel, do you have to indicate [A0, B, A1, C, A2, D] as [1, 0, 1, 0, 1, 0], or indicate as [0, 1, 0, 1, 0, 1] is same for the Conv network?</p>",
          "rawMarkdown": "It's a great idea, and I feel it's like the idea in [CoordConv](https://arxiv.org/abs/1807.03247), to let Conv network to learn some location information instead of only shapes and patterns. \n@tezdhar , may I ask when you add the indicatior channel, do you have to indicate [A0, B, A1, C, A2, D] as [1, 0, 1, 0, 1, 0], or indicate as [0, 1, 0, 1, 0, 1] is same for the Conv network?",
          "votes": 1
        },
        {
          "id": 1482785,
          "postDate": "2021-08-20T09:21:09.307Z",
          "content": "<p><a href=\"https://www.kaggle.com/superchenhao\" target=\"_blank\">@superchenhao</a> I add it as [1, 0, 1, 0, 1, 0], but IMO it shouldn't really matter</p>",
          "rawMarkdown": "@superchenhao I add it as [1, 0, 1, 0, 1, 0], but IMO it shouldn't really matter",
          "votes": 1
        },
        {
          "id": 1482909,
          "postDate": "2021-08-20T10:49:16.860Z",
          "content": "<p>Congratulations Mohsin! Can you explain <code>roll data in inputs by 1 step</code>? I don't understand what you are doing.</p>",
          "rawMarkdown": "Congratulations Mohsin! Can you explain `roll data in inputs by 1 step`? I don't understand what you are doing."
        },
        {
          "id": 1482968,
          "postDate": "2021-08-20T11:21:52.690Z",
          "content": "<p>Taking notation from <a href=\"https://www.kaggle.com/superchenhao\" target=\"_blank\">@superchenhao</a>,  [A0, B, A1, C, A2, D] --&gt; [D, A0, B, A1, C, A2]</p>",
          "rawMarkdown": "Taking notation from @superchenhao,  [A0, B, A1, C, A2, D] --> [D, A0, B, A1, C, A2]",
          "votes": 1
        },
        {
          "id": 1482980,
          "postDate": "2021-08-20T11:31:21.567Z",
          "content": "<p>That's very creative. Does that improve CV LB? </p>\n<p>I did a similar thing that hurt me. I added \"off\" cadence images, i.e. <code>np.vstack( img[1::2] )</code> with <code>target=0</code> as additional train images.(My other train images are <code>np.vstack( img[::2] )</code>). </p>\n<p>This increased CV +0.003 but decreased LB -0.010. I think the host may have added artificial spurious correlation between signal and certain backgrounds, then added different spurious correlations to test data.</p>",
          "rawMarkdown": "That's very creative. Does that improve CV LB? \n\nI did a similar thing that hurt me. I added \"off\" cadence images, i.e. `np.vstack( img[1::2] )` with `target=0` as additional train images.(My other train images are `np.vstack( img[::2] )`). \n\nThis increased CV +0.003 but decreased LB -0.010. I think the host may have added artificial spurious correlation between signal and certain backgrounds, then added different spurious correlations to test data.",
          "votes": 1
        },
        {
          "id": 1482991,
          "postDate": "2021-08-20T11:39:23.470Z",
          "content": "<p>Yes, It improved for me! I guess by not including ON and OFF channels together, it would learn to identify wave like signal stretching from ON to OFF target cadences as non-significant and hurt new class prediction on test set.</p>",
          "rawMarkdown": "Yes, It improved for me! I guess by not including ON and OFF channels together, it would learn to identify wave like signal stretching from ON to OFF target cadences as non-significant and hurt new class prediction on test set.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1489800,
      "postDate": "2021-08-25T09:24:39.820Z",
      "content": "<p>Awesome work, thanks for sharing, you will get Gold medal shortly in your next competition, all the best!</p>",
      "rawMarkdown": "Awesome work, thanks for sharing, you will get Gold medal shortly in your next competition, all the best!",
      "votes": 1
    },
    {
      "id": 1488211,
      "postDate": "2021-08-24T06:55:49.663Z",
      "content": "<p>Thanks for sharing the approach. This will help us a lot in future cases…Eventually, you will get Gold Keep going buddy. <a href=\"https://www.kaggle.com/tezdhar\" target=\"_blank\">@tezdhar</a> </p>",
      "rawMarkdown": "Thanks for sharing the approach. This will help us a lot in future cases...Eventually, you will get Gold Keep going buddy. @tezdhar ",
      "votes": 1
    },
    {
      "id": 1483536,
      "postDate": "2021-08-20T17:34:05.903Z",
      "content": "<p>thx for sharing. new stuff learned.</p>",
      "rawMarkdown": "thx for sharing. new stuff learned.",
      "votes": 1
    },
    {
      "id": 1482583,
      "postDate": "2021-08-20T06:51:25.867Z",
      "content": "<p>Thanks, I also tried to do different distributions with mixup, didnt go so well </p>\n<ol>\n<li>(  there is closing bracket in url ). </li>\n<li>If you back ON/OFF layers as 3ch input you should be able to do any flips? </li>\n<li>And about roll-aug: so AAA cadences become BCD, and even thou there is a signal in them, sample becomes non-anomaly. Am i right? Did it work? </li>\n</ol>",
      "rawMarkdown": "Thanks, I also tried to do different distributions with mixup, didnt go so well \n1. (~~that mixup link is broken can you please fix that?~~  there is closing bracket in url ). \n2. If you back ON/OFF layers as 3ch input you should be able to do any flips? \n3. And about roll-aug: so AAA cadences become BCD, and even thou there is a signal in them, sample becomes non-anomaly. Am i right? Did it work? ",
      "votes": 1,
      "replies": [
        {
          "id": 1482607,
          "postDate": "2021-08-20T07:00:11.277Z",
          "content": "<ol>\n<li>Thanks pointing it broken url, fixed now. </li>\n<li>ON/OFF layers were temporally stacked and since I did add indicator channel (0 and 1 to differentiate ON/OFF), I could have used vertical flip I guess.</li>\n<li>Yes, improves both CV and LB</li>\n</ol>",
          "rawMarkdown": "1. Thanks pointing it broken url, fixed now. \n2. ON/OFF layers were temporally stacked and since I did add indicator channel (0 and 1 to differentiate ON/OFF), I could have used vertical flip I guess.\n3. Yes, improves both CV and LB",
          "votes": 1
        },
        {
          "id": 1482682,
          "postDate": "2021-08-20T07:35:07.737Z",
          "content": "<p>Implementation for untied mixup is quite tricky, I had multiple bugs before I finally got right!<br>\n<a href=\"https://github.com/mohsinkhn/seti-kaggle-tezdhar/blob/dd3b1dcb8729ed73c5bb6ce0041cd385c412156e/src/models/plmodels.py#L108\" target=\"_blank\">https://github.com/mohsinkhn/seti-kaggle-tezdhar/blob/dd3b1dcb8729ed73c5bb6ce0041cd385c412156e/src/models/plmodels.py#L108</a></p>\n<pre><code>    def _mixup_data(self, x, t, use_mixup):\n        \"\"\"Returns mixed inputs, pairs of targets, and lambda\"\"\"\n        if not use_mixup:\n            return x, t\n\n        alpha = self.hparams[\"mixup_alpha\"]\n        beta = self.hparams[\"mixup_beta\"]\n        ualpha = self.hparams[\"mixup_ualpha\"]\n        ubeta = self.hparams[\"mixup_ubeta\"]\n\n        untied = self.hparams[\"mixup_untied\"]\n\n        if alpha &gt; 0:\n            dist1 = betad(alpha, beta)\n            dist2 = betad(ualpha, ubeta)\n            dist3 = betad(ubeta, ualpha)\n            rng = np.random.rand()\n            lam1 = dist1.ppf(rng)\n            if untied:\n                lam2 = dist2.ppf(rng)\n                lam3 = dist3.ppf(rng)\n            else:\n                lam2 = lam1\n        else:\n            lam1 = 1\n            lam2 = 1\n\n        batch_size = x.size()[0]\n        index = torch.randperm(batch_size).to(self.device)\n        mixed_x = lam1 * x + (1 - lam1) * x[index, :]\n        if untied:\n            mixed_y = torch.clamp((lam2 * t) + (1 - lam3) * t[index], 0, 1)  # 0.3, 1 --&gt; 0.5, 1; 0.7, 0 --&gt; 0.5, 1\n        else:\n            mixed_y = lam2 * t + (1 - lam2) * t[index]\n        return mixed_x, mixed_y\n</code></pre>",
          "rawMarkdown": "Implementation for untied mixup is quite tricky, I had multiple bugs before I finally got right!\nhttps://github.com/mohsinkhn/seti-kaggle-tezdhar/blob/dd3b1dcb8729ed73c5bb6ce0041cd385c412156e/src/models/plmodels.py#L108\n\n```\n    def _mixup_data(self, x, t, use_mixup):\n        \"\"\"Returns mixed inputs, pairs of targets, and lambda\"\"\"\n        if not use_mixup:\n            return x, t\n\n        alpha = self.hparams[\"mixup_alpha\"]\n        beta = self.hparams[\"mixup_beta\"]\n        ualpha = self.hparams[\"mixup_ualpha\"]\n        ubeta = self.hparams[\"mixup_ubeta\"]\n\n        untied = self.hparams[\"mixup_untied\"]\n\n        if alpha > 0:\n            dist1 = betad(alpha, beta)\n            dist2 = betad(ualpha, ubeta)\n            dist3 = betad(ubeta, ualpha)\n            rng = np.random.rand()\n            lam1 = dist1.ppf(rng)\n            if untied:\n                lam2 = dist2.ppf(rng)\n                lam3 = dist3.ppf(rng)\n            else:\n                lam2 = lam1\n        else:\n            lam1 = 1\n            lam2 = 1\n\n        batch_size = x.size()[0]\n        index = torch.randperm(batch_size).to(self.device)\n        mixed_x = lam1 * x + (1 - lam1) * x[index, :]\n        if untied:\n            mixed_y = torch.clamp((lam2 * t) + (1 - lam3) * t[index], 0, 1)  # 0.3, 1 --> 0.5, 1; 0.7, 0 --> 0.5, 1\n        else:\n            mixed_y = lam2 * t + (1 - lam2) * t[index]\n        return mixed_x, mixed_y\n```",
          "votes": 1
        }
      ]
    },
    {
      "id": 1482652,
      "postDate": "2021-08-20T07:21:57.413Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/tezdhar\" target=\"_blank\">@tezdhar</a> , especially with the code. You are always impressive to climb up the leaderboard so fast at late stage of competition, I beleive you will get another solo gold and become Grandmaster soon. 🎉</p>\n<p>Regarding your roll-aug, if I understand correctly, you roll the positive sample like [A0, B, A1, C, A2, D] to [B, A1, C, A2, D, A0], then this sample's target will become 0, right? If so there will be fewer postive training sample, make the dataset even more imblance, in my understanding this would hurt the performance, could you please explain a bit here? Thanks.</p>",
      "rawMarkdown": "Thanks for sharing @tezdhar , especially with the code. You are always impressive to climb up the leaderboard so fast at late stage of competition, I beleive you will get another solo gold and become Grandmaster soon. 🎉\n\nRegarding your roll-aug, if I understand correctly, you roll the positive sample like [A0, B, A1, C, A2, D] to [B, A1, C, A2, D, A0], then this sample's target will become 0, right? If so there will be fewer postive training sample, make the dataset even more imblance, in my understanding this would hurt the performance, could you please explain a bit here? Thanks.",
      "votes": 2,
      "replies": [
        {
          "id": 1482688,
          "postDate": "2021-08-20T07:38:22.573Z",
          "content": "<p>Thank you for well wishes! </p>\n<p>Yes, your understanding is correct. The probability of this happening was kept pretty low at 10-20%, which meant only 1-2% positive samples were impacted. <br>\nI think my it gave small improvement on LB 0.786--&gt;0.789. But looking at probabilities of new class in test set, I could see them move from 0.7-0.8 to 0.9ish</p>",
          "rawMarkdown": "Thank you for well wishes! \n\nYes, your understanding is correct. The probability of this happening was kept pretty low at 10-20%, which meant only 1-2% positive samples were impacted. \nI think my it gave small improvement on LB 0.786-->0.789. But looking at probabilities of new class in test set, I could see them move from 0.7-0.8 to 0.9ish",
          "votes": 2
        }
      ]
    },
    {
      "id": 1495848,
      "postDate": "2021-08-29T20:41:43.650Z",
      "content": "<p>Good job =))</p>",
      "rawMarkdown": "Good job =))"
    },
    {
      "id": 1482529,
      "postDate": "2021-08-20T06:10:43.197Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1482696,
      "author_name": "James Howard",
      "author_url": "",
      "post_date": "2021-08-20T07:47:02.013000",
      "content": "<blockquote>\n  <p>Augmentations: To detect unknown classes, I wanted to enforce the condition that signal should be present in only ON channels and NOT OFF inputs. To this end, I roll data in inputs by 1 step and make target 0. The image in input 0 becomes image in input 1, image in input 1 becomes image in input 2 and so on. Doing so, along with SpecAug, I am able to successfully predict new class in test data</p>\n</blockquote>\n<p>OK I love this. It's like the trick I used of adding an indicator channel <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266460\" target=\"_blank\">(link)</a>, but better, because it FORCES the network to be aware of the location of the needle in terms of on/off, rather than just overfitting to its shape. Damn, wish I'd thought of this!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1482706,
          "author_name": "Hao",
          "author_url": "",
          "post_date": "2021-08-20T08:00:10.307000",
          "content": "<p>It's a great idea, and I feel it's like the idea in <a href=\"https://arxiv.org/abs/1807.03247\" target=\"_blank\">CoordConv</a>, to let Conv network to learn some location information instead of only shapes and patterns. <br>\n<a href=\"https://www.kaggle.com/tezdhar\" target=\"_blank\">@tezdhar</a> , may I ask when you add the indicatior channel, do you have to indicate [A0, B, A1, C, A2, D] as [1, 0, 1, 0, 1, 0], or indicate as [0, 1, 0, 1, 0, 1] is same for the Conv network?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1482785,
          "author_name": "Mohsin hasan",
          "author_url": "",
          "post_date": "2021-08-20T09:21:09.307000",
          "content": "<p><a href=\"https://www.kaggle.com/superchenhao\" target=\"_blank\">@superchenhao</a> I add it as [1, 0, 1, 0, 1, 0], but IMO it shouldn't really matter</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1482909,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-20T10:49:16.860000",
          "content": "<p>Congratulations Mohsin! Can you explain <code>roll data in inputs by 1 step</code>? I don't understand what you are doing.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1482968,
          "author_name": "Mohsin hasan",
          "author_url": "",
          "post_date": "2021-08-20T11:21:52.690000",
          "content": "<p>Taking notation from <a href=\"https://www.kaggle.com/superchenhao\" target=\"_blank\">@superchenhao</a>,  [A0, B, A1, C, A2, D] --&gt; [D, A0, B, A1, C, A2]</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1482980,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-20T11:31:21.567000",
          "content": "<p>That's very creative. Does that improve CV LB? </p>\n<p>I did a similar thing that hurt me. I added \"off\" cadence images, i.e. <code>np.vstack( img[1::2] )</code> with <code>target=0</code> as additional train images.(My other train images are <code>np.vstack( img[::2] )</code>). </p>\n<p>This increased CV +0.003 but decreased LB -0.010. I think the host may have added artificial spurious correlation between signal and certain backgrounds, then added different spurious correlations to test data.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1482991,
          "author_name": "Mohsin hasan",
          "author_url": "",
          "post_date": "2021-08-20T11:39:23.470000",
          "content": "<p>Yes, It improved for me! I guess by not including ON and OFF channels together, it would learn to identify wave like signal stretching from ON to OFF target cadences as non-significant and hurt new class prediction on test set.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1489800,
      "author_name": "Old Monk",
      "author_url": "",
      "post_date": "2021-08-25T09:24:39.820000",
      "content": "<p>Awesome work, thanks for sharing, you will get Gold medal shortly in your next competition, all the best!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1488211,
      "author_name": "Saurav Solanki",
      "author_url": "",
      "post_date": "2021-08-24T06:55:49.663000",
      "content": "<p>Thanks for sharing the approach. This will help us a lot in future cases…Eventually, you will get Gold Keep going buddy. <a href=\"https://www.kaggle.com/tezdhar\" target=\"_blank\">@tezdhar</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1483536,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2021-08-20T17:34:05.903000",
      "content": "<p>thx for sharing. new stuff learned.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1482583,
      "author_name": "Gleb",
      "author_url": "",
      "post_date": "2021-08-20T06:51:25.867000",
      "content": "<p>Thanks, I also tried to do different distributions with mixup, didnt go so well </p>\n<ol>\n<li>(  there is closing bracket in url ). </li>\n<li>If you back ON/OFF layers as 3ch input you should be able to do any flips? </li>\n<li>And about roll-aug: so AAA cadences become BCD, and even thou there is a signal in them, sample becomes non-anomaly. Am i right? Did it work? </li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 1482607,
          "author_name": "Mohsin hasan",
          "author_url": "",
          "post_date": "2021-08-20T07:00:11.277000",
          "content": "<ol>\n<li>Thanks pointing it broken url, fixed now. </li>\n<li>ON/OFF layers were temporally stacked and since I did add indicator channel (0 and 1 to differentiate ON/OFF), I could have used vertical flip I guess.</li>\n<li>Yes, improves both CV and LB</li>\n</ol>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1482682,
          "author_name": "Mohsin hasan",
          "author_url": "",
          "post_date": "2021-08-20T07:35:07.737000",
          "content": "<p>Implementation for untied mixup is quite tricky, I had multiple bugs before I finally got right!<br>\n<a href=\"https://github.com/mohsinkhn/seti-kaggle-tezdhar/blob/dd3b1dcb8729ed73c5bb6ce0041cd385c412156e/src/models/plmodels.py#L108\" target=\"_blank\">https://github.com/mohsinkhn/seti-kaggle-tezdhar/blob/dd3b1dcb8729ed73c5bb6ce0041cd385c412156e/src/models/plmodels.py#L108</a></p>\n<pre><code>    def _mixup_data(self, x, t, use_mixup):\n        \"\"\"Returns mixed inputs, pairs of targets, and lambda\"\"\"\n        if not use_mixup:\n            return x, t\n\n        alpha = self.hparams[\"mixup_alpha\"]\n        beta = self.hparams[\"mixup_beta\"]\n        ualpha = self.hparams[\"mixup_ualpha\"]\n        ubeta = self.hparams[\"mixup_ubeta\"]\n\n        untied = self.hparams[\"mixup_untied\"]\n\n        if alpha &gt; 0:\n            dist1 = betad(alpha, beta)\n            dist2 = betad(ualpha, ubeta)\n            dist3 = betad(ubeta, ualpha)\n            rng = np.random.rand()\n            lam1 = dist1.ppf(rng)\n            if untied:\n                lam2 = dist2.ppf(rng)\n                lam3 = dist3.ppf(rng)\n            else:\n                lam2 = lam1\n        else:\n            lam1 = 1\n            lam2 = 1\n\n        batch_size = x.size()[0]\n        index = torch.randperm(batch_size).to(self.device)\n        mixed_x = lam1 * x + (1 - lam1) * x[index, :]\n        if untied:\n            mixed_y = torch.clamp((lam2 * t) + (1 - lam3) * t[index], 0, 1)  # 0.3, 1 --&gt; 0.5, 1; 0.7, 0 --&gt; 0.5, 1\n        else:\n            mixed_y = lam2 * t + (1 - lam2) * t[index]\n        return mixed_x, mixed_y\n</code></pre>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1482652,
      "author_name": "Hao",
      "author_url": "",
      "post_date": "2021-08-20T07:21:57.413000",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/tezdhar\" target=\"_blank\">@tezdhar</a> , especially with the code. You are always impressive to climb up the leaderboard so fast at late stage of competition, I beleive you will get another solo gold and become Grandmaster soon. 🎉</p>\n<p>Regarding your roll-aug, if I understand correctly, you roll the positive sample like [A0, B, A1, C, A2, D] to [B, A1, C, A2, D, A0], then this sample's target will become 0, right? If so there will be fewer postive training sample, make the dataset even more imblance, in my understanding this would hurt the performance, could you please explain a bit here? Thanks.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1482688,
          "author_name": "Mohsin hasan",
          "author_url": "",
          "post_date": "2021-08-20T07:38:22.573000",
          "content": "<p>Thank you for well wishes! </p>\n<p>Yes, your understanding is correct. The probability of this happening was kept pretty low at 10-20%, which meant only 1-2% positive samples were impacted. <br>\nI think my it gave small improvement on LB 0.786--&gt;0.789. But looking at probabilities of new class in test set, I could see them move from 0.7-0.8 to 0.9ish</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1495848,
      "author_name": "Fuco",
      "author_url": "",
      "post_date": "2021-08-29T20:41:43.650000",
      "content": "<p>Good job =))</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1482529,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-20T06:10:43.197000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1482454": "Thanks to hosts, all competitors and Kaggle for once again providing a great learning opportunity.\n\n### Models\nFor most experiments I used `efficientnet_b0`. During final days, I added `efficientnetv2s` and `efficientnetb1`, `efficientnetb2` to ensemble, which is a geometric mean of different runs.\n(I missed the trick here by not adding bigger models).\n\n### Input\nFor most of my experiments I used all 6 inputs stacked by time and resized to either 512x512 or 768x512.\nFor most most experiments, original stacked image is used as channel 0, frequency normalized image is used as channel 1, and 3rd channel is 0 and 1 basis whether that part of input is ON target or OFF target.\n\n### Training setup\nI extensively use config files and [hydra](https://hydra.cc/) to manage experiments. [Weights and biases](https://wandb.ai/site) for logging and monitoring. [Pytorch lightning](https://github.com/PyTorchLightning/pytorch-lightning) for boiler plate training code. Most of the code is **flake8** linted and auto formatted by [black](https://github.com/psf/black). Great thanks to `https://github.com/ashleve/lightning-hydra-template` for providing a good starter code base. Thanks to all the developers of those libraries!\n\nAll my codes are available here: https://github.com/mohsinkhn/seti-kaggle-tezdhar\n\n\n### Experiments:\n* Including off channels gives small boost on CV but a bigger one (0.1-.2) on LB for some of my early models. It could be explained by narrowband drift signals existing in both on-off channels (more so in test data).\n\n* Untied mixup ( https://openreview.net/pdf?id=SkgjKR4YwH ). The basic idea is to have different distributions for mixing of inputs and targets. For inputs I use beta distribution with alpha=1.0, beta=1.0 while for targets I use beta distribution with alpha=1.0, beta=0.5.\n\n* Augmentations: To detect unknown classes, I wanted to enforce the condition that signal should be present in only ON channels and NOT OFF inputs. To this end, I roll data in inputs by 1 step ( [A0, B, A1, C, A2, D] --> [D, A0, B, A1, C, A2]) and make target 0. The image in input 0 becomes image in input 1, image in input 1 becomes image in input 2 and so on. **Doing so, along with SpecAug, I am able to successfully predict new class in test data**\n\n* Other than above augmentation, I used hflip, roll, specAug, meanshift, verticalshift(2 steps as 6 input image can't be vertically flipped)\n\n* I tried many normalisation techniques, mean absolute deviation scaling, sigmoid scaling, random power after min-max scaling etc. but none of them gave significant boost.\n\n* Tried pseudo labelling towards the end with only confident predictions, but couldn't get much benefit.\n\n\nAlthough, feels bad to loose solo gold by 1 rank, it was really good learning experience for me, especially my experimentation stack has improved significantly. .\n\n\n\n\n",
    "1482696": "> Augmentations: To detect unknown classes, I wanted to enforce the condition that signal should be present in only ON channels and NOT OFF inputs. To this end, I roll data in inputs by 1 step and make target 0. The image in input 0 becomes image in input 1, image in input 1 becomes image in input 2 and so on. Doing so, along with SpecAug, I am able to successfully predict new class in test data\n\nOK I love this. It's like the trick I used of adding an indicator channel [(link)](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266460), but better, because it FORCES the network to be aware of the location of the needle in terms of on/off, rather than just overfitting to its shape. Damn, wish I'd thought of this!",
    "1489800": "Awesome work, thanks for sharing, you will get Gold medal shortly in your next competition, all the best!",
    "1488211": "Thanks for sharing the approach. This will help us a lot in future cases...Eventually, you will get Gold Keep going buddy. @tezdhar ",
    "1483536": "thx for sharing. new stuff learned.",
    "1482583": "Thanks, I also tried to do different distributions with mixup, didnt go so well \n1. (~~that mixup link is broken can you please fix that?~~  there is closing bracket in url ). \n2. If you back ON/OFF layers as 3ch input you should be able to do any flips? \n3. And about roll-aug: so AAA cadences become BCD, and even thou there is a signal in them, sample becomes non-anomaly. Am i right? Did it work? ",
    "1482652": "Thanks for sharing @tezdhar , especially with the code. You are always impressive to climb up the leaderboard so fast at late stage of competition, I beleive you will get another solo gold and become Grandmaster soon. 🎉\n\nRegarding your roll-aug, if I understand correctly, you roll the positive sample like [A0, B, A1, C, A2, D] to [B, A1, C, A2, D, A0], then this sample's target will become 0, right? If so there will be fewer postive training sample, make the dataset even more imblance, in my understanding this would hurt the performance, could you please explain a bit here? Thanks.",
    "1495848": "Good job =))",
    "1482529": ""
  }
}