{
  "id": 266396,
  "title": "4th Place Solution. Steven Signal",
  "url": "/competitions/seti-breakthrough-listen/discussion/266396",
  "author_name": "Gleb",
  "post_date": "2021-08-19T01:46:25.533000",
  "votes": 30,
  "comment_count": 5,
  "views": 0,
  "content": "<h1>Intro</h1>\n<p>First of all, we want to thank the organizers and the Kaggle team for hosting the competition. We believe that someday AI will open a big road to exploring the Universe.</p>\n<p>Besides that, we want to point out the following projects:</p>\n<ul>\n<li><a href=\"https://pypi.org/project/timm/\" target=\"_blank\">timm, rwightman</a></li>\n<li><a href=\"https://pytorch.org/\" target=\"_blank\">pytorch</a></li>\n<li><a href=\"https://github.com/qubvel/\" target=\"_blank\">qubvel, TTA, cls models, seg models, etc. Great stuff</a></li>\n<li>feel free to mention any project/author we forget about. We are happy to update this post with references</li>\n</ul>\n<h1>TLDR</h1>\n<p>Classification ensemble with MSDA and focal loss. 2 models types: nf-regnetb1, HRNet18. Spatial cadence with guide layer for keeping ABACAD information</p>\n<h1>Input, preprocessing</h1>\n<p>Most of our experiments were on original resolution (img/sec reasons). Within the last couple of weeks, we scaled up to current image sizes.</p>\n<ul>\n<li><p>We use spatial vstack, 6 frames side by side, as most of you did. Samples upscaled with cubic interpolation x3 by frequency axis, scaled down to 6x160 by time axis. The resulting sample size is 960x768</p></li>\n<li><p>We use 'guide layer' as the second input layer, mask 960x768  of -1/1 structure: 1 for 'A' frames,-1 for 'B', 'C', 'D' frames</p></li>\n<li><p>NN input is ~70% random crop from 960x768x2</p></li>\n</ul>\n<p>Augmentation (on batches, hand-coded and bit of kornia):</p>\n<ul>\n<li>Random Crop</li>\n<li>HFlip, VFlip</li>\n<li>Special tuned MSDA</li>\n</ul>\n<p>CV split 15k / 45k, cv is correlated with LB</p>\n<h3>MSDA</h3>\n<p>We feel like MSDA is one of the key parts of this solution.</p>\n<ul>\n<li><p><a href=\"https://arxiv.org/abs/1710.09412\" target=\"_blank\">MixUp</a> Used on all batches, alpha=2. <strong>But</strong> mixing any label within an anomaly will result in an anomaly label. Labels and permuted labels are the same. It was done not to penalize NN with loss from the mixed anomaly. So one can say it is not MixUp at all.</p>\n<ul>\n<li><a href=\"https://arxiv.org/abs/1905.04899\" target=\"_blank\">CutMix</a> Used on all batches, with parameters from official implementation. <strong>But</strong> mixing happening in two separate sub-batches: the first is only anomaly batch, and the second is rest. It was done due to the inability to distinguish the position of the signal inside of the frame. So one can say it is not CutMix at all.</li></ul>\n<p>We tried other techniques like AugMix, but it's hard to scale up image processing speed with a larger image size. Therefore, we end up using only batch-wise augmentations.</p></li>\n</ul>\n<h1>Train</h1>\n<p>Used models: timm nf-regnetb1 - our light one, timm HRNet18 - larger one. Such pick is random, but we feel like regnet ideas are correct, and HRnet looks like a good choice for this competition.</p>\n<p>Training time for nf-regnetb1 is ~7hrs, for HRNet ~24hrs (4xV100)</p>\n<h3>Train Features</h3>\n<ul>\n<li><a href=\"https://arxiv.org/abs/1708.02002\" target=\"_blank\">FocalLoss</a>, gamma=2, alpha=.7</li>\n<li><a href=\"https://github.com/pytorch/examples/blob/master/distributed/ddp/README.md\" target=\"_blank\">DDP</a> multiple GPUs</li>\n<li><a href=\"https://pytorch.org/docs/stable/notes/amp_examples.html\" target=\"_blank\">AMP</a> to reduces memory requirements for training models, enabling larger models or larger batches</li>\n<li><a href=\"https://fastai.github.io/timmdocs/training_modelEMA\" target=\"_blank\">EMA</a> (timm)</li>\n<li>usual LR stuff, warmup, cosine decay </li>\n<li>early in competition SAM, but with AMP and DDP it requires much careful tuning, fails with numerical instability (and x2 training time)</li>\n<li>all data in the tensor dataset, sharded between GPUs (which is not a good idea accuracy-wise, but it is fast)</li>\n</ul>\n<h1>Inference, postprocess</h1>\n<p>TTA: flips, multiplier, sometimes additional Gaussian noise(mean=0, std=[0,.1,.2]).</p>\n<p>Ensemble, merge is simple average:</p>\n<ul>\n<li><a href=\"https://arxiv.org/pdf/2101.08692.pdf\" target=\"_blank\">nf-regnetb1</a>, 384, 768, split1 </li>\n<li><a href=\"https://arxiv.org/pdf/2101.08692.pdf\" target=\"_blank\">nf-regnetb1</a>, 384, 768, split2, minor MSDA changes</li>\n<li><a href=\"https://github.com/HRNet/HRNet-Image-Classification\" target=\"_blank\">HRNet18</a>, 448, 896, split1</li>\n<li><a href=\"https://github.com/HRNet/HRNet-Image-Classification\" target=\"_blank\">HRNet18</a>, 448, 896, split1, minor MSDA changes, Focal with lower gamma and alpha</li>\n</ul>\n<p>CV score for ensemble parts: .901, .901, .908, .908</p>\n<p>LB: .794, .794, .797, .797</p>\n<h1>LB basic roadmap</h1>\n<ul>\n<li>.780 simple cnn, but with MSDA</li>\n<li>.783 +guide layer</li>\n<li>.788 +focal </li>\n<li>.794 +upscale resolution</li>\n<li>.797 +heavy model</li>\n<li>.800 +ensembles</li>\n</ul>\n<p>Update. Best single model  HRNet18 488x896 PB 0.80095; LB 0.79766</p>\n<h1>Shortlist of things that did not work</h1>\n<p>Channel-wise cadence, consistency loss, AUC proxy loss, frequency normalizing, whitening, dynamic sampling, selective sampling, autoencoders, VAE, VQVAE, MEMAE, more of XXXAE and unsupervised methods, <em>PSEUDO LABELING, best was non-harming one</em>, KNN, DBSCAN clustering of train/ test with adversarial models, frames decomposition, transformers, logits ensembles, shuffling cadences, creating additional data from test B, C, D frames, space2depth and other stems.</p>\n<p>Probably a personal record.</p>\n<hr>\n<h3>P.S.</h3>\n<p>There was a belief in finding some data insights that could give that jump to the final score. However, spending nearly all of our time and efforts were not worth it. </p>\n<p>As we fail to use almost all of our data-related discoveries (such as difference between train and test) we left that writing to other teams.</p>\n<h3>Upd after reading 1-st place, congrats btw</h3>\n<ol>\n<li>We used only new data, old data did give a little boost to us, so little that we dropped that.</li>\n<li>We also find sin-shaped signal and also tried to add it with setigen/handcode, but with no improvements to score</li>\n<li>We also added near-1 probabilities as pseudo labels, but score immediately  drop to .75. So we did DBSCAN clustering on test images and select certain cluster, that 'base', (not a cluster). Adding it as pseudo labels dropped score only slightly, and basically we could not go further then that.</li>\n<li>Cleaning images for us ended at trimmed normalization by frequency and time, which visually helped, but not scoring-wise </li>\n</ol>",
  "messages": [
    {
      "id": 1480361,
      "postDate": "2021-08-19T01:46:25.533Z",
      "content": "<h1>Intro</h1>\n<p>First of all, we want to thank the organizers and the Kaggle team for hosting the competition. We believe that someday AI will open a big road to exploring the Universe.</p>\n<p>Besides that, we want to point out the following projects:</p>\n<ul>\n<li><a href=\"https://pypi.org/project/timm/\" target=\"_blank\">timm, rwightman</a></li>\n<li><a href=\"https://pytorch.org/\" target=\"_blank\">pytorch</a></li>\n<li><a href=\"https://github.com/qubvel/\" target=\"_blank\">qubvel, TTA, cls models, seg models, etc. Great stuff</a></li>\n<li>feel free to mention any project/author we forget about. We are happy to update this post with references</li>\n</ul>\n<h1>TLDR</h1>\n<p>Classification ensemble with MSDA and focal loss. 2 models types: nf-regnetb1, HRNet18. Spatial cadence with guide layer for keeping ABACAD information</p>\n<h1>Input, preprocessing</h1>\n<p>Most of our experiments were on original resolution (img/sec reasons). Within the last couple of weeks, we scaled up to current image sizes.</p>\n<ul>\n<li><p>We use spatial vstack, 6 frames side by side, as most of you did. Samples upscaled with cubic interpolation x3 by frequency axis, scaled down to 6x160 by time axis. The resulting sample size is 960x768</p></li>\n<li><p>We use 'guide layer' as the second input layer, mask 960x768  of -1/1 structure: 1 for 'A' frames,-1 for 'B', 'C', 'D' frames</p></li>\n<li><p>NN input is ~70% random crop from 960x768x2</p></li>\n</ul>\n<p>Augmentation (on batches, hand-coded and bit of kornia):</p>\n<ul>\n<li>Random Crop</li>\n<li>HFlip, VFlip</li>\n<li>Special tuned MSDA</li>\n</ul>\n<p>CV split 15k / 45k, cv is correlated with LB</p>\n<h3>MSDA</h3>\n<p>We feel like MSDA is one of the key parts of this solution.</p>\n<ul>\n<li><p><a href=\"https://arxiv.org/abs/1710.09412\" target=\"_blank\">MixUp</a> Used on all batches, alpha=2. <strong>But</strong> mixing any label within an anomaly will result in an anomaly label. Labels and permuted labels are the same. It was done not to penalize NN with loss from the mixed anomaly. So one can say it is not MixUp at all.</p>\n<ul>\n<li><a href=\"https://arxiv.org/abs/1905.04899\" target=\"_blank\">CutMix</a> Used on all batches, with parameters from official implementation. <strong>But</strong> mixing happening in two separate sub-batches: the first is only anomaly batch, and the second is rest. It was done due to the inability to distinguish the position of the signal inside of the frame. So one can say it is not CutMix at all.</li></ul>\n<p>We tried other techniques like AugMix, but it's hard to scale up image processing speed with a larger image size. Therefore, we end up using only batch-wise augmentations.</p></li>\n</ul>\n<h1>Train</h1>\n<p>Used models: timm nf-regnetb1 - our light one, timm HRNet18 - larger one. Such pick is random, but we feel like regnet ideas are correct, and HRnet looks like a good choice for this competition.</p>\n<p>Training time for nf-regnetb1 is ~7hrs, for HRNet ~24hrs (4xV100)</p>\n<h3>Train Features</h3>\n<ul>\n<li><a href=\"https://arxiv.org/abs/1708.02002\" target=\"_blank\">FocalLoss</a>, gamma=2, alpha=.7</li>\n<li><a href=\"https://github.com/pytorch/examples/blob/master/distributed/ddp/README.md\" target=\"_blank\">DDP</a> multiple GPUs</li>\n<li><a href=\"https://pytorch.org/docs/stable/notes/amp_examples.html\" target=\"_blank\">AMP</a> to reduces memory requirements for training models, enabling larger models or larger batches</li>\n<li><a href=\"https://fastai.github.io/timmdocs/training_modelEMA\" target=\"_blank\">EMA</a> (timm)</li>\n<li>usual LR stuff, warmup, cosine decay </li>\n<li>early in competition SAM, but with AMP and DDP it requires much careful tuning, fails with numerical instability (and x2 training time)</li>\n<li>all data in the tensor dataset, sharded between GPUs (which is not a good idea accuracy-wise, but it is fast)</li>\n</ul>\n<h1>Inference, postprocess</h1>\n<p>TTA: flips, multiplier, sometimes additional Gaussian noise(mean=0, std=[0,.1,.2]).</p>\n<p>Ensemble, merge is simple average:</p>\n<ul>\n<li><a href=\"https://arxiv.org/pdf/2101.08692.pdf\" target=\"_blank\">nf-regnetb1</a>, 384, 768, split1 </li>\n<li><a href=\"https://arxiv.org/pdf/2101.08692.pdf\" target=\"_blank\">nf-regnetb1</a>, 384, 768, split2, minor MSDA changes</li>\n<li><a href=\"https://github.com/HRNet/HRNet-Image-Classification\" target=\"_blank\">HRNet18</a>, 448, 896, split1</li>\n<li><a href=\"https://github.com/HRNet/HRNet-Image-Classification\" target=\"_blank\">HRNet18</a>, 448, 896, split1, minor MSDA changes, Focal with lower gamma and alpha</li>\n</ul>\n<p>CV score for ensemble parts: .901, .901, .908, .908</p>\n<p>LB: .794, .794, .797, .797</p>\n<h1>LB basic roadmap</h1>\n<ul>\n<li>.780 simple cnn, but with MSDA</li>\n<li>.783 +guide layer</li>\n<li>.788 +focal </li>\n<li>.794 +upscale resolution</li>\n<li>.797 +heavy model</li>\n<li>.800 +ensembles</li>\n</ul>\n<p>Update. Best single model  HRNet18 488x896 PB 0.80095; LB 0.79766</p>\n<h1>Shortlist of things that did not work</h1>\n<p>Channel-wise cadence, consistency loss, AUC proxy loss, frequency normalizing, whitening, dynamic sampling, selective sampling, autoencoders, VAE, VQVAE, MEMAE, more of XXXAE and unsupervised methods, <em>PSEUDO LABELING, best was non-harming one</em>, KNN, DBSCAN clustering of train/ test with adversarial models, frames decomposition, transformers, logits ensembles, shuffling cadences, creating additional data from test B, C, D frames, space2depth and other stems.</p>\n<p>Probably a personal record.</p>\n<hr>\n<h3>P.S.</h3>\n<p>There was a belief in finding some data insights that could give that jump to the final score. However, spending nearly all of our time and efforts were not worth it. </p>\n<p>As we fail to use almost all of our data-related discoveries (such as difference between train and test) we left that writing to other teams.</p>\n<h3>Upd after reading 1-st place, congrats btw</h3>\n<ol>\n<li>We used only new data, old data did give a little boost to us, so little that we dropped that.</li>\n<li>We also find sin-shaped signal and also tried to add it with setigen/handcode, but with no improvements to score</li>\n<li>We also added near-1 probabilities as pseudo labels, but score immediately  drop to .75. So we did DBSCAN clustering on test images and select certain cluster, that 'base', (not a cluster). Adding it as pseudo labels dropped score only slightly, and basically we could not go further then that.</li>\n<li>Cleaning images for us ended at trimmed normalization by frequency and time, which visually helped, but not scoring-wise </li>\n</ol>",
      "rawMarkdown": "# Intro\n\nFirst of all, we want to thank the organizers and the Kaggle team for hosting the competition. We believe that someday AI will open a big road to exploring the Universe.\n\nBesides that, we want to point out the following projects:\n- [timm, rwightman](https://pypi.org/project/timm/)\n- [pytorch](https://pytorch.org/)\n- [qubvel, TTA, cls models, seg models, etc. Great stuff](https://github.com/qubvel/)\n- feel free to mention any project/author we forget about. We are happy to update this post with references\n\n# TLDR\n\nClassification ensemble with MSDA and focal loss. 2 models types: nf-regnetb1, HRNet18. Spatial cadence with guide layer for keeping ABACAD information\n\n\n# Input, preprocessing\n\nMost of our experiments were on original resolution (img/sec reasons). Within the last couple of weeks, we scaled up to current image sizes.\n\n - We use spatial vstack, 6 frames side by side, as most of you did. Samples upscaled with cubic interpolation x3 by frequency axis, scaled down to 6x160 by time axis. The resulting sample size is 960x768\n\n - We use 'guide layer' as the second input layer, mask 960x768  of -1/1 structure: 1 for 'A' frames,-1 for 'B', 'C', 'D' frames\n\n - NN input is ~70% random crop from 960x768x2\n\nAugmentation (on batches, hand-coded and bit of kornia):\n- Random Crop\n- HFlip, VFlip\n- Special tuned MSDA\n\nCV split 15k / 45k, cv is correlated with LB\n\n### MSDA\n\nWe feel like MSDA is one of the key parts of this solution.\n\n - [MixUp](https://arxiv.org/abs/1710.09412) Used on all batches, alpha=2. **But** mixing any label within an anomaly will result in an anomaly label. Labels and permuted labels are the same. It was done not to penalize NN with loss from the mixed anomaly. So one can say it is not MixUp at all.\n\n- [CutMix](https://arxiv.org/abs/1905.04899) Used on all batches, with parameters from official implementation. **But** mixing happening in two separate sub-batches: the first is only anomaly batch, and the second is rest. It was done due to the inability to distinguish the position of the signal inside of the frame. So one can say it is not CutMix at all.\n \n We tried other techniques like AugMix, but it's hard to scale up image processing speed with a larger image size. Therefore, we end up using only batch-wise augmentations.\n\n# Train\nUsed models: timm nf-regnetb1 - our light one, timm HRNet18 - larger one. Such pick is random, but we feel like regnet ideas are correct, and HRnet looks like a good choice for this competition.\n\nTraining time for nf-regnetb1 is ~7hrs, for HRNet ~24hrs (4xV100)\n\n### Train Features\n\n - [FocalLoss](https://arxiv.org/abs/1708.02002), gamma=2, alpha=.7\n - [DDP](https://github.com/pytorch/examples/blob/master/distributed/ddp/README.md) multiple GPUs\n - [AMP](https://pytorch.org/docs/stable/notes/amp_examples.html) to reduces memory requirements for training models, enabling larger models or larger batches\n - [EMA](https://fastai.github.io/timmdocs/training_modelEMA) (timm)\n - usual LR stuff, warmup, cosine decay \n - early in competition SAM, but with AMP and DDP it requires much careful tuning, fails with numerical instability (and x2 training time)\n - all data in the tensor dataset, sharded between GPUs (which is not a good idea accuracy-wise, but it is fast)\n\n\n# Inference, postprocess\n\nTTA: flips, multiplier, sometimes additional Gaussian noise(mean=0, std=[0,.1,.2]).\n\nEnsemble, merge is simple average:\n\n - [nf-regnetb1](https://arxiv.org/pdf/2101.08692.pdf), 384, 768, split1 \n - [nf-regnetb1](https://arxiv.org/pdf/2101.08692.pdf), 384, 768, split2, minor MSDA changes\n - [HRNet18](https://github.com/HRNet/HRNet-Image-Classification), 448, 896, split1\n - [HRNet18](https://github.com/HRNet/HRNet-Image-Classification), 448, 896, split1, minor MSDA changes, Focal with lower gamma and alpha\n\nCV score for ensemble parts: .901, .901, .908, .908\n\nLB: .794, .794, .797, .797\n\n# LB basic roadmap\n\n- .780 simple cnn, but with MSDA\n- .783 +guide layer\n- .788 +focal \n- .794 +upscale resolution\n- .797 +heavy model\n- .800 +ensembles\n\nUpdate. Best single model  HRNet18 488x896 PB 0.80095; LB 0.79766\n\n# Shortlist of things that did not work\n\nChannel-wise cadence, consistency loss, AUC proxy loss, frequency normalizing, whitening, dynamic sampling, selective sampling, autoencoders, VAE, VQVAE, MEMAE, more of XXXAE and unsupervised methods, *PSEUDO LABELING, best was non-harming one*, KNN, DBSCAN clustering of train/ test with adversarial models, frames decomposition, transformers, logits ensembles, shuffling cadences, creating additional data from test B, C, D frames, space2depth and other stems.\n\nProbably a personal record.\n\n_______________\n\n### P.S.\nThere was a belief in finding some data insights that could give that jump to the final score. However, spending nearly all of our time and efforts were not worth it. \n\nAs we fail to use almost all of our data-related discoveries (such as difference between train and test) we left that writing to other teams.\n\n### Upd after reading 1-st place, congrats btw\n\n1. We used only new data, old data did give a little boost to us, so little that we dropped that.\n2. We also find sin-shaped signal and also tried to add it with setigen/handcode, but with no improvements to score\n3. We also added near-1 probabilities as pseudo labels, but score immediately  drop to .75. So we did DBSCAN clustering on test images and select certain cluster, that 'base', (not a cluster). Adding it as pseudo labels dropped score only slightly, and basically we could not go further then that.\n4. Cleaning images for us ended at trimmed normalization by frequency and time, which visually helped, but not scoring-wise \n",
      "votes": 30
    },
    {
      "id": 1481078,
      "postDate": "2021-08-19T09:49:43.693Z",
      "content": "<p>Wow! It's really technical solution.<br>\nGreat work!</p>",
      "rawMarkdown": "Wow! It's really technical solution.\nGreat work!",
      "votes": 1
    },
    {
      "id": 1480728,
      "postDate": "2021-08-19T06:41:00.060Z",
      "content": "<p>Congrats on 4th Place<br>\nVery interesting solution <a href=\"https://www.kaggle.com/bakeryproducts\" target=\"_blank\">@bakeryproducts</a> especially the upscaling and preprocessing part , thanks for sharing , can you share the code for your solution as well, it will be great for us to learn</p>",
      "rawMarkdown": "Congrats on 4th Place\nVery interesting solution @bakeryproducts especially the upscaling and preprocessing part , thanks for sharing , can you share the code for your solution as well, it will be great for us to learn",
      "replies": [
        {
          "id": 1481666,
          "postDate": "2021-08-19T16:09:44.237Z",
          "content": "<p>Thanks, we didnt plan to, but maybe in future</p>",
          "rawMarkdown": "Thanks, we didnt plan to, but maybe in future"
        }
      ]
    },
    {
      "id": 1480472,
      "postDate": "2021-08-19T03:27:02.647Z",
      "content": "<p>Very cool. Soloing the comp especially the last 3 weeks where discussions and kernels were all but dead (except for that one funny ensemble rickroll kernel), sometimes it felt like there's nothing interesting to learn. However reading through your solution has put a few good things on my radar. Thank you for sharing.</p>",
      "rawMarkdown": "Very cool. Soloing the comp especially the last 3 weeks where discussions and kernels were all but dead (except for that one funny ensemble rickroll kernel), sometimes it felt like there's nothing interesting to learn. However reading through your solution has put a few good things on my radar. Thank you for sharing.",
      "replies": [
        {
          "id": 1481662,
          "postDate": "2021-08-19T16:09:11.950Z",
          "content": "<p>And this is what i dont like about platform. We cant really start discussion until its over and people are tired after a month or more of intensive research. Im happy you get something from my write-up.</p>",
          "rawMarkdown": "And this is what i dont like about platform. We cant really start discussion until its over and people are tired after a month or more of intensive research. Im happy you get something from my write-up."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1481078,
      "author_name": "WOOSUNG YOON",
      "author_url": "",
      "post_date": "2021-08-19T09:49:43.693000",
      "content": "<p>Wow! It's really technical solution.<br>\nGreat work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1480728,
      "author_name": "Mr_KnowNothing",
      "author_url": "",
      "post_date": "2021-08-19T06:41:00.060000",
      "content": "<p>Congrats on 4th Place<br>\nVery interesting solution <a href=\"https://www.kaggle.com/bakeryproducts\" target=\"_blank\">@bakeryproducts</a> especially the upscaling and preprocessing part , thanks for sharing , can you share the code for your solution as well, it will be great for us to learn</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1481666,
          "author_name": "Gleb",
          "author_url": "",
          "post_date": "2021-08-19T16:09:44.237000",
          "content": "<p>Thanks, we didnt plan to, but maybe in future</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1480472,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2021-08-19T03:27:02.647000",
      "content": "<p>Very cool. Soloing the comp especially the last 3 weeks where discussions and kernels were all but dead (except for that one funny ensemble rickroll kernel), sometimes it felt like there's nothing interesting to learn. However reading through your solution has put a few good things on my radar. Thank you for sharing.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1481662,
          "author_name": "Gleb",
          "author_url": "",
          "post_date": "2021-08-19T16:09:11.950000",
          "content": "<p>And this is what i dont like about platform. We cant really start discussion until its over and people are tired after a month or more of intensive research. Im happy you get something from my write-up.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1480361": "# Intro\n\nFirst of all, we want to thank the organizers and the Kaggle team for hosting the competition. We believe that someday AI will open a big road to exploring the Universe.\n\nBesides that, we want to point out the following projects:\n- [timm, rwightman](https://pypi.org/project/timm/)\n- [pytorch](https://pytorch.org/)\n- [qubvel, TTA, cls models, seg models, etc. Great stuff](https://github.com/qubvel/)\n- feel free to mention any project/author we forget about. We are happy to update this post with references\n\n# TLDR\n\nClassification ensemble with MSDA and focal loss. 2 models types: nf-regnetb1, HRNet18. Spatial cadence with guide layer for keeping ABACAD information\n\n\n# Input, preprocessing\n\nMost of our experiments were on original resolution (img/sec reasons). Within the last couple of weeks, we scaled up to current image sizes.\n\n - We use spatial vstack, 6 frames side by side, as most of you did. Samples upscaled with cubic interpolation x3 by frequency axis, scaled down to 6x160 by time axis. The resulting sample size is 960x768\n\n - We use 'guide layer' as the second input layer, mask 960x768  of -1/1 structure: 1 for 'A' frames,-1 for 'B', 'C', 'D' frames\n\n - NN input is ~70% random crop from 960x768x2\n\nAugmentation (on batches, hand-coded and bit of kornia):\n- Random Crop\n- HFlip, VFlip\n- Special tuned MSDA\n\nCV split 15k / 45k, cv is correlated with LB\n\n### MSDA\n\nWe feel like MSDA is one of the key parts of this solution.\n\n - [MixUp](https://arxiv.org/abs/1710.09412) Used on all batches, alpha=2. **But** mixing any label within an anomaly will result in an anomaly label. Labels and permuted labels are the same. It was done not to penalize NN with loss from the mixed anomaly. So one can say it is not MixUp at all.\n\n- [CutMix](https://arxiv.org/abs/1905.04899) Used on all batches, with parameters from official implementation. **But** mixing happening in two separate sub-batches: the first is only anomaly batch, and the second is rest. It was done due to the inability to distinguish the position of the signal inside of the frame. So one can say it is not CutMix at all.\n \n We tried other techniques like AugMix, but it's hard to scale up image processing speed with a larger image size. Therefore, we end up using only batch-wise augmentations.\n\n# Train\nUsed models: timm nf-regnetb1 - our light one, timm HRNet18 - larger one. Such pick is random, but we feel like regnet ideas are correct, and HRnet looks like a good choice for this competition.\n\nTraining time for nf-regnetb1 is ~7hrs, for HRNet ~24hrs (4xV100)\n\n### Train Features\n\n - [FocalLoss](https://arxiv.org/abs/1708.02002), gamma=2, alpha=.7\n - [DDP](https://github.com/pytorch/examples/blob/master/distributed/ddp/README.md) multiple GPUs\n - [AMP](https://pytorch.org/docs/stable/notes/amp_examples.html) to reduces memory requirements for training models, enabling larger models or larger batches\n - [EMA](https://fastai.github.io/timmdocs/training_modelEMA) (timm)\n - usual LR stuff, warmup, cosine decay \n - early in competition SAM, but with AMP and DDP it requires much careful tuning, fails with numerical instability (and x2 training time)\n - all data in the tensor dataset, sharded between GPUs (which is not a good idea accuracy-wise, but it is fast)\n\n\n# Inference, postprocess\n\nTTA: flips, multiplier, sometimes additional Gaussian noise(mean=0, std=[0,.1,.2]).\n\nEnsemble, merge is simple average:\n\n - [nf-regnetb1](https://arxiv.org/pdf/2101.08692.pdf), 384, 768, split1 \n - [nf-regnetb1](https://arxiv.org/pdf/2101.08692.pdf), 384, 768, split2, minor MSDA changes\n - [HRNet18](https://github.com/HRNet/HRNet-Image-Classification), 448, 896, split1\n - [HRNet18](https://github.com/HRNet/HRNet-Image-Classification), 448, 896, split1, minor MSDA changes, Focal with lower gamma and alpha\n\nCV score for ensemble parts: .901, .901, .908, .908\n\nLB: .794, .794, .797, .797\n\n# LB basic roadmap\n\n- .780 simple cnn, but with MSDA\n- .783 +guide layer\n- .788 +focal \n- .794 +upscale resolution\n- .797 +heavy model\n- .800 +ensembles\n\nUpdate. Best single model  HRNet18 488x896 PB 0.80095; LB 0.79766\n\n# Shortlist of things that did not work\n\nChannel-wise cadence, consistency loss, AUC proxy loss, frequency normalizing, whitening, dynamic sampling, selective sampling, autoencoders, VAE, VQVAE, MEMAE, more of XXXAE and unsupervised methods, *PSEUDO LABELING, best was non-harming one*, KNN, DBSCAN clustering of train/ test with adversarial models, frames decomposition, transformers, logits ensembles, shuffling cadences, creating additional data from test B, C, D frames, space2depth and other stems.\n\nProbably a personal record.\n\n_______________\n\n### P.S.\nThere was a belief in finding some data insights that could give that jump to the final score. However, spending nearly all of our time and efforts were not worth it. \n\nAs we fail to use almost all of our data-related discoveries (such as difference between train and test) we left that writing to other teams.\n\n### Upd after reading 1-st place, congrats btw\n\n1. We used only new data, old data did give a little boost to us, so little that we dropped that.\n2. We also find sin-shaped signal and also tried to add it with setigen/handcode, but with no improvements to score\n3. We also added near-1 probabilities as pseudo labels, but score immediately  drop to .75. So we did DBSCAN clustering on test images and select certain cluster, that 'base', (not a cluster). Adding it as pseudo labels dropped score only slightly, and basically we could not go further then that.\n4. Cleaning images for us ended at trimmed normalization by frequency and time, which visually helped, but not scoring-wise \n",
    "1481078": "Wow! It's really technical solution.\nGreat work!",
    "1480728": "Congrats on 4th Place\nVery interesting solution @bakeryproducts especially the upscaling and preprocessing part , thanks for sharing , can you share the code for your solution as well, it will be great for us to learn",
    "1480472": "Very cool. Soloing the comp especially the last 3 weeks where discussions and kernels were all but dead (except for that one funny ensemble rickroll kernel), sometimes it felt like there's nothing interesting to learn. However reading through your solution has put a few good things on my radar. Thank you for sharing."
  }
}