{
  "id": 118186,
  "title": "Quick Silver - The Late Joiners' Journey",
  "url": "/competitions/understanding_cloud_organization/discussion/118186",
  "author_name": "Borys Tymchenko",
  "post_date": "2019-11-19T23:23:33.287000",
  "votes": 13,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hello and congrats to all participants!\nThanks, organizers for this incredible competition!</p>\n\n<p>We entered this competition after a good rest from Severstal competition.\nActually, I made the repo for it 12 days before the final deadline.</p>\n\n<p>We used our pipeline from Severstal competition and changed only separate parts of it.\nAt first, I tried the best scoring config from Severstal competition for this data, only hoping for \"easy gold\" but got the rank around ~900.\nThis is my second competition, where the 'fit_predict' could not get into the bronze zone.</p>\n\n<p><strong>At first, here are the things that we tried, and that did not work at all:</strong></p>\n\n<ul>\n<li>Higher than 352x576 resolution,</li>\n<li>Adding heavy encoders (everything larger than Densenet169/SEResNeXt50 overfitted like crazy)</li>\n<li>Masking that black stripe with mean value or overlaying with a black rectangle,</li>\n<li>Training on crops,</li>\n<li>Separate classifier,</li>\n<li>Unet decoder,</li>\n<li>PSP decoder,</li>\n<li>Hard color augmentations,</li>\n<li>Decoder from <a href=\"/hengck23\">@hengck23</a> (but was great to learn from!),</li>\n<li>BCE Loss, Focal Loss,  Dice loss or Lovasz loss (alone),</li>\n<li>Symmetric CrossEntropy (<a href=\"https://arxiv.org/pdf/1908.06112.pdf\">https://arxiv.org/pdf/1908.06112.pdf</a>)</li>\n</ul>\n\n<p><strong>Things that worked:</strong>\n<a href=\"/spsancti\">@spsancti</a>:\n- Combination of Trimmed BCE and Dice loss,\n- Separate classification head,\n- EfficientNet-B0 and B1 encoders,\n- Pseudolabeling and Knowledge distillation</p>\n\n<p><a href=\"/hakuryuu\">@hakuryuu</a>:\n- Small networks (Resnet34, EfficientNet-B0...B2),\n- Grid shuffle augmentation,\n- Sum of  symmetric Lovasz, Trimmed Focal, Dice and BCE losses</p>\n\n<p><strong>We both used:</strong>\n- Cosine Annealing LR schedule,\n- FPN decoder,\n- Hard geometric augmentations,\n- Over9000 optimizer (better than RAdam on all runs)\n- Threshold selection on the validation</p>\n\n<p><strong>Libraries used:</strong>\n- Catalyst and Apex\n- Albumentations\n- <a href=\"/pavel92\">@pavel92</a> 's awesome segmentation models library</p>\n\n<p>We trained single-fold models, with different split, to get more boost after merging them.</p>\n\n<p><strong>Pseudolabelling and Knowledge distillation</strong></p>\n\n<p>One day before the deadline, I decided to try to use pseudolabelling. \nI took our most successful submit for that time (0.6575 public) and used it as hard pseudo labels. I did not select any confident predictions assuming that the source markings are noisy enough to tolerate some more noise from pseudo labels. \nI pretrained the model on pseudo labels and continued to the regular train. \nImmediately there was an improvement on local validation while training on pseudo labels, which quickly disappeared when I switched to the regular train. So, I trained several models with this scheme only to confirm this result.</p>\n\n<p>This made me think that noisy corners in original markings confuse models too much, and I decided to relabel training data with the same ensemble that produced this submission. \nWith relabeled training data mixed with pseudo labels, the single model could reach 0.68 on local validation. I submitted it and got the first of my single models to break the 0.66 barrier on public LB. </p>\n\n<p>These results looked too optimistic, so we thought that I created a selection bias, so we decided to select one ensemble with models trained with pseudo labels and KD and one without.</p>\n\n<p><strong>Ensembling</strong>\nIn our final submission we had 2 ensembles:\nThe first one was trained without pseudo labels and consisted of 10 models.\nThe second one was the same as first, but with 5 more models trained on KD mixed with pseudo labels. </p>\n\n<p>Every model in ensemble was wrapped with 8xTTA:\n- HorizontalFlip\n- VerticalFip\n- BrightnessUp\n- BrightnessDown</p>\n\n<p>Probably, the assumption of selection bias was correct, as ensemble trained with KD and pseudo labels scored 0.65982 on private LB, which is 0.0034 less than an ensemble without them. </p>\n\n<p><strong>Hardware</strong>\nWe used 2 servers with 1xV100 and 1 workstation with 2x1080Ti. \nWe thank VITech Lab and 3DLook for providing computing resources for this competition.</p>",
  "messages": [
    {
      "id": 677210,
      "postDate": "2019-11-19T23:23:33.287Z",
      "content": "<p>Hello and congrats to all participants!\nThanks, organizers for this incredible competition!</p>\n\n<p>We entered this competition after a good rest from Severstal competition.\nActually, I made the repo for it 12 days before the final deadline.</p>\n\n<p>We used our pipeline from Severstal competition and changed only separate parts of it.\nAt first, I tried the best scoring config from Severstal competition for this data, only hoping for \"easy gold\" but got the rank around ~900.\nThis is my second competition, where the 'fit_predict' could not get into the bronze zone.</p>\n\n<p><strong>At first, here are the things that we tried, and that did not work at all:</strong></p>\n\n<ul>\n<li>Higher than 352x576 resolution,</li>\n<li>Adding heavy encoders (everything larger than Densenet169/SEResNeXt50 overfitted like crazy)</li>\n<li>Masking that black stripe with mean value or overlaying with a black rectangle,</li>\n<li>Training on crops,</li>\n<li>Separate classifier,</li>\n<li>Unet decoder,</li>\n<li>PSP decoder,</li>\n<li>Hard color augmentations,</li>\n<li>Decoder from <a href=\"/hengck23\">@hengck23</a> (but was great to learn from!),</li>\n<li>BCE Loss, Focal Loss,  Dice loss or Lovasz loss (alone),</li>\n<li>Symmetric CrossEntropy (<a href=\"https://arxiv.org/pdf/1908.06112.pdf\">https://arxiv.org/pdf/1908.06112.pdf</a>)</li>\n</ul>\n\n<p><strong>Things that worked:</strong>\n<a href=\"/spsancti\">@spsancti</a>:\n- Combination of Trimmed BCE and Dice loss,\n- Separate classification head,\n- EfficientNet-B0 and B1 encoders,\n- Pseudolabeling and Knowledge distillation</p>\n\n<p><a href=\"/hakuryuu\">@hakuryuu</a>:\n- Small networks (Resnet34, EfficientNet-B0...B2),\n- Grid shuffle augmentation,\n- Sum of  symmetric Lovasz, Trimmed Focal, Dice and BCE losses</p>\n\n<p><strong>We both used:</strong>\n- Cosine Annealing LR schedule,\n- FPN decoder,\n- Hard geometric augmentations,\n- Over9000 optimizer (better than RAdam on all runs)\n- Threshold selection on the validation</p>\n\n<p><strong>Libraries used:</strong>\n- Catalyst and Apex\n- Albumentations\n- <a href=\"/pavel92\">@pavel92</a> 's awesome segmentation models library</p>\n\n<p>We trained single-fold models, with different split, to get more boost after merging them.</p>\n\n<p><strong>Pseudolabelling and Knowledge distillation</strong></p>\n\n<p>One day before the deadline, I decided to try to use pseudolabelling. \nI took our most successful submit for that time (0.6575 public) and used it as hard pseudo labels. I did not select any confident predictions assuming that the source markings are noisy enough to tolerate some more noise from pseudo labels. \nI pretrained the model on pseudo labels and continued to the regular train. \nImmediately there was an improvement on local validation while training on pseudo labels, which quickly disappeared when I switched to the regular train. So, I trained several models with this scheme only to confirm this result.</p>\n\n<p>This made me think that noisy corners in original markings confuse models too much, and I decided to relabel training data with the same ensemble that produced this submission. \nWith relabeled training data mixed with pseudo labels, the single model could reach 0.68 on local validation. I submitted it and got the first of my single models to break the 0.66 barrier on public LB. </p>\n\n<p>These results looked too optimistic, so we thought that I created a selection bias, so we decided to select one ensemble with models trained with pseudo labels and KD and one without.</p>\n\n<p><strong>Ensembling</strong>\nIn our final submission we had 2 ensembles:\nThe first one was trained without pseudo labels and consisted of 10 models.\nThe second one was the same as first, but with 5 more models trained on KD mixed with pseudo labels. </p>\n\n<p>Every model in ensemble was wrapped with 8xTTA:\n- HorizontalFlip\n- VerticalFip\n- BrightnessUp\n- BrightnessDown</p>\n\n<p>Probably, the assumption of selection bias was correct, as ensemble trained with KD and pseudo labels scored 0.65982 on private LB, which is 0.0034 less than an ensemble without them. </p>\n\n<p><strong>Hardware</strong>\nWe used 2 servers with 1xV100 and 1 workstation with 2x1080Ti. \nWe thank VITech Lab and 3DLook for providing computing resources for this competition.</p>",
      "rawMarkdown": "Hello and congrats to all participants!\nThanks, organizers for this incredible competition!\n\nWe entered this competition after a good rest from Severstal competition.\nActually, I made the repo for it 12 days before the final deadline.\n\nWe used our pipeline from Severstal competition and changed only separate parts of it.\nAt first, I tried the best scoring config from Severstal competition for this data, only hoping for \"easy gold\" but got the rank around ~900.\nThis is my second competition, where the 'fit_predict' could not get into the bronze zone.\n\n**At first, here are the things that we tried, and that did not work at all:**\n\n- Higher than 352x576 resolution,\n- Adding heavy encoders (everything larger than Densenet169/SEResNeXt50 overfitted like crazy)\n- Masking that black stripe with mean value or overlaying with a black rectangle,\n- Training on crops,\n- Separate classifier,\n- Unet decoder,\n- PSP decoder,\n- Hard color augmentations,\n- Decoder from @hengck23 (but was great to learn from!),\n- BCE Loss, Focal Loss,  Dice loss or Lovasz loss (alone),\n- Symmetric CrossEntropy (https://arxiv.org/pdf/1908.06112.pdf)\n\n**Things that worked:**\n@spsancti:\n- Combination of Trimmed BCE and Dice loss,\n- Separate classification head,\n- EfficientNet-B0 and B1 encoders,\n- Pseudolabeling and Knowledge distillation\n\n@hakuryuu:\n- Small networks (Resnet34, EfficientNet-B0...B2),\n- Grid shuffle augmentation,\n- Sum of  symmetric Lovasz, Trimmed Focal, Dice and BCE losses\n\n__We both used:__\n- Cosine Annealing LR schedule,\n- FPN decoder,\n- Hard geometric augmentations,\n- Over9000 optimizer (better than RAdam on all runs)\n- Threshold selection on the validation\n\n__Libraries used:__\n- Catalyst and Apex\n- Albumentations\n- @pavel92 's awesome segmentation models library\n\nWe trained single-fold models, with different split, to get more boost after merging them.\n\n**Pseudolabelling and Knowledge distillation**\n\nOne day before the deadline, I decided to try to use pseudolabelling. \nI took our most successful submit for that time (0.6575 public) and used it as hard pseudo labels. I did not select any confident predictions assuming that the source markings are noisy enough to tolerate some more noise from pseudo labels. \nI pretrained the model on pseudo labels and continued to the regular train. \nImmediately there was an improvement on local validation while training on pseudo labels, which quickly disappeared when I switched to the regular train. So, I trained several models with this scheme only to confirm this result.\n\nThis made me think that noisy corners in original markings confuse models too much, and I decided to relabel training data with the same ensemble that produced this submission. \nWith relabeled training data mixed with pseudo labels, the single model could reach 0.68 on local validation. I submitted it and got the first of my single models to break the 0.66 barrier on public LB. \n\nThese results looked too optimistic, so we thought that I created a selection bias, so we decided to select one ensemble with models trained with pseudo labels and KD and one without.\n\n**Ensembling**\nIn our final submission we had 2 ensembles:\nThe first one was trained without pseudo labels and consisted of 10 models.\nThe second one was the same as first, but with 5 more models trained on KD mixed with pseudo labels. \n\nEvery model in ensemble was wrapped with 8xTTA:\n- HorizontalFlip\n- VerticalFip\n- BrightnessUp\n- BrightnessDown\n\nProbably, the assumption of selection bias was correct, as ensemble trained with KD and pseudo labels scored 0.65982 on private LB, which is 0.0034 less than an ensemble without them. \n\n**Hardware**\nWe used 2 servers with 1xV100 and 1 workstation with 2x1080Ti. \nWe thank VITech Lab and 3DLook for providing computing resources for this competition.\n",
      "votes": 13
    },
    {
      "id": 677406,
      "postDate": "2019-11-20T05:34:56.533Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 677406,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-20T05:34:56.533000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "677210": "Hello and congrats to all participants!\nThanks, organizers for this incredible competition!\n\nWe entered this competition after a good rest from Severstal competition.\nActually, I made the repo for it 12 days before the final deadline.\n\nWe used our pipeline from Severstal competition and changed only separate parts of it.\nAt first, I tried the best scoring config from Severstal competition for this data, only hoping for \"easy gold\" but got the rank around ~900.\nThis is my second competition, where the 'fit_predict' could not get into the bronze zone.\n\n**At first, here are the things that we tried, and that did not work at all:**\n\n- Higher than 352x576 resolution,\n- Adding heavy encoders (everything larger than Densenet169/SEResNeXt50 overfitted like crazy)\n- Masking that black stripe with mean value or overlaying with a black rectangle,\n- Training on crops,\n- Separate classifier,\n- Unet decoder,\n- PSP decoder,\n- Hard color augmentations,\n- Decoder from @hengck23 (but was great to learn from!),\n- BCE Loss, Focal Loss,  Dice loss or Lovasz loss (alone),\n- Symmetric CrossEntropy (https://arxiv.org/pdf/1908.06112.pdf)\n\n**Things that worked:**\n@spsancti:\n- Combination of Trimmed BCE and Dice loss,\n- Separate classification head,\n- EfficientNet-B0 and B1 encoders,\n- Pseudolabeling and Knowledge distillation\n\n@hakuryuu:\n- Small networks (Resnet34, EfficientNet-B0...B2),\n- Grid shuffle augmentation,\n- Sum of  symmetric Lovasz, Trimmed Focal, Dice and BCE losses\n\n__We both used:__\n- Cosine Annealing LR schedule,\n- FPN decoder,\n- Hard geometric augmentations,\n- Over9000 optimizer (better than RAdam on all runs)\n- Threshold selection on the validation\n\n__Libraries used:__\n- Catalyst and Apex\n- Albumentations\n- @pavel92 's awesome segmentation models library\n\nWe trained single-fold models, with different split, to get more boost after merging them.\n\n**Pseudolabelling and Knowledge distillation**\n\nOne day before the deadline, I decided to try to use pseudolabelling. \nI took our most successful submit for that time (0.6575 public) and used it as hard pseudo labels. I did not select any confident predictions assuming that the source markings are noisy enough to tolerate some more noise from pseudo labels. \nI pretrained the model on pseudo labels and continued to the regular train. \nImmediately there was an improvement on local validation while training on pseudo labels, which quickly disappeared when I switched to the regular train. So, I trained several models with this scheme only to confirm this result.\n\nThis made me think that noisy corners in original markings confuse models too much, and I decided to relabel training data with the same ensemble that produced this submission. \nWith relabeled training data mixed with pseudo labels, the single model could reach 0.68 on local validation. I submitted it and got the first of my single models to break the 0.66 barrier on public LB. \n\nThese results looked too optimistic, so we thought that I created a selection bias, so we decided to select one ensemble with models trained with pseudo labels and KD and one without.\n\n**Ensembling**\nIn our final submission we had 2 ensembles:\nThe first one was trained without pseudo labels and consisted of 10 models.\nThe second one was the same as first, but with 5 more models trained on KD mixed with pseudo labels. \n\nEvery model in ensemble was wrapped with 8xTTA:\n- HorizontalFlip\n- VerticalFip\n- BrightnessUp\n- BrightnessDown\n\nProbably, the assumption of selection bias was correct, as ensemble trained with KD and pseudo labels scored 0.65982 on private LB, which is 0.0034 less than an ensemble without them. \n\n**Hardware**\nWe used 2 servers with 1xV100 and 1 workstation with 2x1080Ti. \nWe thank VITech Lab and 3DLook for providing computing resources for this competition.\n",
    "677406": ""
  }
}