{
  "id": 118060,
  "title": "86-th(bronze) writeup - solution description and lessons learned.",
  "url": "/competitions/understanding_cloud_organization/discussion/118060",
  "author_name": "Victor Zaguskin",
  "post_date": "2019-11-19T11:49:41.114000",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p><strong>solution summary</strong>\nMy final submit is a voting average of 5 best submissions by public LB. Each of this 5 submissions is a 5-fold voting average of an efficientnet-b4 Unet or FPN, \ntrained with BCE Dice/only bce loss with or without tta on image sizes 512x352, 640x320. \nVoting average settings - minimum 4 of 5 nonempty masks to consider mask non-empty, minimum 2 positive pixels to consider a pixel positive for non-empty mask.</p>\n\n<p><strong>Pipeline</strong>\nThe code is based on this great <a href=\"https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools\">kernel </a> with few important changes that allowed to get single model \npublic score 0.66+:\n- efficientnet-b4 backbone\n- border mode cv2.BORDER_REFLECT_101(default) in albumentations ShiftScaleRotate\n- Radam optimizer\n- Customer learning rate scheduler(decoder starts at 1e-3 and decleaning to 1e-4 in 10 epochs, encoder starts at 0 and increasing to 1e-4 in 10 epochs, then both\ndecleaning in steps to 1e-5 for next 30 epochs everu 3 epochs). Usually the best checkpoint is at epoch ~22.</p>\n\n<p>Side note: I switched from keras to pytorch in this competition and quite happy with that. The most important reasons are:\n1. Mixed precision training that can be enabled with 1 line of code in catalyst\n2. Parameter groups in optimizer that allow fine control over model learning.</p>\n\n<p><strong>Here are some highlights of the lessons learned.</strong>\n1. Even in competition like this where CV/LB discrepancy is small, single model score improvement after a hyperparameter change doesn't prove anything. N-fold should be used for validation every time.\n2. Too much parameter fitting on validation set(threshold, min_size) easily leads to overfitting to validation. Using constant values eventually is better.\n3. Models with tta inference don't improve public LB most of the time, but consistently better on private.</p>\n\n<p><strong>Things that didn't work for me:</strong>\n1. Pseudo labeling\n2. Loss functions beyond BCEDice/BCE.\n3. Classification(stopped adding value after LB reached 0.66)</p>\n\n<p><strong>What things I wish I've tried:</strong>\n1. Implement final thresholded dice as a metric and use it for checkpointing\n2. triplet thresholding(<a href=\"https://www.kaggle.com/c/siim-acr-pneumothorax-segmentation/discussion/107824\">https://www.kaggle.com/c/siim-acr-pneumothorax-segmentation/discussion/107824</a>), or double flat thresholds like Heng did.\n3. implement good k-fold validation scheme early and make a grid search for image size/loss/augmentations/tta\n4. pretraining - <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/118017#latest-676633\">https://www.kaggle.com/c/understanding_cloud_organization/discussion/118017#latest-676633</a></p>\n\n<p><strong>Conclusion</strong>\nI got my first kaggle medal, was quite close to silver zone and didn't suffer a major shakeup( actually enjoyed it with getting +16 positions) so result is quite positive for me. But still have to learn a lot.</p>",
  "messages": [
    {
      "id": 676654,
      "postDate": "2019-11-19T11:49:41.113Z",
      "content": "<p><strong>solution summary</strong>\nMy final submit is a voting average of 5 best submissions by public LB. Each of this 5 submissions is a 5-fold voting average of an efficientnet-b4 Unet or FPN, \ntrained with BCE Dice/only bce loss with or without tta on image sizes 512x352, 640x320. \nVoting average settings - minimum 4 of 5 nonempty masks to consider mask non-empty, minimum 2 positive pixels to consider a pixel positive for non-empty mask.</p>\n\n<p><strong>Pipeline</strong>\nThe code is based on this great <a href=\"https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools\">kernel </a> with few important changes that allowed to get single model \npublic score 0.66+:\n- efficientnet-b4 backbone\n- border mode cv2.BORDER_REFLECT_101(default) in albumentations ShiftScaleRotate\n- Radam optimizer\n- Customer learning rate scheduler(decoder starts at 1e-3 and decleaning to 1e-4 in 10 epochs, encoder starts at 0 and increasing to 1e-4 in 10 epochs, then both\ndecleaning in steps to 1e-5 for next 30 epochs everu 3 epochs). Usually the best checkpoint is at epoch ~22.</p>\n\n<p>Side note: I switched from keras to pytorch in this competition and quite happy with that. The most important reasons are:\n1. Mixed precision training that can be enabled with 1 line of code in catalyst\n2. Parameter groups in optimizer that allow fine control over model learning.</p>\n\n<p><strong>Here are some highlights of the lessons learned.</strong>\n1. Even in competition like this where CV/LB discrepancy is small, single model score improvement after a hyperparameter change doesn't prove anything. N-fold should be used for validation every time.\n2. Too much parameter fitting on validation set(threshold, min_size) easily leads to overfitting to validation. Using constant values eventually is better.\n3. Models with tta inference don't improve public LB most of the time, but consistently better on private.</p>\n\n<p><strong>Things that didn't work for me:</strong>\n1. Pseudo labeling\n2. Loss functions beyond BCEDice/BCE.\n3. Classification(stopped adding value after LB reached 0.66)</p>\n\n<p><strong>What things I wish I've tried:</strong>\n1. Implement final thresholded dice as a metric and use it for checkpointing\n2. triplet thresholding(<a href=\"https://www.kaggle.com/c/siim-acr-pneumothorax-segmentation/discussion/107824\">https://www.kaggle.com/c/siim-acr-pneumothorax-segmentation/discussion/107824</a>), or double flat thresholds like Heng did.\n3. implement good k-fold validation scheme early and make a grid search for image size/loss/augmentations/tta\n4. pretraining - <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/118017#latest-676633\">https://www.kaggle.com/c/understanding_cloud_organization/discussion/118017#latest-676633</a></p>\n\n<p><strong>Conclusion</strong>\nI got my first kaggle medal, was quite close to silver zone and didn't suffer a major shakeup( actually enjoyed it with getting +16 positions) so result is quite positive for me. But still have to learn a lot.</p>",
      "rawMarkdown": "**solution summary**\nMy final submit is a voting average of 5 best submissions by public LB. Each of this 5 submissions is a 5-fold voting average of an efficientnet-b4 Unet or FPN, \ntrained with BCE Dice/only bce loss with or without tta on image sizes 512x352, 640x320. \nVoting average settings - minimum 4 of 5 nonempty masks to consider mask non-empty, minimum 2 positive pixels to consider a pixel positive for non-empty mask.\n\n**Pipeline**\nThe code is based on this great [kernel ](https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools) with few important changes that allowed to get single model \npublic score 0.66+:\n- efficientnet-b4 backbone\n- border mode cv2.BORDER_REFLECT_101(default) in albumentations ShiftScaleRotate\n- Radam optimizer\n- Customer learning rate scheduler(decoder starts at 1e-3 and decleaning to 1e-4 in 10 epochs, encoder starts at 0 and increasing to 1e-4 in 10 epochs, then both\ndecleaning in steps to 1e-5 for next 30 epochs everu 3 epochs). Usually the best checkpoint is at epoch ~22.\n\nSide note: I switched from keras to pytorch in this competition and quite happy with that. The most important reasons are:\n1. Mixed precision training that can be enabled with 1 line of code in catalyst\n2. Parameter groups in optimizer that allow fine control over model learning.\n\n**Here are some highlights of the lessons learned.**\n1. Even in competition like this where CV/LB discrepancy is small, single model score improvement after a hyperparameter change doesn't prove anything. N-fold should be used for validation every time.\n2. Too much parameter fitting on validation set(threshold, min_size) easily leads to overfitting to validation. Using constant values eventually is better.\n3. Models with tta inference don't improve public LB most of the time, but consistently better on private.\n\n**Things that didn't work for me:**\n1. Pseudo labeling\n2. Loss functions beyond BCEDice/BCE.\n3. Classification(stopped adding value after LB reached 0.66)\n\n**What things I wish I've tried:**\n1. Implement final thresholded dice as a metric and use it for checkpointing\n2. triplet thresholding(https://www.kaggle.com/c/siim-acr-pneumothorax-segmentation/discussion/107824), or double flat thresholds like Heng did.\n3. implement good k-fold validation scheme early and make a grid search for image size/loss/augmentations/tta\n4. pretraining - https://www.kaggle.com/c/understanding_cloud_organization/discussion/118017#latest-676633\n\n**Conclusion**\nI got my first kaggle medal, was quite close to silver zone and didn't suffer a major shakeup( actually enjoyed it with getting +16 positions) so result is quite positive for me. But still have to learn a lot.",
      "votes": 5
    },
    {
      "id": 677082,
      "postDate": "2019-11-19T19:15:28.680Z",
      "content": "<p>Congratulations Victor ! </p>\n\n<p>When you say \"Each of this 5 submissions is a 5-fold voting average\", in the folding you made 5 folds, each of them containing 20% of the data, then made all combinations for training on 4 of them(80% of data) and validate on 1(20% of data). Then after training 5 segmentation network in this scheme, you have take the mean of 5 predictions, right ?</p>",
      "rawMarkdown": "Congratulations Victor ! \n\nWhen you say \"Each of this 5 submissions is a 5-fold voting average\", in the folding you made 5 folds, each of them containing 20% of the data, then made all combinations for training on 4 of them(80% of data) and validate on 1(20% of data). Then after training 5 segmentation network in this scheme, you have take the mean of 5 predictions, right ?",
      "replies": [
        {
          "id": 677089,
          "postDate": "2019-11-19T19:31:14.837Z",
          "content": "<p>Thanks, Vlad.\nYes, almost correct. I train 5 models as you describe, validate and create 5 submissions.\nThen I have to blend 5 binary masks for each image/label. This is done the following way:\n1. If less then 4 masks are non-empty(2 or more empty) - I submit an empty mask.\n2. If I have 4+ non-empty masks - then I set each pixel to 1 if it is 1 in two or more masks. </p>\n\n<p>This way performed consistently better then prediction averaging for me.</p>",
          "rawMarkdown": "Thanks, Vlad.\nYes, almost correct. I train 5 models as you describe, validate and create 5 submissions.\nThen I have to blend 5 binary masks for each image/label. This is done the following way:\n1. If less then 4 masks are non-empty(2 or more empty) - I submit an empty mask.\n2. If I have 4+ non-empty masks - then I set each pixel to 1 if it is 1 in two or more masks. \n\nThis way performed consistently better then prediction averaging for me.",
          "votes": 1
        },
        {
          "id": 677115,
          "postDate": "2019-11-19T20:16:49.260Z",
          "content": "<p>Interesting approach.\nOne more thing, on the 5-fold voting average did you start each train from scratch or from the best checkpoint of the previous fold? </p>",
          "rawMarkdown": "Interesting approach.\nOne more thing, on the 5-fold voting average did you start each train from scratch or from the best checkpoint of the previous fold? "
        },
        {
          "id": 677162,
          "postDate": "2019-11-19T21:40:44.130Z",
          "content": "<p>From scratch. \nUsing checkpoint from previous fold would introduce a leak and break validation.</p>",
          "rawMarkdown": "From scratch. \nUsing checkpoint from previous fold would introduce a leak and break validation.",
          "votes": 1
        },
        {
          "id": 677167,
          "postDate": "2019-11-19T21:45:00.033Z",
          "content": "<p>Yes, this is what I was thinking also.</p>",
          "rawMarkdown": "Yes, this is what I was thinking also."
        }
      ]
    },
    {
      "id": 676853,
      "postDate": "2019-11-19T14:47:55.207Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 677082,
      "author_name": "Vlad Vaduva",
      "author_url": "",
      "post_date": "2019-11-19T19:15:28.680000",
      "content": "<p>Congratulations Victor ! </p>\n\n<p>When you say \"Each of this 5 submissions is a 5-fold voting average\", in the folding you made 5 folds, each of them containing 20% of the data, then made all combinations for training on 4 of them(80% of data) and validate on 1(20% of data). Then after training 5 segmentation network in this scheme, you have take the mean of 5 predictions, right ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 677089,
          "author_name": "Victor Zaguskin",
          "author_url": "",
          "post_date": "2019-11-19T19:31:14.837000",
          "content": "<p>Thanks, Vlad.\nYes, almost correct. I train 5 models as you describe, validate and create 5 submissions.\nThen I have to blend 5 binary masks for each image/label. This is done the following way:\n1. If less then 4 masks are non-empty(2 or more empty) - I submit an empty mask.\n2. If I have 4+ non-empty masks - then I set each pixel to 1 if it is 1 in two or more masks. </p>\n\n<p>This way performed consistently better then prediction averaging for me.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 677115,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2019-11-19T20:16:49.260000",
          "content": "<p>Interesting approach.\nOne more thing, on the 5-fold voting average did you start each train from scratch or from the best checkpoint of the previous fold? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 677162,
          "author_name": "Victor Zaguskin",
          "author_url": "",
          "post_date": "2019-11-19T21:40:44.130000",
          "content": "<p>From scratch. \nUsing checkpoint from previous fold would introduce a leak and break validation.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 677167,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2019-11-19T21:45:00.033000",
          "content": "<p>Yes, this is what I was thinking also.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 676853,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-19T14:47:55.207000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "676654": "**solution summary**\nMy final submit is a voting average of 5 best submissions by public LB. Each of this 5 submissions is a 5-fold voting average of an efficientnet-b4 Unet or FPN, \ntrained with BCE Dice/only bce loss with or without tta on image sizes 512x352, 640x320. \nVoting average settings - minimum 4 of 5 nonempty masks to consider mask non-empty, minimum 2 positive pixels to consider a pixel positive for non-empty mask.\n\n**Pipeline**\nThe code is based on this great [kernel ](https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools) with few important changes that allowed to get single model \npublic score 0.66+:\n- efficientnet-b4 backbone\n- border mode cv2.BORDER_REFLECT_101(default) in albumentations ShiftScaleRotate\n- Radam optimizer\n- Customer learning rate scheduler(decoder starts at 1e-3 and decleaning to 1e-4 in 10 epochs, encoder starts at 0 and increasing to 1e-4 in 10 epochs, then both\ndecleaning in steps to 1e-5 for next 30 epochs everu 3 epochs). Usually the best checkpoint is at epoch ~22.\n\nSide note: I switched from keras to pytorch in this competition and quite happy with that. The most important reasons are:\n1. Mixed precision training that can be enabled with 1 line of code in catalyst\n2. Parameter groups in optimizer that allow fine control over model learning.\n\n**Here are some highlights of the lessons learned.**\n1. Even in competition like this where CV/LB discrepancy is small, single model score improvement after a hyperparameter change doesn't prove anything. N-fold should be used for validation every time.\n2. Too much parameter fitting on validation set(threshold, min_size) easily leads to overfitting to validation. Using constant values eventually is better.\n3. Models with tta inference don't improve public LB most of the time, but consistently better on private.\n\n**Things that didn't work for me:**\n1. Pseudo labeling\n2. Loss functions beyond BCEDice/BCE.\n3. Classification(stopped adding value after LB reached 0.66)\n\n**What things I wish I've tried:**\n1. Implement final thresholded dice as a metric and use it for checkpointing\n2. triplet thresholding(https://www.kaggle.com/c/siim-acr-pneumothorax-segmentation/discussion/107824), or double flat thresholds like Heng did.\n3. implement good k-fold validation scheme early and make a grid search for image size/loss/augmentations/tta\n4. pretraining - https://www.kaggle.com/c/understanding_cloud_organization/discussion/118017#latest-676633\n\n**Conclusion**\nI got my first kaggle medal, was quite close to silver zone and didn't suffer a major shakeup( actually enjoyed it with getting +16 positions) so result is quite positive for me. But still have to learn a lot.",
    "677082": "Congratulations Victor ! \n\nWhen you say \"Each of this 5 submissions is a 5-fold voting average\", in the folding you made 5 folds, each of them containing 20% of the data, then made all combinations for training on 4 of them(80% of data) and validate on 1(20% of data). Then after training 5 segmentation network in this scheme, you have take the mean of 5 predictions, right ?",
    "676853": ""
  }
}