{
  "id": 118325,
  "title": "240 place with simple model, no kfold, no combining networks",
  "url": "/competitions/understanding_cloud_organization/writeups/vlad-vaduva-240-place-with-simple-model-no-kfold-n",
  "author_name": "",
  "post_date": "2019-11-20T18:45:44.270415300Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi guys, </p>\n\n<p>For this competition I created and tested single models without any kfolding or combining multiple networks architectures.\nThe parameters which I tested were:</p>\n\n<p><strong>Convolutional Networks Arhitectures:</strong>\n- Unet \n- FPN</p>\n\n<p><strong>Pretrained Networks:</strong>\n- resnet18\n- resnet34\n- resnet50\n- resnet101\n- resnet152\n- se _ resnext50_32x4d\n- se _ resnext101_32x4d\n- efficientnet-b0 \n- efficientnet-b1\n- efficientnet-b2\n- efficientnet-b7</p>\n\n<p><strong>Batch sizes</strong>\n- Starting from 1 to 9\n- Also used accumulation_steps(steps=2 and 3)</p>\n\n<p><strong>Preprocessing</strong>\n- Resize to (640, 320) for segmentation input data. Then resized the masks to (525, 350)\n- HorizontalFlip(p=0.25),\n- VerticalFlip(p=0.25),\n- ShiftScaleRotate(scale_limit=0.5, rotate_limit=0, shift_limit=0.1, p=0.5, border_mode=0),\n- GridDistortion(p=0.25)</p>\n\n<p><strong>Optimizers</strong>\n- Adam\n- RAdam\n- SGD</p>\n\n<p><strong>Losses</strong>\n- BCEDiceLoss\n- IoULoss\n- FocalLossBinary\n- Custom loss(BCEDiceLoss*0.4 + IoULoss*0.2 + FocalLossBinary*0.4)</p>\n\n<p><strong>Post processing</strong>\n- Finding optimum threshold between in the interval 0.3-1 with a 0.005 step for each category\n- Finding minimum pixels value for considering the mask a non-Zero one (9000-25000) with a 1000 step for each category</p>\n\n<p><strong>BEST MODEL FOUND</strong>\nBest model obtain 0.65803 on public score and <strong>0.65038 on private leaderboard</strong></p>\n\n<p>The configuration for this single model without any k-folds or combination with another arhitecture  was:</p>\n\n<p><strong>FPN</strong>+\n<strong>se _ resnext101_32x4d</strong>+\n<strong>batch size 6(accumulate gradient=2)</strong>+\n<strong>RAdam</strong>+\n<strong>BCEDiceLoss</strong>+</p>\n\n<p><strong>Preprocessing</strong>: \nResize to (640, 320) for segmentation input data. Then resized the masks to (525, 350)\nHorizontalFlip(p=0.25)+\nVerticalFlip(p=0.25)+\nShiftScaleRotate(scale_limit=0.5, rotate_limit=0, shift_limit=0.1, p=0.5, border_mode=0)+\nGridDistortion(p=0.25)</p>\n\n<p><strong>Post processing thresholds</strong>(cat1: thres=0.335, min_pixels=21000, cat2: thres=0.605, min_pixels=15000, cat3:thres=0.640, min_pixels=20000 , cat4: thres=0.565, min_pixels=16000)</p>\n\n<p><strong>Other things that I wish I had tried but not had time:</strong>  </p>\n\n<ul>\n<li>Instead of resizing initially to (640, 320) as input for segmentation network and then resizing the resulting masks to (525, 350) I would had iniatially resized to (525, 350) and then use padding for creating a (640, 320) imagine. And for submiting the mask, I would had remove the pixels offsets from the padding in that way avoiding loss due to resizing results from (640, 320) to (525, 350) </li>\n<li>Insist with more custom weights on my custom loss (BCEDiceLoss*0.4 + IoULoss*0.2 + FocalLossBinary*0.4)</li>\n<li>use AdamW (<a href=\"https://towardsdatascience.com/why-adamw-matters-736223f31b5d\">https://towardsdatascience.com/why-adamw-matters-736223f31b5d</a>)</li>\n<li>TTA</li>\n<li>Mixed precision training (to see how much increased batch size will help compared with accumulate gradient methodology) and also evaluate fp16</li>\n<li>Use and evaluate lovasz loss</li>\n</ul>",
  "messages": [
    {
      "id": "677910",
      "postDate": "11/20/2019 18:45:44",
      "content": "<p>Hi guys, </p>\n\n<p>For this competition I created and tested single models without any kfolding or combining multiple networks architectures.\nThe parameters which I tested were:</p>\n\n<p><strong>Convolutional Networks Arhitectures:</strong>\n- Unet \n- FPN</p>\n\n<p><strong>Pretrained Networks:</strong>\n- resnet18\n- resnet34\n- resnet50\n- resnet101\n- resnet152\n- se _ resnext50_32x4d\n- se _ resnext101_32x4d\n- efficientnet-b0 \n- efficientnet-b1\n- efficientnet-b2\n- efficientnet-b7</p>\n\n<p><strong>Batch sizes</strong>\n- Starting from 1 to 9\n- Also used accumulation_steps(steps=2 and 3)</p>\n\n<p><strong>Preprocessing</strong>\n- Resize to (640, 320) for segmentation input data. Then resized the masks to (525, 350)\n- HorizontalFlip(p=0.25),\n- VerticalFlip(p=0.25),\n- ShiftScaleRotate(scale_limit=0.5, rotate_limit=0, shift_limit=0.1, p=0.5, border_mode=0),\n- GridDistortion(p=0.25)</p>\n\n<p><strong>Optimizers</strong>\n- Adam\n- RAdam\n- SGD</p>\n\n<p><strong>Losses</strong>\n- BCEDiceLoss\n- IoULoss\n- FocalLossBinary\n- Custom loss(BCEDiceLoss*0.4 + IoULoss*0.2 + FocalLossBinary*0.4)</p>\n\n<p><strong>Post processing</strong>\n- Finding optimum threshold between in the interval 0.3-1 with a 0.005 step for each category\n- Finding minimum pixels value for considering the mask a non-Zero one (9000-25000) with a 1000 step for each category</p>\n\n<p><strong>BEST MODEL FOUND</strong>\nBest model obtain 0.65803 on public score and <strong>0.65038 on private leaderboard</strong></p>\n\n<p>The configuration for this single model without any k-folds or combination with another arhitecture  was:</p>\n\n<p><strong>FPN</strong>+\n<strong>se _ resnext101_32x4d</strong>+\n<strong>batch size 6(accumulate gradient=2)</strong>+\n<strong>RAdam</strong>+\n<strong>BCEDiceLoss</strong>+</p>\n\n<p><strong>Preprocessing</strong>: \nResize to (640, 320) for segmentation input data. Then resized the masks to (525, 350)\nHorizontalFlip(p=0.25)+\nVerticalFlip(p=0.25)+\nShiftScaleRotate(scale_limit=0.5, rotate_limit=0, shift_limit=0.1, p=0.5, border_mode=0)+\nGridDistortion(p=0.25)</p>\n\n<p><strong>Post processing thresholds</strong>(cat1: thres=0.335, min_pixels=21000, cat2: thres=0.605, min_pixels=15000, cat3:thres=0.640, min_pixels=20000 , cat4: thres=0.565, min_pixels=16000)</p>\n\n<p><strong>Other things that I wish I had tried but not had time:</strong>  </p>\n\n<ul>\n<li>Instead of resizing initially to (640, 320) as input for segmentation network and then resizing the resulting masks to (525, 350) I would had iniatially resized to (525, 350) and then use padding for creating a (640, 320) imagine. And for submiting the mask, I would had remove the pixels offsets from the padding in that way avoiding loss due to resizing results from (640, 320) to (525, 350) </li>\n<li>Insist with more custom weights on my custom loss (BCEDiceLoss*0.4 + IoULoss*0.2 + FocalLossBinary*0.4)</li>\n<li>use AdamW (<a href=\"https://towardsdatascience.com/why-adamw-matters-736223f31b5d\">https://towardsdatascience.com/why-adamw-matters-736223f31b5d</a>)</li>\n<li>TTA</li>\n<li>Mixed precision training (to see how much increased batch size will help compared with accumulate gradient methodology) and also evaluate fp16</li>\n<li>Use and evaluate lovasz loss</li>\n</ul>",
      "rawMarkdown": "Hi guys, \n\nFor this competition I created and tested single models without any kfolding or combining multiple networks architectures.\nThe parameters which I tested were:\n\n**Convolutional Networks Arhitectures:**\n- Unet \n- FPN\n\n**Pretrained Networks:**\n- resnet18\n- resnet34\n- resnet50\n- resnet101\n- resnet152\n- se _ resnext50_32x4d\n- se _ resnext101_32x4d\n- efficientnet-b0 \n- efficientnet-b1\n- efficientnet-b2\n- efficientnet-b7\n\n**Batch sizes**\n- Starting from 1 to 9\n- Also used accumulation_steps(steps=2 and 3)\n\n**Preprocessing**\n- Resize to (640, 320) for segmentation input data. Then resized the masks to (525, 350)\n- HorizontalFlip(p=0.25),\n- VerticalFlip(p=0.25),\n- ShiftScaleRotate(scale_limit=0.5, rotate_limit=0, shift_limit=0.1, p=0.5, border_mode=0),\n- GridDistortion(p=0.25)\n\n\n**Optimizers**\n- Adam\n- RAdam\n- SGD\n\n**Losses**\n- BCEDiceLoss\n- IoULoss\n- FocalLossBinary\n- Custom loss(BCEDiceLoss*0.4 + IoULoss*0.2 + FocalLossBinary*0.4)\n\n**Post processing**\n- Finding optimum threshold between in the interval 0.3-1 with a 0.005 step for each category\n- Finding minimum pixels value for considering the mask a non-Zero one (9000-25000) with a 1000 step for each category\n\n**BEST MODEL FOUND**\nBest model obtain 0.65803 on public score and **0.65038 on private leaderboard**\n\nThe configuration for this single model without any k-folds or combination with another arhitecture  was:\n\n**FPN**+\n**se _ resnext101_32x4d**+\n**batch size 6(accumulate gradient=2)**+\n**RAdam**+\n**BCEDiceLoss**+\n\n**Preprocessing**: \nResize to (640, 320) for segmentation input data. Then resized the masks to (525, 350)\nHorizontalFlip(p=0.25)+\nVerticalFlip(p=0.25)+\nShiftScaleRotate(scale_limit=0.5, rotate_limit=0, shift_limit=0.1, p=0.5, border_mode=0)+\nGridDistortion(p=0.25)\n\n**Post processing thresholds**(cat1: thres=0.335, min_pixels=21000, cat2: thres=0.605, min_pixels=15000, cat3:thres=0.640, min_pixels=20000 , cat4: thres=0.565, min_pixels=16000)\n  \n\n**Other things that I wish I had tried but not had time:**  \n\n\n- Instead of resizing initially to (640, 320) as input for segmentation network and then resizing the resulting masks to (525, 350) I would had iniatially resized to (525, 350) and then use padding for creating a (640, 320) imagine. And for submiting the mask, I would had remove the pixels offsets from the padding in that way avoiding loss due to resizing results from (640, 320) to (525, 350) \n- Insist with more custom weights on my custom loss (BCEDiceLoss*0.4 + IoULoss*0.2 + FocalLossBinary*0.4)\n- use AdamW (https://towardsdatascience.com/why-adamw-matters-736223f31b5d)\n- TTA\n- Mixed precision training (to see how much increased batch size will help compared with accumulate gradient methodology) and also evaluate fp16\n-  Use and evaluate lovasz loss",
      "votes": null
    },
    {
      "id": "677937",
      "postDate": "11/20/2019 19:47:45",
      "content": "<p>Great job Vlad. That's a great model. I like how you systematically searched to find it. If you add TTA, K-Fold, self ensemble different seeds, and classifier (to remove false positive) that would easily finish in top 50.</p>",
      "rawMarkdown": "Great job Vlad. That's a great model. I like how you systematically searched to find it. If you add TTA, K-Fold, self ensemble different seeds, and classifier (to remove false positive) that would easily finish in top 50.",
      "votes": null
    },
    {
      "id": "677979",
      "postDate": "11/20/2019 20:59:21",
      "content": "<p>Thank you Chris !</p>",
      "rawMarkdown": "Thank you Chris !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 677937,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "11/20/2019 19:47:45",
      "content": "<p>Great job Vlad. That's a great model. I like how you systematically searched to find it. If you add TTA, K-Fold, self ensemble different seeds, and classifier (to remove false positive) that would easily finish in top 50.</p>",
      "votes": null,
      "replies": [
        {
          "id": 677979,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "11/20/2019 20:59:21",
          "content": "<p>Thank you Chris !</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "677910": "Hi guys, \n\nFor this competition I created and tested single models without any kfolding or combining multiple networks architectures.\nThe parameters which I tested were:\n\n**Convolutional Networks Arhitectures:**\n- Unet \n- FPN\n\n**Pretrained Networks:**\n- resnet18\n- resnet34\n- resnet50\n- resnet101\n- resnet152\n- se _ resnext50_32x4d\n- se _ resnext101_32x4d\n- efficientnet-b0 \n- efficientnet-b1\n- efficientnet-b2\n- efficientnet-b7\n\n**Batch sizes**\n- Starting from 1 to 9\n- Also used accumulation_steps(steps=2 and 3)\n\n**Preprocessing**\n- Resize to (640, 320) for segmentation input data. Then resized the masks to (525, 350)\n- HorizontalFlip(p=0.25),\n- VerticalFlip(p=0.25),\n- ShiftScaleRotate(scale_limit=0.5, rotate_limit=0, shift_limit=0.1, p=0.5, border_mode=0),\n- GridDistortion(p=0.25)\n\n\n**Optimizers**\n- Adam\n- RAdam\n- SGD\n\n**Losses**\n- BCEDiceLoss\n- IoULoss\n- FocalLossBinary\n- Custom loss(BCEDiceLoss*0.4 + IoULoss*0.2 + FocalLossBinary*0.4)\n\n**Post processing**\n- Finding optimum threshold between in the interval 0.3-1 with a 0.005 step for each category\n- Finding minimum pixels value for considering the mask a non-Zero one (9000-25000) with a 1000 step for each category\n\n**BEST MODEL FOUND**\nBest model obtain 0.65803 on public score and **0.65038 on private leaderboard**\n\nThe configuration for this single model without any k-folds or combination with another arhitecture  was:\n\n**FPN**+\n**se _ resnext101_32x4d**+\n**batch size 6(accumulate gradient=2)**+\n**RAdam**+\n**BCEDiceLoss**+\n\n**Preprocessing**: \nResize to (640, 320) for segmentation input data. Then resized the masks to (525, 350)\nHorizontalFlip(p=0.25)+\nVerticalFlip(p=0.25)+\nShiftScaleRotate(scale_limit=0.5, rotate_limit=0, shift_limit=0.1, p=0.5, border_mode=0)+\nGridDistortion(p=0.25)\n\n**Post processing thresholds**(cat1: thres=0.335, min_pixels=21000, cat2: thres=0.605, min_pixels=15000, cat3:thres=0.640, min_pixels=20000 , cat4: thres=0.565, min_pixels=16000)\n  \n\n**Other things that I wish I had tried but not had time:**  \n\n\n- Instead of resizing initially to (640, 320) as input for segmentation network and then resizing the resulting masks to (525, 350) I would had iniatially resized to (525, 350) and then use padding for creating a (640, 320) imagine. And for submiting the mask, I would had remove the pixels offsets from the padding in that way avoiding loss due to resizing results from (640, 320) to (525, 350) \n- Insist with more custom weights on my custom loss (BCEDiceLoss*0.4 + IoULoss*0.2 + FocalLossBinary*0.4)\n- use AdamW (https://towardsdatascience.com/why-adamw-matters-736223f31b5d)\n- TTA\n- Mixed precision training (to see how much increased batch size will help compared with accumulate gradient methodology) and also evaluate fp16\n-  Use and evaluate lovasz loss",
    "677937": "Great job Vlad. That's a great model. I like how you systematically searched to find it. If you add TTA, K-Fold, self ensemble different seeds, and classifier (to remove false positive) that would easily finish in top 50.",
    "677979": "Thank you Chris !"
  },
  "source": "meta"
}