{
  "id": 226764,
  "title": "76th Place Simple Solution",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/226764",
  "author_name": "Andy",
  "post_date": "2021-03-17T16:37:45.358000",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>First of all, thanks to Kaggle and the RANZCR team for hosting a such an interesting competition. This was a big step for me personally since this was my first image competition. I want to thank <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> for the amazing <a href=\"https://www.kaggle.com/underwearfitting/single-fold-training-of-resnet200d-lb0-965\" target=\"_blank\">notebook</a> which was a huge help for me as a beginner learning PyTorch. </p>\n<h1>1. Preprocessing and Augmentations:</h1>\n<p>My final solution was an ensemble of three <code>resnet200d</code> with slightly different preprocessing methods. Two of them were pretrained on the ChestX dataset and one of them was pretrained on imagenet. <br>\nFor augmentations, I found that two augmentations especially helped to reduce overfitting</p>\n<ul>\n<li>Large dropout. For one of my models, I used <code>max_h_size=int(image_size * 0.3), max_w_size=int(image_size * 0.3), num_holes=1, p=0.5</code>. For the other two, I used <code>max_h_size=int(image_size * 0.1), max_w_size=int(image_size * 0.1), num_holes=4, p=0.5</code></li>\n<li>Motion Blur. It seemed like that motion blur can mimic the fuzziness of some xrays. For two of my models, I used  <code>albumentations.MotionBlur(blur_limit=(7, 15), p=0.5)</code>.</li>\n</ul>\n<p>I found that applying CLAHE to images as a preprocessing method seems to make the xray clearer, I did that for one of my models and compared to ones without, it seems to improve cv by 0.0002~.</p>\n<h1>2. Models and Cross-Validation:</h1>\n<p>Very early on, I experimented with some tf models and none of them were able to reach cv &gt; 95.7 and LB &gt; 95.9. Then I put down the competition for a while before coming back in the final week deciding that PyTorch would be the way to go. </p>\n<p>I used <code>StratifiedGroupKFold</code> for two of my models and <code>StratifiedKFold</code> for one. <code>StratifiedGroupKFold</code> got higher LB usually, but <code>StratifiedKFold</code> had more correlation between CV and LB. </p>\n<ul>\n<li><p>First resnet200d:</p>\n<ul>\n<li>This one was very similar to the one shown in this notebook with some different hyper-parameters. My cv was around 95.4~</li></ul></li>\n<li><p>Second resnet200d:</p>\n<ul>\n<li>This one was pretrained on ChestX and with CLAHE applied to the images before feeding in the model. I used 5 fold <code>StratifiedGroupKFold</code> and the cv was 96.5~. The image size was 575x575</li></ul></li>\n<li><p>Third resnet200d:</p>\n<ul>\n<li>This one was also pretrained on ChestX with a larger droupout compared to the previous two <code>max_h_size=int(image_size * 0.3), max_w_size=int(image_size * 0.3), num_holes=1, p=0.5</code>. I used 5 fold <code>StratifiedKFold</code> and the cv was 96.6~. The image size was 575x575.</li></ul></li>\n</ul>\n<h1>3. PostProcessing (kinda):</h1>\n<p>For inference, I used a larger images size compared to train (100 pixels larger) this lets the model utilize the details of the image. I learned this handy trick from <a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/188299\" target=\"_blank\">here</a>. It improved LB by about 0.0001~ I used simple average between the three models. </p>",
  "messages": [
    {
      "id": 1242490,
      "postDate": "2021-03-17T16:37:45.360Z",
      "content": "<p>First of all, thanks to Kaggle and the RANZCR team for hosting a such an interesting competition. This was a big step for me personally since this was my first image competition. I want to thank <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> for the amazing <a href=\"https://www.kaggle.com/underwearfitting/single-fold-training-of-resnet200d-lb0-965\" target=\"_blank\">notebook</a> which was a huge help for me as a beginner learning PyTorch. </p>\n<h1>1. Preprocessing and Augmentations:</h1>\n<p>My final solution was an ensemble of three <code>resnet200d</code> with slightly different preprocessing methods. Two of them were pretrained on the ChestX dataset and one of them was pretrained on imagenet. <br>\nFor augmentations, I found that two augmentations especially helped to reduce overfitting</p>\n<ul>\n<li>Large dropout. For one of my models, I used <code>max_h_size=int(image_size * 0.3), max_w_size=int(image_size * 0.3), num_holes=1, p=0.5</code>. For the other two, I used <code>max_h_size=int(image_size * 0.1), max_w_size=int(image_size * 0.1), num_holes=4, p=0.5</code></li>\n<li>Motion Blur. It seemed like that motion blur can mimic the fuzziness of some xrays. For two of my models, I used  <code>albumentations.MotionBlur(blur_limit=(7, 15), p=0.5)</code>.</li>\n</ul>\n<p>I found that applying CLAHE to images as a preprocessing method seems to make the xray clearer, I did that for one of my models and compared to ones without, it seems to improve cv by 0.0002~.</p>\n<h1>2. Models and Cross-Validation:</h1>\n<p>Very early on, I experimented with some tf models and none of them were able to reach cv &gt; 95.7 and LB &gt; 95.9. Then I put down the competition for a while before coming back in the final week deciding that PyTorch would be the way to go. </p>\n<p>I used <code>StratifiedGroupKFold</code> for two of my models and <code>StratifiedKFold</code> for one. <code>StratifiedGroupKFold</code> got higher LB usually, but <code>StratifiedKFold</code> had more correlation between CV and LB. </p>\n<ul>\n<li><p>First resnet200d:</p>\n<ul>\n<li>This one was very similar to the one shown in this notebook with some different hyper-parameters. My cv was around 95.4~</li></ul></li>\n<li><p>Second resnet200d:</p>\n<ul>\n<li>This one was pretrained on ChestX and with CLAHE applied to the images before feeding in the model. I used 5 fold <code>StratifiedGroupKFold</code> and the cv was 96.5~. The image size was 575x575</li></ul></li>\n<li><p>Third resnet200d:</p>\n<ul>\n<li>This one was also pretrained on ChestX with a larger droupout compared to the previous two <code>max_h_size=int(image_size * 0.3), max_w_size=int(image_size * 0.3), num_holes=1, p=0.5</code>. I used 5 fold <code>StratifiedKFold</code> and the cv was 96.6~. The image size was 575x575.</li></ul></li>\n</ul>\n<h1>3. PostProcessing (kinda):</h1>\n<p>For inference, I used a larger images size compared to train (100 pixels larger) this lets the model utilize the details of the image. I learned this handy trick from <a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/188299\" target=\"_blank\">here</a>. It improved LB by about 0.0001~ I used simple average between the three models. </p>",
      "rawMarkdown": "First of all, thanks to Kaggle and the RANZCR team for hosting a such an interesting competition. This was a big step for me personally since this was my first image competition. I want to thank @underwearfitting for the amazing [notebook](https://www.kaggle.com/underwearfitting/single-fold-training-of-resnet200d-lb0-965) which was a huge help for me as a beginner learning PyTorch. \n\n\n\n\n# 1. Preprocessing and Augmentations:\n\n\nMy final solution was an ensemble of three `resnet200d` with slightly different preprocessing methods. Two of them were pretrained on the ChestX dataset and one of them was pretrained on imagenet. \nFor augmentations, I found that two augmentations especially helped to reduce overfitting\n- Large dropout. For one of my models, I used `max_h_size=int(image_size * 0.3), max_w_size=int(image_size * 0.3), num_holes=1, p=0.5`. For the other two, I used `max_h_size=int(image_size * 0.1), max_w_size=int(image_size * 0.1), num_holes=4, p=0.5`\n- Motion Blur. It seemed like that motion blur can mimic the fuzziness of some xrays. For two of my models, I used  `   albumentations.MotionBlur(blur_limit=(7, 15), p=0.5)`.\n\nI found that applying CLAHE to images as a preprocessing method seems to make the xray clearer, I did that for one of my models and compared to ones without, it seems to improve cv by 0.0002~.\n\n\n# 2. Models and Cross-Validation:\n\nVery early on, I experimented with some tf models and none of them were able to reach cv > 95.7 and LB > 95.9. Then I put down the competition for a while before coming back in the final week deciding that PyTorch would be the way to go. \n\nI used `StratifiedGroupKFold` for two of my models and `StratifiedKFold` for one. `StratifiedGroupKFold` got higher LB usually, but `StratifiedKFold` had more correlation between CV and LB. \n- First resnet200d:\n       - This one was very similar to the one shown in this notebook with some different hyper-parameters. My cv was around 95.4~\n\n- Second resnet200d:\n       - This one was pretrained on ChestX and with CLAHE applied to the images before feeding in the model. I used 5 fold `StratifiedGroupKFold` and the cv was 96.5~. The image size was 575x575\n\n- Third resnet200d:\n       - This one was also pretrained on ChestX with a larger droupout compared to the previous two `max_h_size=int(image_size * 0.3), max_w_size=int(image_size * 0.3), num_holes=1, p=0.5`. I used 5 fold `StratifiedKFold` and the cv was 96.6~. The image size was 575x575.\n\n# 3. PostProcessing (kinda):\n\nFor inference, I used a larger images size compared to train (100 pixels larger) this lets the model utilize the details of the image. I learned this handy trick from [here](https://www.kaggle.com/c/landmark-recognition-2020/discussion/188299). It improved LB by about 0.0001~ I used simple average between the three models. \n\n\n",
      "votes": 6
    },
    {
      "id": 1243587,
      "postDate": "2021-03-18T10:31:34.903Z",
      "content": "<p><a href=\"https://www.kaggle.com/andy1010\" target=\"_blank\">@andy1010</a> Thanks so much for sharing the solution . Congratulations on Silver Finish</p>",
      "rawMarkdown": "@andy1010 Thanks so much for sharing the solution . Congratulations on Silver Finish",
      "votes": 1
    },
    {
      "id": 1243023,
      "postDate": "2021-03-18T00:54:30.207Z",
      "content": "<p>Thank you for sharing your solutions. Very impressive.</p>",
      "rawMarkdown": "Thank you for sharing your solutions. Very impressive.",
      "votes": 2
    },
    {
      "id": 1246820,
      "postDate": "2021-03-21T06:19:37.437Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1243587,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2021-03-18T10:31:34.903000",
      "content": "<p><a href=\"https://www.kaggle.com/andy1010\" target=\"_blank\">@andy1010</a> Thanks so much for sharing the solution . Congratulations on Silver Finish</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1243023,
      "author_name": "Miyatti",
      "author_url": "",
      "post_date": "2021-03-18T00:54:30.207000",
      "content": "<p>Thank you for sharing your solutions. Very impressive.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1246820,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-21T06:19:37.437000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1242490": "First of all, thanks to Kaggle and the RANZCR team for hosting a such an interesting competition. This was a big step for me personally since this was my first image competition. I want to thank @underwearfitting for the amazing [notebook](https://www.kaggle.com/underwearfitting/single-fold-training-of-resnet200d-lb0-965) which was a huge help for me as a beginner learning PyTorch. \n\n\n\n\n# 1. Preprocessing and Augmentations:\n\n\nMy final solution was an ensemble of three `resnet200d` with slightly different preprocessing methods. Two of them were pretrained on the ChestX dataset and one of them was pretrained on imagenet. \nFor augmentations, I found that two augmentations especially helped to reduce overfitting\n- Large dropout. For one of my models, I used `max_h_size=int(image_size * 0.3), max_w_size=int(image_size * 0.3), num_holes=1, p=0.5`. For the other two, I used `max_h_size=int(image_size * 0.1), max_w_size=int(image_size * 0.1), num_holes=4, p=0.5`\n- Motion Blur. It seemed like that motion blur can mimic the fuzziness of some xrays. For two of my models, I used  `   albumentations.MotionBlur(blur_limit=(7, 15), p=0.5)`.\n\nI found that applying CLAHE to images as a preprocessing method seems to make the xray clearer, I did that for one of my models and compared to ones without, it seems to improve cv by 0.0002~.\n\n\n# 2. Models and Cross-Validation:\n\nVery early on, I experimented with some tf models and none of them were able to reach cv > 95.7 and LB > 95.9. Then I put down the competition for a while before coming back in the final week deciding that PyTorch would be the way to go. \n\nI used `StratifiedGroupKFold` for two of my models and `StratifiedKFold` for one. `StratifiedGroupKFold` got higher LB usually, but `StratifiedKFold` had more correlation between CV and LB. \n- First resnet200d:\n       - This one was very similar to the one shown in this notebook with some different hyper-parameters. My cv was around 95.4~\n\n- Second resnet200d:\n       - This one was pretrained on ChestX and with CLAHE applied to the images before feeding in the model. I used 5 fold `StratifiedGroupKFold` and the cv was 96.5~. The image size was 575x575\n\n- Third resnet200d:\n       - This one was also pretrained on ChestX with a larger droupout compared to the previous two `max_h_size=int(image_size * 0.3), max_w_size=int(image_size * 0.3), num_holes=1, p=0.5`. I used 5 fold `StratifiedKFold` and the cv was 96.6~. The image size was 575x575.\n\n# 3. PostProcessing (kinda):\n\nFor inference, I used a larger images size compared to train (100 pixels larger) this lets the model utilize the details of the image. I learned this handy trick from [here](https://www.kaggle.com/c/landmark-recognition-2020/discussion/188299). It improved LB by about 0.0001~ I used simple average between the three models. \n\n\n",
    "1243587": "@andy1010 Thanks so much for sharing the solution . Congratulations on Silver Finish",
    "1243023": "Thank you for sharing your solutions. Very impressive.",
    "1246820": ""
  }
}