{
  "id": 70505,
  "title": "18th solution: SENet-DeepLabV3+",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/70505",
  "author_name": "OsciiArt",
  "post_date": "2018-11-04T17:41:29.882000",
  "votes": 41,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Here I would like to share my solution.</p>\n\n<p>My score is Stage 1 Public LB 0.211 (42nd)/Stage 2 Private LB 0.218 (18th)</p>\n\n<h3>Models</h3>\n\n<p>I tackle this competition with semantic segmentation approach because,  </p>\n\n<ul>\n<li>There are a few objects and each object is well separated, so splitting objects from segmentation mask must be easy.</li>\n<li>Shapes of opacity areas are ambiguous, so rough masks generated from bounding boxes are not so unnatural.</li>\n<li>Training a segmentation model is easier than an object detection model, I believe.</li>\n<li>A segmentation model suits for ensemble compared to an object detection model.</li>\n</ul>\n\n<p>I choose DeepLabV3+ model, the state of the art for semantic segmentation. I take the DeepLabV3 implementation from <br>\n<a href=\"https://github.com/jfzhang95/pytorch-deeplab-xception\">https://github.com/jfzhang95/pytorch-deeplab-xception</a> <br>\nI change the model head from Xception to SENet or SE-ResNext101, because Xception head is not trained stably somehow. I take the SENet and SE-ResNext101 implementation from <br>\n<a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a>  </p>\n\n<p>I train 3 models.</p>\n\n<ol>\n<li>SE-ResNext101-DeepLabV3+ (resized input)</li>\n<li>SE-ResNext101-DeepLabV3+ (cropped input)</li>\n<li>SENet-DeepLabV3+ (resized input)</li>\n</ol>\n\n<p>For resized-input models, input images are resized into 448x448. For cropped-input models, input images are cropped into 448x448 in training and original size image is used in prediction. Using 2 input patterns, I hope global and local features are learned.</p>\n\n<h3>Preprocessing</h3>\n\n<p>I use ellipses inscribed to bounding boxes as true masks. rounded masks are more natural than rectangles and applicable for rotate augmentation.</p>\n\n<h3>Training</h3>\n\n<ul>\n<li>5 fold CV</li>\n<li>Adam optimizer</li>\n<li>batch size: 8</li>\n</ul>\n\n<p>Learning rate is scheduled from 1e-3 to 1e-6 by cosine annealing with 3 cycles, 16 epochs per 1 cycle.\nIn each cycle, the training condition is modified like below,</p>\n\n<ul>\n<li>Cycle 1: train only with opacity images.</li>\n<li>Cycle 2: train with opacity and no-opacity images with appearance rate 1:1.</li>\n<li>Cycle 3: train with opacity and no-opacity images with appearance rate 1:1 and many augmentations.</li>\n</ul>\n\n<p>I use cross-entropy loss with class weight; background:opacity = 1:2.\nI try focal loss and lovasz loss but they don't work,\nmaybe because mask shape is rough and so losses more aware with mask edge are not preferable.</p>\n\n<h3>Augmentations</h3>\n\n<p>I use <a href=\"https://github.com/albu/albumentations\">Albumentations</a> for augmentation.</p>\n\n<ul>\n<li>Cycle1 and 2 with cropped input: random cropping, value shifting, and horizontal flip</li>\n<li>Cycle3 with cropped input: +\nshifting, scaling, rotation, CLAHE, contrast, brightness, gamma, Gaussian noise, and CutOut</li>\n<li>Cycle1 and 2 with resized input: shifting, scaling, rotation, value shifting, and horizontal flip</li>\n<li>Cycle3 with resized input: +\nCLAHE, contrast, brightness, gamma, Gaussian noise, and CutOut</li>\n</ul>\n\n<h3>TTA and ensemble</h3>\n\n<ul>\n<li>3models</li>\n<li>5 fold CV</li>\n<li>horizontal flip</li>\n<li>Cycle 2 weight, Cycle 3 weight and Cycle 3 weight with CLAHE input</li>\n</ul>\n\n<p>In total, 3x5x2x3=90 predictions are averaged.</p>\n\n<h3>Postprocessing</h3>\n\n<p>I generate predicted mask by thresholding model output.\nbounding boxes are generated by this predicted mask.\nThe peak value of model output is used as a confidence score.\nA bounding box with confidence score under a threshold is removed.\nPreferable confidence-threshold are very different between local CV and Stage 1 public LB.\nAs described <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723\">here</a>,\nThe most of stage 1 train labels are verified by 1 doctor and test data of stage 1 and 2 are verified by three doctors including certificated radiologist.\nIt looks like many images regarded as no opacity in train label criteria are regarded as opacity in test label criteria.\nFor the train data, the best mask threshold = 0.50 and the best confidence threshold = 0.77.\nFor the stage 1 test, the best mask threshold = 0.51 and the best confidence threshold = 0.62.\nWhile Stage 1, I searched the best threshold based on public LB score\nand I select mask threshold = 0.50 and confidence threshold = 0.60.</p>\n\n<p>Here, I show some results.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/415221/10600/sample.png\" alt=\"result\"></p>\n\n<h3>Things do not work</h3>\n\n<p>I try to use classification models to improve segmentation prediction.\nFor classification 3 class label and NIH Chest X-ray 14 data can be used and it looks promising.\nI use SE-ResNeXt model.\nAs a modification, I try mean teacher semi-supervised learning using NIH Chest X-ray 14 images.\nI also try multi-task training using 3 class label and 14 class label of NIH Chest X-ray 14 dataset.\nAUC of each model is as below.</p>\n\n<p>model / AUC <br>\nSE-ResNeXt baseline / 0.853 <br>\nSE-ResNeXt mean teacher / 0.897 <br>\nSE-ResNeXt multi task / 0.880  </p>\n\n<p>I try to use prediction of classification for selecting bounding boxes but it makes a score decrease.\nThe peak value of segmentation prediction can classify images with about AUC 0.89,\nso maybe there is no room to improve by classification models.</p>\n\n<p>I try to apply <a href=\"https://arxiv.org/abs/1703.01780\">mean teacher</a> to segmentation model.\nIt may be promising as described in <a href=\"https://arxiv.org/abs/1807.04657\">this paper</a>.\nBut mean teacher makes model bad. In this competition task, prediction quality is low so that self-teaching may not work well.</p>\n\n<p>I try to classify predicted bounding box is hit or not by LGBM as like <a href=\"https://www.kaggle.com/c/data-science-bowl-2018/discussion/54741\">the DSB 2018 1st solution</a>.</p>\n\n<h3>Score History</h3>\n\n<ul>\n<li>Xception-DeepLabV3+ (crop input) -&gt; LB 0.077  </li>\n<li>remove low bounding box with low confidence score  -&gt; LB 0.146  </li>\n<li>change mask shape from rectangle to ellipse  -&gt; LB 0.163  </li>\n<li>SE-ResNeXt-DeepLabV3+ (crop input)  -&gt; LB 0.186  </li>\n<li>add SE-ResNeXt-DeepLabV3+ (resize input) -&gt; LB 0.195  </li>\n<li>cycle 2 training  -&gt; LB 0.207  </li>\n<li>add SE-ResNeXt-DeepLabV3+ (resize input) and flip TTA  -&gt; LB 0.215  </li>\n<li>bug fix (lol)  -&gt; LB 0.219  </li>\n<li>add cycle 3 training and CLAHE TTA (final model)  -&gt; LB  0.221 -&gt; stage 2 LB 0.219 </li>\n<li>resized 87.5% (following 1st solution, late submission) -&gt; stage 2 LB 0.234 (Wow!)  </li>\n</ul>",
  "messages": [
    {
      "id": 415221,
      "postDate": "2018-11-04T17:41:29.883Z",
      "content": "<p>Here I would like to share my solution.</p>\n\n<p>My score is Stage 1 Public LB 0.211 (42nd)/Stage 2 Private LB 0.218 (18th)</p>\n\n<h3>Models</h3>\n\n<p>I tackle this competition with semantic segmentation approach because,  </p>\n\n<ul>\n<li>There are a few objects and each object is well separated, so splitting objects from segmentation mask must be easy.</li>\n<li>Shapes of opacity areas are ambiguous, so rough masks generated from bounding boxes are not so unnatural.</li>\n<li>Training a segmentation model is easier than an object detection model, I believe.</li>\n<li>A segmentation model suits for ensemble compared to an object detection model.</li>\n</ul>\n\n<p>I choose DeepLabV3+ model, the state of the art for semantic segmentation. I take the DeepLabV3 implementation from <br>\n<a href=\"https://github.com/jfzhang95/pytorch-deeplab-xception\">https://github.com/jfzhang95/pytorch-deeplab-xception</a> <br>\nI change the model head from Xception to SENet or SE-ResNext101, because Xception head is not trained stably somehow. I take the SENet and SE-ResNext101 implementation from <br>\n<a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a>  </p>\n\n<p>I train 3 models.</p>\n\n<ol>\n<li>SE-ResNext101-DeepLabV3+ (resized input)</li>\n<li>SE-ResNext101-DeepLabV3+ (cropped input)</li>\n<li>SENet-DeepLabV3+ (resized input)</li>\n</ol>\n\n<p>For resized-input models, input images are resized into 448x448. For cropped-input models, input images are cropped into 448x448 in training and original size image is used in prediction. Using 2 input patterns, I hope global and local features are learned.</p>\n\n<h3>Preprocessing</h3>\n\n<p>I use ellipses inscribed to bounding boxes as true masks. rounded masks are more natural than rectangles and applicable for rotate augmentation.</p>\n\n<h3>Training</h3>\n\n<ul>\n<li>5 fold CV</li>\n<li>Adam optimizer</li>\n<li>batch size: 8</li>\n</ul>\n\n<p>Learning rate is scheduled from 1e-3 to 1e-6 by cosine annealing with 3 cycles, 16 epochs per 1 cycle.\nIn each cycle, the training condition is modified like below,</p>\n\n<ul>\n<li>Cycle 1: train only with opacity images.</li>\n<li>Cycle 2: train with opacity and no-opacity images with appearance rate 1:1.</li>\n<li>Cycle 3: train with opacity and no-opacity images with appearance rate 1:1 and many augmentations.</li>\n</ul>\n\n<p>I use cross-entropy loss with class weight; background:opacity = 1:2.\nI try focal loss and lovasz loss but they don't work,\nmaybe because mask shape is rough and so losses more aware with mask edge are not preferable.</p>\n\n<h3>Augmentations</h3>\n\n<p>I use <a href=\"https://github.com/albu/albumentations\">Albumentations</a> for augmentation.</p>\n\n<ul>\n<li>Cycle1 and 2 with cropped input: random cropping, value shifting, and horizontal flip</li>\n<li>Cycle3 with cropped input: +\nshifting, scaling, rotation, CLAHE, contrast, brightness, gamma, Gaussian noise, and CutOut</li>\n<li>Cycle1 and 2 with resized input: shifting, scaling, rotation, value shifting, and horizontal flip</li>\n<li>Cycle3 with resized input: +\nCLAHE, contrast, brightness, gamma, Gaussian noise, and CutOut</li>\n</ul>\n\n<h3>TTA and ensemble</h3>\n\n<ul>\n<li>3models</li>\n<li>5 fold CV</li>\n<li>horizontal flip</li>\n<li>Cycle 2 weight, Cycle 3 weight and Cycle 3 weight with CLAHE input</li>\n</ul>\n\n<p>In total, 3x5x2x3=90 predictions are averaged.</p>\n\n<h3>Postprocessing</h3>\n\n<p>I generate predicted mask by thresholding model output.\nbounding boxes are generated by this predicted mask.\nThe peak value of model output is used as a confidence score.\nA bounding box with confidence score under a threshold is removed.\nPreferable confidence-threshold are very different between local CV and Stage 1 public LB.\nAs described <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723\">here</a>,\nThe most of stage 1 train labels are verified by 1 doctor and test data of stage 1 and 2 are verified by three doctors including certificated radiologist.\nIt looks like many images regarded as no opacity in train label criteria are regarded as opacity in test label criteria.\nFor the train data, the best mask threshold = 0.50 and the best confidence threshold = 0.77.\nFor the stage 1 test, the best mask threshold = 0.51 and the best confidence threshold = 0.62.\nWhile Stage 1, I searched the best threshold based on public LB score\nand I select mask threshold = 0.50 and confidence threshold = 0.60.</p>\n\n<p>Here, I show some results.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/415221/10600/sample.png\" alt=\"result\"></p>\n\n<h3>Things do not work</h3>\n\n<p>I try to use classification models to improve segmentation prediction.\nFor classification 3 class label and NIH Chest X-ray 14 data can be used and it looks promising.\nI use SE-ResNeXt model.\nAs a modification, I try mean teacher semi-supervised learning using NIH Chest X-ray 14 images.\nI also try multi-task training using 3 class label and 14 class label of NIH Chest X-ray 14 dataset.\nAUC of each model is as below.</p>\n\n<p>model / AUC <br>\nSE-ResNeXt baseline / 0.853 <br>\nSE-ResNeXt mean teacher / 0.897 <br>\nSE-ResNeXt multi task / 0.880  </p>\n\n<p>I try to use prediction of classification for selecting bounding boxes but it makes a score decrease.\nThe peak value of segmentation prediction can classify images with about AUC 0.89,\nso maybe there is no room to improve by classification models.</p>\n\n<p>I try to apply <a href=\"https://arxiv.org/abs/1703.01780\">mean teacher</a> to segmentation model.\nIt may be promising as described in <a href=\"https://arxiv.org/abs/1807.04657\">this paper</a>.\nBut mean teacher makes model bad. In this competition task, prediction quality is low so that self-teaching may not work well.</p>\n\n<p>I try to classify predicted bounding box is hit or not by LGBM as like <a href=\"https://www.kaggle.com/c/data-science-bowl-2018/discussion/54741\">the DSB 2018 1st solution</a>.</p>\n\n<h3>Score History</h3>\n\n<ul>\n<li>Xception-DeepLabV3+ (crop input) -&gt; LB 0.077  </li>\n<li>remove low bounding box with low confidence score  -&gt; LB 0.146  </li>\n<li>change mask shape from rectangle to ellipse  -&gt; LB 0.163  </li>\n<li>SE-ResNeXt-DeepLabV3+ (crop input)  -&gt; LB 0.186  </li>\n<li>add SE-ResNeXt-DeepLabV3+ (resize input) -&gt; LB 0.195  </li>\n<li>cycle 2 training  -&gt; LB 0.207  </li>\n<li>add SE-ResNeXt-DeepLabV3+ (resize input) and flip TTA  -&gt; LB 0.215  </li>\n<li>bug fix (lol)  -&gt; LB 0.219  </li>\n<li>add cycle 3 training and CLAHE TTA (final model)  -&gt; LB  0.221 -&gt; stage 2 LB 0.219 </li>\n<li>resized 87.5% (following 1st solution, late submission) -&gt; stage 2 LB 0.234 (Wow!)  </li>\n</ul>",
      "rawMarkdown": "Here I would like to share my solution.\n\nMy score is Stage 1 Public LB 0.211 (42nd)/Stage 2 Private LB 0.218 (18th)\n\n\n### Models\nI tackle this competition with semantic segmentation approach because,  \n\n- There are a few objects and each object is well separated, so splitting objects from segmentation mask must be easy.\n- Shapes of opacity areas are ambiguous, so rough masks generated from bounding boxes are not so unnatural.\n- Training a segmentation model is easier than an object detection model, I believe.\n- A segmentation model suits for ensemble compared to an object detection model.\n\nI choose DeepLabV3+ model, the state of the art for semantic segmentation. I take the DeepLabV3 implementation from  \nhttps://github.com/jfzhang95/pytorch-deeplab-xception  \nI change the model head from Xception to SENet or SE-ResNext101, because Xception head is not trained stably somehow. I take the SENet and SE-ResNext101 implementation from  \nhttps://github.com/Cadene/pretrained-models.pytorch  \n\nI train 3 models.\n\n1. SE-ResNext101-DeepLabV3+ (resized input)\n2. SE-ResNext101-DeepLabV3+ (cropped input)\n3. SENet-DeepLabV3+ (resized input)\n\nFor resized-input models, input images are resized into 448x448. For cropped-input models, input images are cropped into 448x448 in training and original size image is used in prediction. Using 2 input patterns, I hope global and local features are learned.\n\n### Preprocessing\nI use ellipses inscribed to bounding boxes as true masks. rounded masks are more natural than rectangles and applicable for rotate augmentation.\n\n### Training\n- 5 fold CV\n- Adam optimizer\n- batch size: 8\n\nLearning rate is scheduled from 1e-3 to 1e-6 by cosine annealing with 3 cycles, 16 epochs per 1 cycle.\nIn each cycle, the training condition is modified like below,\n\n- Cycle 1: train only with opacity images.\n- Cycle 2: train with opacity and no-opacity images with appearance rate 1:1.\n- Cycle 3: train with opacity and no-opacity images with appearance rate 1:1 and many augmentations.\n\nI use cross-entropy loss with class weight; background:opacity = 1:2.\nI try focal loss and lovasz loss but they don't work,\nmaybe because mask shape is rough and so losses more aware with mask edge are not preferable.\n\n### Augmentations\nI use [Albumentations](https://github.com/albu/albumentations) for augmentation.\n\n- Cycle1 and 2 with cropped input: random cropping, value shifting, and horizontal flip\n- Cycle3 with cropped input: +\nshifting, scaling, rotation, CLAHE, contrast, brightness, gamma, Gaussian noise, and CutOut\n- Cycle1 and 2 with resized input: shifting, scaling, rotation, value shifting, and horizontal flip\n- Cycle3 with resized input: +\nCLAHE, contrast, brightness, gamma, Gaussian noise, and CutOut\n\n### TTA and ensemble\n- 3models\n- 5 fold CV\n- horizontal flip\n- Cycle 2 weight, Cycle 3 weight and Cycle 3 weight with CLAHE input\n\nIn total, 3x5x2x3=90 predictions are averaged.\n\n\n### Postprocessing\nI generate predicted mask by thresholding model output.\nbounding boxes are generated by this predicted mask.\nThe peak value of model output is used as a confidence score.\nA bounding box with confidence score under a threshold is removed.\nPreferable confidence-threshold are very different between local CV and Stage 1 public LB.\nAs described [here](https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723),\nThe most of stage 1 train labels are verified by 1 doctor and test data of stage 1 and 2 are verified by three doctors including certificated radiologist.\nIt looks like many images regarded as no opacity in train label criteria are regarded as opacity in test label criteria.\nFor the train data, the best mask threshold = 0.50 and the best confidence threshold = 0.77.\nFor the stage 1 test, the best mask threshold = 0.51 and the best confidence threshold = 0.62.\nWhile Stage 1, I searched the best threshold based on public LB score\nand I select mask threshold = 0.50 and confidence threshold = 0.60.\n\nHere, I show some results.\n\n![result][1]\n\n### Things do not work\nI try to use classification models to improve segmentation prediction.\nFor classification 3 class label and NIH Chest X-ray 14 data can be used and it looks promising.\nI use SE-ResNeXt model.\nAs a modification, I try mean teacher semi-supervised learning using NIH Chest X-ray 14 images.\nI also try multi-task training using 3 class label and 14 class label of NIH Chest X-ray 14 dataset.\nAUC of each model is as below.\n\nmodel / AUC  \nSE-ResNeXt baseline / 0.853  \nSE-ResNeXt mean teacher / 0.897  \nSE-ResNeXt multi task / 0.880  \n\nI try to use prediction of classification for selecting bounding boxes but it makes a score decrease.\nThe peak value of segmentation prediction can classify images with about AUC 0.89,\nso maybe there is no room to improve by classification models.\n\nI try to apply [mean teacher](https://arxiv.org/abs/1703.01780) to segmentation model.\nIt may be promising as described in [this paper](https://arxiv.org/abs/1807.04657).\nBut mean teacher makes model bad. In this competition task, prediction quality is low so that self-teaching may not work well.\n\nI try to classify predicted bounding box is hit or not by LGBM as like [the DSB 2018 1st solution](https://www.kaggle.com/c/data-science-bowl-2018/discussion/54741).\n\n\n### Score History\n- Xception-DeepLabV3+ (crop input) -&gt; LB 0.077  \n- remove low bounding box with low confidence score  -&gt; LB 0.146  \n- change mask shape from rectangle to ellipse  -&gt; LB 0.163  \n- SE-ResNeXt-DeepLabV3+ (crop input)  -&gt; LB 0.186  \n- add SE-ResNeXt-DeepLabV3+ (resize input) -&gt; LB 0.195  \n- cycle 2 training  -&gt; LB 0.207  \n- add SE-ResNeXt-DeepLabV3+ (resize input) and flip TTA  -&gt; LB 0.215  \n- bug fix (lol)  -&gt; LB 0.219  \n- add cycle 3 training and CLAHE TTA (final model)  -&gt; LB  0.221 -&gt; stage 2 LB 0.219 \n- resized 87.5% (following 1st solution, late submission) -&gt; stage 2 LB 0.234 (Wow!)  \n\n [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/415221/10600/sample.png",
      "votes": 40
    },
    {
      "id": 415364,
      "postDate": "2018-11-05T01:55:56.553Z",
      "content": "<p>Thanks for sharing. It's great to see a segmentation solution perform so well. My intuition is that segmentation might lead to smaller bounding box predictions. Did you find this to be true? </p>",
      "rawMarkdown": "Thanks for sharing. It's great to see a segmentation solution perform so well. My intuition is that segmentation might lead to smaller bounding box predictions. Did you find this to be true? ",
      "votes": 1,
      "replies": [
        {
          "id": 415622,
          "postDate": "2018-11-05T12:03:12.653Z",
          "content": "<p>In my case, my model tends to predict bigger bounding box than the true one (you can find it from the fig above). Maybe It is because I used class weighted cross-entropy (background:opacity = 1:2). BTW, your resize 87.5% approach give me score 0.234. What a clever solution! </p>",
          "rawMarkdown": "In my case, my model tends to predict bigger bounding box than the true one (you can find it from the fig above). Maybe It is because I used class weighted cross-entropy (background:opacity = 1:2). BTW, your resize 87.5% approach give me score 0.234. What a clever solution! ",
          "votes": 1
        },
        {
          "id": 415724,
          "postDate": "2018-11-05T15:04:27.717Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 415730,
          "postDate": "2018-11-05T15:12:18.277Z",
          "content": "<p>Thanks for sharing @OsciiArt, a very nice work, and actually I prefer the straight forward approach of a segmentation than that of rcnn. Do you have access to stage 2 test set labels?</p>",
          "rawMarkdown": "Thanks for sharing @OsciiArt, a very nice work, and actually I prefer the straight forward approach of a segmentation than that of rcnn. Do you have access to stage 2 test set labels?"
        },
        {
          "id": 415921,
          "postDate": "2018-11-05T23:09:19.183Z",
          "content": "<p>No, the score 0.234 is a result of late submission.</p>",
          "rawMarkdown": "No, the score 0.234 is a result of late submission.",
          "votes": 1
        },
        {
          "id": 417428,
          "postDate": "2018-11-08T08:50:35.083Z",
          "content": "<p>Thank you, @OsciiArt, and another question - what are the values of the learning rates? I think that there was a typo</p>",
          "rawMarkdown": "Thank you, @OsciiArt, and another question - what are the values of the learning rates? I think that there was a typo"
        },
        {
          "id": 418010,
          "postDate": "2018-11-09T05:23:09.637Z",
          "content": "<p>@Hader, thank you for pointing that out. I fixed the typo.</p>",
          "rawMarkdown": "@Hader, thank you for pointing that out. I fixed the typo.",
          "votes": 1
        }
      ]
    },
    {
      "id": 652916,
      "postDate": "2019-10-19T16:05:29.077Z",
      "content": "<p>hi <a href=\"/osciiart\">@osciiart</a>  thanks for posting do you have github location to find the model details  ?</p>",
      "rawMarkdown": "hi @osciiart  thanks for posting do you have github location to find the model details  ?"
    },
    {
      "id": 416536,
      "postDate": "2018-11-06T20:59:45.333Z",
      "content": "<p>Congrats and thanks for sharing. Interesting to see segmentation solution working so well with this data.</p>",
      "rawMarkdown": "Congrats and thanks for sharing. Interesting to see segmentation solution working so well with this data."
    },
    {
      "id": 415600,
      "postDate": "2018-11-05T11:30:30.500Z",
      "content": "<p>Congratulations. That's a very creative solution.</p>",
      "rawMarkdown": "Congratulations. That's a very creative solution."
    },
    {
      "id": 418729,
      "postDate": "2018-11-10T13:57:51.603Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 415960,
      "postDate": "2018-11-06T00:59:12.147Z",
      "content": "<p>Thanks for sharing~</p>",
      "rawMarkdown": "Thanks for sharing~"
    }
  ],
  "comments": [
    {
      "id": 415364,
      "author_name": "Ian Pan",
      "author_url": "",
      "post_date": "2018-11-05T01:55:56.553000",
      "content": "<p>Thanks for sharing. It's great to see a segmentation solution perform so well. My intuition is that segmentation might lead to smaller bounding box predictions. Did you find this to be true? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 415622,
          "author_name": "OsciiArt",
          "author_url": "",
          "post_date": "2018-11-05T12:03:12.653000",
          "content": "<p>In my case, my model tends to predict bigger bounding box than the true one (you can find it from the fig above). Maybe It is because I used class weighted cross-entropy (background:opacity = 1:2). BTW, your resize 87.5% approach give me score 0.234. What a clever solution! </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 415724,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-05T15:04:27.717000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415730,
          "author_name": "Hadar",
          "author_url": "",
          "post_date": "2018-11-05T15:12:18.277000",
          "content": "<p>Thanks for sharing @OsciiArt, a very nice work, and actually I prefer the straight forward approach of a segmentation than that of rcnn. Do you have access to stage 2 test set labels?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415921,
          "author_name": "OsciiArt",
          "author_url": "",
          "post_date": "2018-11-05T23:09:19.183000",
          "content": "<p>No, the score 0.234 is a result of late submission.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 417428,
          "author_name": "Hadar",
          "author_url": "",
          "post_date": "2018-11-08T08:50:35.083000",
          "content": "<p>Thank you, @OsciiArt, and another question - what are the values of the learning rates? I think that there was a typo</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 418010,
          "author_name": "OsciiArt",
          "author_url": "",
          "post_date": "2018-11-09T05:23:09.637000",
          "content": "<p>@Hader, thank you for pointing that out. I fixed the typo.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 652916,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2019-10-19T16:05:29.077000",
      "content": "<p>hi <a href=\"/osciiart\">@osciiart</a>  thanks for posting do you have github location to find the model details  ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 416536,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2018-11-06T20:59:45.333000",
      "content": "<p>Congrats and thanks for sharing. Interesting to see segmentation solution working so well with this data.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 415600,
      "author_name": "Yee Ng",
      "author_url": "",
      "post_date": "2018-11-05T11:30:30.500000",
      "content": "<p>Congratulations. That's a very creative solution.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 418729,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-10T13:57:51.603000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 415960,
      "author_name": "Kent Chiu",
      "author_url": "",
      "post_date": "2018-11-06T00:59:12.147000",
      "content": "<p>Thanks for sharing~</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "415221": "Here I would like to share my solution.\n\nMy score is Stage 1 Public LB 0.211 (42nd)/Stage 2 Private LB 0.218 (18th)\n\n\n### Models\nI tackle this competition with semantic segmentation approach because,  \n\n- There are a few objects and each object is well separated, so splitting objects from segmentation mask must be easy.\n- Shapes of opacity areas are ambiguous, so rough masks generated from bounding boxes are not so unnatural.\n- Training a segmentation model is easier than an object detection model, I believe.\n- A segmentation model suits for ensemble compared to an object detection model.\n\nI choose DeepLabV3+ model, the state of the art for semantic segmentation. I take the DeepLabV3 implementation from  \nhttps://github.com/jfzhang95/pytorch-deeplab-xception  \nI change the model head from Xception to SENet or SE-ResNext101, because Xception head is not trained stably somehow. I take the SENet and SE-ResNext101 implementation from  \nhttps://github.com/Cadene/pretrained-models.pytorch  \n\nI train 3 models.\n\n1. SE-ResNext101-DeepLabV3+ (resized input)\n2. SE-ResNext101-DeepLabV3+ (cropped input)\n3. SENet-DeepLabV3+ (resized input)\n\nFor resized-input models, input images are resized into 448x448. For cropped-input models, input images are cropped into 448x448 in training and original size image is used in prediction. Using 2 input patterns, I hope global and local features are learned.\n\n### Preprocessing\nI use ellipses inscribed to bounding boxes as true masks. rounded masks are more natural than rectangles and applicable for rotate augmentation.\n\n### Training\n- 5 fold CV\n- Adam optimizer\n- batch size: 8\n\nLearning rate is scheduled from 1e-3 to 1e-6 by cosine annealing with 3 cycles, 16 epochs per 1 cycle.\nIn each cycle, the training condition is modified like below,\n\n- Cycle 1: train only with opacity images.\n- Cycle 2: train with opacity and no-opacity images with appearance rate 1:1.\n- Cycle 3: train with opacity and no-opacity images with appearance rate 1:1 and many augmentations.\n\nI use cross-entropy loss with class weight; background:opacity = 1:2.\nI try focal loss and lovasz loss but they don't work,\nmaybe because mask shape is rough and so losses more aware with mask edge are not preferable.\n\n### Augmentations\nI use [Albumentations](https://github.com/albu/albumentations) for augmentation.\n\n- Cycle1 and 2 with cropped input: random cropping, value shifting, and horizontal flip\n- Cycle3 with cropped input: +\nshifting, scaling, rotation, CLAHE, contrast, brightness, gamma, Gaussian noise, and CutOut\n- Cycle1 and 2 with resized input: shifting, scaling, rotation, value shifting, and horizontal flip\n- Cycle3 with resized input: +\nCLAHE, contrast, brightness, gamma, Gaussian noise, and CutOut\n\n### TTA and ensemble\n- 3models\n- 5 fold CV\n- horizontal flip\n- Cycle 2 weight, Cycle 3 weight and Cycle 3 weight with CLAHE input\n\nIn total, 3x5x2x3=90 predictions are averaged.\n\n\n### Postprocessing\nI generate predicted mask by thresholding model output.\nbounding boxes are generated by this predicted mask.\nThe peak value of model output is used as a confidence score.\nA bounding box with confidence score under a threshold is removed.\nPreferable confidence-threshold are very different between local CV and Stage 1 public LB.\nAs described [here](https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723),\nThe most of stage 1 train labels are verified by 1 doctor and test data of stage 1 and 2 are verified by three doctors including certificated radiologist.\nIt looks like many images regarded as no opacity in train label criteria are regarded as opacity in test label criteria.\nFor the train data, the best mask threshold = 0.50 and the best confidence threshold = 0.77.\nFor the stage 1 test, the best mask threshold = 0.51 and the best confidence threshold = 0.62.\nWhile Stage 1, I searched the best threshold based on public LB score\nand I select mask threshold = 0.50 and confidence threshold = 0.60.\n\nHere, I show some results.\n\n![result][1]\n\n### Things do not work\nI try to use classification models to improve segmentation prediction.\nFor classification 3 class label and NIH Chest X-ray 14 data can be used and it looks promising.\nI use SE-ResNeXt model.\nAs a modification, I try mean teacher semi-supervised learning using NIH Chest X-ray 14 images.\nI also try multi-task training using 3 class label and 14 class label of NIH Chest X-ray 14 dataset.\nAUC of each model is as below.\n\nmodel / AUC  \nSE-ResNeXt baseline / 0.853  \nSE-ResNeXt mean teacher / 0.897  \nSE-ResNeXt multi task / 0.880  \n\nI try to use prediction of classification for selecting bounding boxes but it makes a score decrease.\nThe peak value of segmentation prediction can classify images with about AUC 0.89,\nso maybe there is no room to improve by classification models.\n\nI try to apply [mean teacher](https://arxiv.org/abs/1703.01780) to segmentation model.\nIt may be promising as described in [this paper](https://arxiv.org/abs/1807.04657).\nBut mean teacher makes model bad. In this competition task, prediction quality is low so that self-teaching may not work well.\n\nI try to classify predicted bounding box is hit or not by LGBM as like [the DSB 2018 1st solution](https://www.kaggle.com/c/data-science-bowl-2018/discussion/54741).\n\n\n### Score History\n- Xception-DeepLabV3+ (crop input) -&gt; LB 0.077  \n- remove low bounding box with low confidence score  -&gt; LB 0.146  \n- change mask shape from rectangle to ellipse  -&gt; LB 0.163  \n- SE-ResNeXt-DeepLabV3+ (crop input)  -&gt; LB 0.186  \n- add SE-ResNeXt-DeepLabV3+ (resize input) -&gt; LB 0.195  \n- cycle 2 training  -&gt; LB 0.207  \n- add SE-ResNeXt-DeepLabV3+ (resize input) and flip TTA  -&gt; LB 0.215  \n- bug fix (lol)  -&gt; LB 0.219  \n- add cycle 3 training and CLAHE TTA (final model)  -&gt; LB  0.221 -&gt; stage 2 LB 0.219 \n- resized 87.5% (following 1st solution, late submission) -&gt; stage 2 LB 0.234 (Wow!)  \n\n [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/415221/10600/sample.png",
    "415364": "Thanks for sharing. It's great to see a segmentation solution perform so well. My intuition is that segmentation might lead to smaller bounding box predictions. Did you find this to be true? ",
    "652916": "hi @osciiart  thanks for posting do you have github location to find the model details  ?",
    "416536": "Congrats and thanks for sharing. Interesting to see segmentation solution working so well with this data.",
    "415600": "Congratulations. That's a very creative solution.",
    "418729": "",
    "415960": "Thanks for sharing~"
  }
}