{
  "id": 70464,
  "title": "14th place solution [6th if resized the boxes]",
  "url": "/competitions/rsna-pneumonia-detection-challenge/writeups/xin-yi-14th-place-solution-6th-if-resized-the-boxe",
  "author_name": "",
  "post_date": "2019-02-14T22:48:17.303Z",
  "votes": 16,
  "comment_count": 7,
  "views": 0,
  "content": "<p>First, I would like to thank the organizers/annotators who make this data and challenge available. It's gonna be a valuable resource to the medical imaging community. I learned a lot by going through the discussion forum, shared kernels, and most importantly through top winners' solutions. \n<strong>Congratulations to you all!!</strong></p>\n\n<p>Here I would like to share an overview of my solution (1 segmentation net and 4 detection nets) and some general thoughts. Hopefully, it can be useful to someone.  All my models were trained on a machine with a single 1060 GPU.</p>\n\n<p>Stage 1 LB score: 0.220 </p>\n\n<p>Stage 2 LB score: 0.222</p>\n\n<hr>\n\n<h1>Segmentation:</h1>\n\n<p>The first part of my solution is a semantic segmentation network. It follows a UNet-like architecture and trained with dice loss and focal loss from scratch. The training set includes all the negative and positive samples. It's basically a patch-based classifier that gives each pixel a probability value of being pneumonia. I thresholded the raw output with a value of 0.1 cause this gives the highest sensitivity. </p>\n\n<h2>Details:</h2>\n\n<ul>\n<li><p>image size: 256x256</p></li>\n<li><p>preprocessing: contrast limited adaptive histogram equalization</p></li>\n<li><p>training time: ~20 hours</p></li>\n<li><p>train/val: ~24000/~1200 (stratified)</p></li>\n<li><p>post-processing: erosion with a disk of radius 5 (output is too large by visual inspection)</p></li>\n<li><p>optimizer and scheduler: Adam/step scheduler with gamma=0.1 and step_size=20</p></li>\n<li><p>batch size: 8</p></li>\n<li><p>epochs: 26 </p></li>\n<li><p>failures: PSPNet with different backbones; reweight pixels inside each mask according to the intensity</p></li>\n</ul>\n\n<p>I think I've spent roughly 3 weeks trying to make it work, but a single segmentation network alone didn't take me very far. The best score I got in stage 1 was 0.12 using a threshold of 0.9 (local validation score almost 0.3). At roughly the same time, people started to get better results by using Maskrcnn and yoloV3.  Using thresholding on segmentation maps to create submission  just doesn't seem to be as accurate as predicting bounding boxes directly given the evaluation metric of the task. So I switched to detection methods.</p>\n\n<hr>\n\n<h1>Detection:</h1>\n\n<p>I ended up with 4 detection models, one from Maskrcnn, two from fasterrcnn and one from retinaNet. The first two are two-stage methods and the last one is one stage method. I picked them hoping each one can predict from a different perspective. I tried to play around with their hyper-parameters but with no success (could be because of my little experience on detection models) so I used all the default hyper-parameters to create my submission.</p>\n\n<h2>Details:</h2>\n\n<p>Maskrcnn (keras): take directly from Henrique Mendonça's kernel (thanks Henrique!)</p>\n\n<p>fasterrcnn/retinanet (pytorch): trained with code from <a href=\"https://github.com/open-mmlab/mmdetection\">https://github.com/open-mmlab/mmdetection</a></p>\n\n<ul>\n<li><p>image size: 512x512 </p></li>\n<li><p>preprocessing: none</p></li>\n<li><p>post-processing: NMS</p></li>\n<li><p>training time: ~5 hours</p></li>\n<li><p>epochs: 10 for fasterrcnn 4 for retinanet</p></li>\n<li><p>batch size: 2</p></li>\n<li><p>train/val:~5000/~500 (only positive samples, for retinanet as well; each one was trained on a different split)</p></li>\n<li><p>backbone: resnet50</p></li>\n</ul>\n\n<p>My largest performance boost is from combining my segmentation result with Maskrcnn result. Trying simply intersection gave me a score of ~0.19. The performance gain is mainly resulted by the fact that a lot of false positives from the detection network got filtered out. It makes a lot of sense because the detection network didn't see the whole bunch of negative images. Later on, I just incorporated more detection models as mentioned above and I finally got a score of 0.22 in stage 1.</p>\n\n<hr>\n\n<h1>Thoughts</h1>\n\n<p>Generally speaking, you gonna have to find ways to use all the images (both positive and negative samples) to train your network. I also believe using NIH dataset would be beneficial in some way but can't afford to do any experiment. Below is my rate of the importance of various aspects to the success of this challenge.</p>\n\n<ul>\n<li><p><strong>ensemble</strong> ****:  In cases where there's too much variability in the dataset, the ensemble is the right way to go.</p></li>\n<li><p><strong>post-processing</strong> ****:  Both Ian and Dmytro have mentioned resizing the box. I didn't try resizing in my submission though. If I had done it (shrinking box by a factor of 0.875),  my stage2 score would have been 0.236 according to late submission.</p></li>\n<li><p><strong>validation</strong> ****: Many people have lost their rank due to over-fitting to the stage 1 LB. So don't be too obsessed ​with the public score.</p></li>\n<li><p><strong>model adoption</strong> ***: Detection models seem to be better than segm​entation models in this challenge. Among detection models, retinanet is the most elegant one due to its simplest pipeline and ability to handle​ both positive and negative sample. I don't think there's a big difference between the other 2 stage detection models.</p></li>\n<li><p><strong>backbone</strong> ***: se-resnext101 seems to be the best backbone so far on this task according to the other teams' solution share.</p></li>\n<li><p><strong>augmentation</strong> **   : I only tried hori​zontal flip, small scale an​d rotation transfo​rm. The other fancier augmentation operations do not seem to help much.</p></li>\n<li><p><strong>image size</strong>  **: 512x512 is slightly better than 256x256 in detection but worse in segmentation. </p></li>\n<li><p><strong>hyper-parameter tuning</strong> **: used default setting for my detection models</p></li>\n<li><p><strong>preprocessing</strong> *: </p>\n\n<p>If I were to do this challenge again, I would start with retinanet as Dymtro did.</p></li>\n</ul>\n\n<h1>Acknlowledgements</h1>\n\n<p><a href=\"https://www.kaggle.com/chenyc15/mean-average-precision-metric\">Yicheng Chen's metric kernel</a></p>\n\n<p><a href=\"https://www.kaggle.com/hmendonca/mask-rcnn-and-coco-transfer-learning-lb-0-155\">Henrique Mendonça's kernel</a></p>\n\n<p><a href=\"https://github.com/ahrnbom/ensemble-objdet%29\">detection ensemble</a>  mentioned by Ian in a <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/68739#latest-411691\">post</a> </p>",
  "messages": [
    {
      "id": "414992",
      "postDate": "11/04/2018 04:01:01",
      "content": "<p>First, I would like to thank the organizers/annotators who make this data and challenge available. It's gonna be a valuable resource to the medical imaging community. I learned a lot by going through the discussion forum, shared kernels, and most importantly through top winners' solutions. \n<strong>Congratulations to you all!!</strong></p>\n\n<p>Here I would like to share an overview of my solution (1 segmentation net and 4 detection nets) and some general thoughts. Hopefully, it can be useful to someone.  All my models were trained on a machine with a single 1060 GPU.</p>\n\n<p>Stage 1 LB score: 0.220 </p>\n\n<p>Stage 2 LB score: 0.222</p>\n\n<hr>\n\n<h1>Segmentation:</h1>\n\n<p>The first part of my solution is a semantic segmentation network. It follows a UNet-like architecture and trained with dice loss and focal loss from scratch. The training set includes all the negative and positive samples. It's basically a patch-based classifier that gives each pixel a probability value of being pneumonia. I thresholded the raw output with a value of 0.1 cause this gives the highest sensitivity. </p>\n\n<h2>Details:</h2>\n\n<ul>\n<li><p>image size: 256x256</p></li>\n<li><p>preprocessing: contrast limited adaptive histogram equalization</p></li>\n<li><p>training time: ~20 hours</p></li>\n<li><p>train/val: ~24000/~1200 (stratified)</p></li>\n<li><p>post-processing: erosion with a disk of radius 5 (output is too large by visual inspection)</p></li>\n<li><p>optimizer and scheduler: Adam/step scheduler with gamma=0.1 and step_size=20</p></li>\n<li><p>batch size: 8</p></li>\n<li><p>epochs: 26 </p></li>\n<li><p>failures: PSPNet with different backbones; reweight pixels inside each mask according to the intensity</p></li>\n</ul>\n\n<p>I think I've spent roughly 3 weeks trying to make it work, but a single segmentation network alone didn't take me very far. The best score I got in stage 1 was 0.12 using a threshold of 0.9 (local validation score almost 0.3). At roughly the same time, people started to get better results by using Maskrcnn and yoloV3.  Using thresholding on segmentation maps to create submission  just doesn't seem to be as accurate as predicting bounding boxes directly given the evaluation metric of the task. So I switched to detection methods.</p>\n\n<hr>\n\n<h1>Detection:</h1>\n\n<p>I ended up with 4 detection models, one from Maskrcnn, two from fasterrcnn and one from retinaNet. The first two are two-stage methods and the last one is one stage method. I picked them hoping each one can predict from a different perspective. I tried to play around with their hyper-parameters but with no success (could be because of my little experience on detection models) so I used all the default hyper-parameters to create my submission.</p>\n\n<h2>Details:</h2>\n\n<p>Maskrcnn (keras): take directly from Henrique Mendonça's kernel (thanks Henrique!)</p>\n\n<p>fasterrcnn/retinanet (pytorch): trained with code from <a href=\"https://github.com/open-mmlab/mmdetection\">https://github.com/open-mmlab/mmdetection</a></p>\n\n<ul>\n<li><p>image size: 512x512 </p></li>\n<li><p>preprocessing: none</p></li>\n<li><p>post-processing: NMS</p></li>\n<li><p>training time: ~5 hours</p></li>\n<li><p>epochs: 10 for fasterrcnn 4 for retinanet</p></li>\n<li><p>batch size: 2</p></li>\n<li><p>train/val:~5000/~500 (only positive samples, for retinanet as well; each one was trained on a different split)</p></li>\n<li><p>backbone: resnet50</p></li>\n</ul>\n\n<p>My largest performance boost is from combining my segmentation result with Maskrcnn result. Trying simply intersection gave me a score of ~0.19. The performance gain is mainly resulted by the fact that a lot of false positives from the detection network got filtered out. It makes a lot of sense because the detection network didn't see the whole bunch of negative images. Later on, I just incorporated more detection models as mentioned above and I finally got a score of 0.22 in stage 1.</p>\n\n<hr>\n\n<h1>Thoughts</h1>\n\n<p>Generally speaking, you gonna have to find ways to use all the images (both positive and negative samples) to train your network. I also believe using NIH dataset would be beneficial in some way but can't afford to do any experiment. Below is my rate of the importance of various aspects to the success of this challenge.</p>\n\n<ul>\n<li><p><strong>ensemble</strong> ****:  In cases where there's too much variability in the dataset, the ensemble is the right way to go.</p></li>\n<li><p><strong>post-processing</strong> ****:  Both Ian and Dmytro have mentioned resizing the box. I didn't try resizing in my submission though. If I had done it (shrinking box by a factor of 0.875),  my stage2 score would have been 0.236 according to late submission.</p></li>\n<li><p><strong>validation</strong> ****: Many people have lost their rank due to over-fitting to the stage 1 LB. So don't be too obsessed ​with the public score.</p></li>\n<li><p><strong>model adoption</strong> ***: Detection models seem to be better than segm​entation models in this challenge. Among detection models, retinanet is the most elegant one due to its simplest pipeline and ability to handle​ both positive and negative sample. I don't think there's a big difference between the other 2 stage detection models.</p></li>\n<li><p><strong>backbone</strong> ***: se-resnext101 seems to be the best backbone so far on this task according to the other teams' solution share.</p></li>\n<li><p><strong>augmentation</strong> **   : I only tried hori​zontal flip, small scale an​d rotation transfo​rm. The other fancier augmentation operations do not seem to help much.</p></li>\n<li><p><strong>image size</strong>  **: 512x512 is slightly better than 256x256 in detection but worse in segmentation. </p></li>\n<li><p><strong>hyper-parameter tuning</strong> **: used default setting for my detection models</p></li>\n<li><p><strong>preprocessing</strong> *: </p>\n\n<p>If I were to do this challenge again, I would start with retinanet as Dymtro did.</p></li>\n</ul>\n\n<h1>Acknlowledgements</h1>\n\n<p><a href=\"https://www.kaggle.com/chenyc15/mean-average-precision-metric\">Yicheng Chen's metric kernel</a></p>\n\n<p><a href=\"https://www.kaggle.com/hmendonca/mask-rcnn-and-coco-transfer-learning-lb-0-155\">Henrique Mendonça's kernel</a></p>\n\n<p><a href=\"https://github.com/ahrnbom/ensemble-objdet%29\">detection ensemble</a>  mentioned by Ian in a <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/68739#latest-411691\">post</a> </p>",
      "rawMarkdown": "First, I would like to thank the organizers/annotators who make this data and challenge available. It's gonna be a valuable resource to the medical imaging community. I learned a lot by going through the discussion forum, shared kernels, and most importantly through top winners' solutions. \n**Congratulations to you all!!**\n\nHere I would like to share an overview of my solution (1 segmentation net and 4 detection nets) and some general thoughts. Hopefully, it can be useful to someone.  All my models were trained on a machine with a single 1060 GPU.\n\nStage 1 LB score: 0.220 \n\nStage 2 LB score: 0.222\n\n---\n\n# Segmentation:\n\nThe first part of my solution is a semantic segmentation network. It follows a UNet-like architecture and trained with dice loss and focal loss from scratch. The training set includes all the negative and positive samples. It's basically a patch-based classifier that gives each pixel a probability value of being pneumonia. I thresholded the raw output with a value of 0.1 cause this gives the highest sensitivity. \n\n## Details:\n\n- image size: 256x256\n\n- preprocessing: contrast limited adaptive histogram equalization\n\n- training time: ~20 hours\n\n- train/val: ~24000/~1200 (stratified)\n\n- post-processing: erosion with a disk of radius 5 (output is too large by visual inspection)\n\n- optimizer and scheduler: Adam/step scheduler with gamma=0.1 and step_size=20\n\n- batch size: 8\n\n- epochs: 26 \n\n- failures: PSPNet with different backbones; reweight pixels inside each mask according to the intensity\n\nI think I've spent roughly 3 weeks trying to make it work, but a single segmentation network alone didn't take me very far. The best score I got in stage 1 was 0.12 using a threshold of 0.9 (local validation score almost 0.3). At roughly the same time, people started to get better results by using Maskrcnn and yoloV3.  Using thresholding on segmentation maps to create submission  just doesn't seem to be as accurate as predicting bounding boxes directly given the evaluation metric of the task. So I switched to detection methods.\n\n---\n\n# Detection:\n\nI ended up with 4 detection models, one from Maskrcnn, two from fasterrcnn and one from retinaNet. The first two are two-stage methods and the last one is one stage method. I picked them hoping each one can predict from a different perspective. I tried to play around with their hyper-parameters but with no success (could be because of my little experience on detection models) so I used all the default hyper-parameters to create my submission.\n\n## Details:\n\nMaskrcnn (keras): take directly from Henrique Mendonça's kernel (thanks Henrique!)\n\nfasterrcnn/retinanet (pytorch): trained with code from https://github.com/open-mmlab/mmdetection\n\n- image size: 512x512 \n\n- preprocessing: none\n\n- post-processing: NMS\n\n- training time: ~5 hours\n\n- epochs: 10 for fasterrcnn 4 for retinanet\n\n- batch size: 2\n\n- train/val:~5000/~500 (only positive samples, for retinanet as well; each one was trained on a different split)\n\n- backbone: resnet50\n\n\nMy largest performance boost is from combining my segmentation result with Maskrcnn result. Trying simply intersection gave me a score of ~0.19. The performance gain is mainly resulted by the fact that a lot of false positives from the detection network got filtered out. It makes a lot of sense because the detection network didn't see the whole bunch of negative images. Later on, I just incorporated more detection models as mentioned above and I finally got a score of 0.22 in stage 1.\n\n---\n\n# Thoughts\n\nGenerally speaking, you gonna have to find ways to use all the images (both positive and negative samples) to train your network. I also believe using NIH dataset would be beneficial in some way but can't afford to do any experiment. Below is my rate of the importance of various aspects to the success of this challenge.\n\n-  **ensemble** \\****:  In cases where there's too much variability in the dataset, the ensemble is the right way to go.\n\n-   **post-processing** \\****:  Both Ian and Dmytro have mentioned resizing the box. I didn't try resizing in my submission though. If I had done it (shrinking box by a factor of 0.875),  my stage2 score would have been 0.236 according to late submission.\n\n-   **validation** \\****: Many people have lost their rank due to over-fitting to the stage 1 LB. So don't be too obsessed ​with the public score.\n\n-    **model adoption** \\***: Detection models seem to be better than segm​entation models in this challenge. Among detection models, retinanet is the most elegant one due to its simplest pipeline and ability to handle​ both positive and negative sample. I don't think there's a big difference between the other 2 stage detection models.\n\n- **backbone** \\***: se-resnext101 seems to be the best backbone so far on this task according to the other teams' solution share.\n\n- **augmentation** \\**   : I only tried hori​zontal flip, small scale an​d rotation transfo​rm. The other fancier augmentation operations do not seem to help much.\n\n-  **image size**  \\**: 512x512 is slightly better than 256x256 in detection but worse in segmentation. \n\n- **hyper-parameter tuning** \\**: used default setting for my detection models\n\n-  **preprocessing** \\*: \n\n If I were to do this challenge again, I would start with retinanet as Dymtro did.\n\n# Acknlowledgements\n[Yicheng Chen's metric kernel][1]\n\n[Henrique Mendonça's kernel][2]\n\n[detection ensemble][3]  mentioned by Ian in a [post][4] \n\n\n  [1]: https://www.kaggle.com/chenyc15/mean-average-precision-metric\n  [2]: https://www.kaggle.com/hmendonca/mask-rcnn-and-coco-transfer-learning-lb-0-155\n  [3]: https://github.com/ahrnbom/ensemble-objdet)\n  [4]: https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/68739#latest-411691",
      "votes": null
    },
    {
      "id": "415048",
      "postDate": "11/04/2018 08:55:58",
      "content": "<p>Thanks for sharing.\nLots of team (including mine) use classification model to filter out the false positive samples, new to know that U-net segmentation is another approach.</p>",
      "rawMarkdown": "Thanks for sharing.\nLots of team (including mine) use classification model to filter out the false positive samples, new to know that U-net segmentation is another approach.",
      "votes": null
    },
    {
      "id": "415079",
      "postDate": "11/04/2018 11:13:07",
      "content": "<p>can you please share your code.</p>",
      "rawMarkdown": "can you please share your code.",
      "votes": null
    },
    {
      "id": "415321",
      "postDate": "11/04/2018 22:37:14",
      "content": "<p>I'll consider releasing my code if more people are interested.</p>",
      "rawMarkdown": "I'll consider releasing my code if more people are interested.",
      "votes": null
    },
    {
      "id": "415772",
      "postDate": "11/05/2018 17:07:59",
      "content": "<p>great work!</p>",
      "rawMarkdown": "great work!",
      "votes": null
    },
    {
      "id": "416839",
      "postDate": "11/07/2018 10:49:09",
      "content": "<p>Would love to go through and try out your code :)</p>",
      "rawMarkdown": "Would love to go through and try out your code :)",
      "votes": null
    },
    {
      "id": "526808",
      "postDate": "05/03/2019 20:30:52",
      "content": "<p>Please share your code.</p>",
      "rawMarkdown": "Please share your code.",
      "votes": null
    },
    {
      "id": "789731",
      "postDate": "03/28/2020 23:44:34",
      "content": "<p>share your code plz</p>",
      "rawMarkdown": "share your code plz",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 415048,
      "author_name": "lkhphuc",
      "author_url": "",
      "post_date": "11/04/2018 08:55:58",
      "content": "<p>Thanks for sharing.\nLots of team (including mine) use classification model to filter out the false positive samples, new to know that U-net segmentation is another approach.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 415079,
      "author_name": "",
      "author_url": "",
      "post_date": "11/04/2018 11:13:07",
      "content": "<p>can you please share your code.</p>",
      "votes": null,
      "replies": [
        {
          "id": 415321,
          "author_name": "xinario",
          "author_url": "",
          "post_date": "11/04/2018 22:37:14",
          "content": "<p>I'll consider releasing my code if more people are interested.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 416839,
          "author_name": "rishabhiitbhu",
          "author_url": "",
          "post_date": "11/07/2018 10:49:09",
          "content": "<p>Would love to go through and try out your code :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526808,
          "author_name": "usamaraja125",
          "author_url": "",
          "post_date": "05/03/2019 20:30:52",
          "content": "<p>Please share your code.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 415772,
      "author_name": "haimin777",
      "author_url": "",
      "post_date": "11/05/2018 17:07:59",
      "content": "<p>great work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 789731,
      "author_name": "globalicon",
      "author_url": "",
      "post_date": "03/28/2020 23:44:34",
      "content": "<p>share your code plz</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "414992": "First, I would like to thank the organizers/annotators who make this data and challenge available. It's gonna be a valuable resource to the medical imaging community. I learned a lot by going through the discussion forum, shared kernels, and most importantly through top winners' solutions. \n**Congratulations to you all!!**\n\nHere I would like to share an overview of my solution (1 segmentation net and 4 detection nets) and some general thoughts. Hopefully, it can be useful to someone.  All my models were trained on a machine with a single 1060 GPU.\n\nStage 1 LB score: 0.220 \n\nStage 2 LB score: 0.222\n\n---\n\n# Segmentation:\n\nThe first part of my solution is a semantic segmentation network. It follows a UNet-like architecture and trained with dice loss and focal loss from scratch. The training set includes all the negative and positive samples. It's basically a patch-based classifier that gives each pixel a probability value of being pneumonia. I thresholded the raw output with a value of 0.1 cause this gives the highest sensitivity. \n\n## Details:\n\n- image size: 256x256\n\n- preprocessing: contrast limited adaptive histogram equalization\n\n- training time: ~20 hours\n\n- train/val: ~24000/~1200 (stratified)\n\n- post-processing: erosion with a disk of radius 5 (output is too large by visual inspection)\n\n- optimizer and scheduler: Adam/step scheduler with gamma=0.1 and step_size=20\n\n- batch size: 8\n\n- epochs: 26 \n\n- failures: PSPNet with different backbones; reweight pixels inside each mask according to the intensity\n\nI think I've spent roughly 3 weeks trying to make it work, but a single segmentation network alone didn't take me very far. The best score I got in stage 1 was 0.12 using a threshold of 0.9 (local validation score almost 0.3). At roughly the same time, people started to get better results by using Maskrcnn and yoloV3.  Using thresholding on segmentation maps to create submission  just doesn't seem to be as accurate as predicting bounding boxes directly given the evaluation metric of the task. So I switched to detection methods.\n\n---\n\n# Detection:\n\nI ended up with 4 detection models, one from Maskrcnn, two from fasterrcnn and one from retinaNet. The first two are two-stage methods and the last one is one stage method. I picked them hoping each one can predict from a different perspective. I tried to play around with their hyper-parameters but with no success (could be because of my little experience on detection models) so I used all the default hyper-parameters to create my submission.\n\n## Details:\n\nMaskrcnn (keras): take directly from Henrique Mendonça's kernel (thanks Henrique!)\n\nfasterrcnn/retinanet (pytorch): trained with code from https://github.com/open-mmlab/mmdetection\n\n- image size: 512x512 \n\n- preprocessing: none\n\n- post-processing: NMS\n\n- training time: ~5 hours\n\n- epochs: 10 for fasterrcnn 4 for retinanet\n\n- batch size: 2\n\n- train/val:~5000/~500 (only positive samples, for retinanet as well; each one was trained on a different split)\n\n- backbone: resnet50\n\n\nMy largest performance boost is from combining my segmentation result with Maskrcnn result. Trying simply intersection gave me a score of ~0.19. The performance gain is mainly resulted by the fact that a lot of false positives from the detection network got filtered out. It makes a lot of sense because the detection network didn't see the whole bunch of negative images. Later on, I just incorporated more detection models as mentioned above and I finally got a score of 0.22 in stage 1.\n\n---\n\n# Thoughts\n\nGenerally speaking, you gonna have to find ways to use all the images (both positive and negative samples) to train your network. I also believe using NIH dataset would be beneficial in some way but can't afford to do any experiment. Below is my rate of the importance of various aspects to the success of this challenge.\n\n-  **ensemble** \\****:  In cases where there's too much variability in the dataset, the ensemble is the right way to go.\n\n-   **post-processing** \\****:  Both Ian and Dmytro have mentioned resizing the box. I didn't try resizing in my submission though. If I had done it (shrinking box by a factor of 0.875),  my stage2 score would have been 0.236 according to late submission.\n\n-   **validation** \\****: Many people have lost their rank due to over-fitting to the stage 1 LB. So don't be too obsessed ​with the public score.\n\n-    **model adoption** \\***: Detection models seem to be better than segm​entation models in this challenge. Among detection models, retinanet is the most elegant one due to its simplest pipeline and ability to handle​ both positive and negative sample. I don't think there's a big difference between the other 2 stage detection models.\n\n- **backbone** \\***: se-resnext101 seems to be the best backbone so far on this task according to the other teams' solution share.\n\n- **augmentation** \\**   : I only tried hori​zontal flip, small scale an​d rotation transfo​rm. The other fancier augmentation operations do not seem to help much.\n\n-  **image size**  \\**: 512x512 is slightly better than 256x256 in detection but worse in segmentation. \n\n- **hyper-parameter tuning** \\**: used default setting for my detection models\n\n-  **preprocessing** \\*: \n\n If I were to do this challenge again, I would start with retinanet as Dymtro did.\n\n# Acknlowledgements\n[Yicheng Chen's metric kernel][1]\n\n[Henrique Mendonça's kernel][2]\n\n[detection ensemble][3]  mentioned by Ian in a [post][4] \n\n\n  [1]: https://www.kaggle.com/chenyc15/mean-average-precision-metric\n  [2]: https://www.kaggle.com/hmendonca/mask-rcnn-and-coco-transfer-learning-lb-0-155\n  [3]: https://github.com/ahrnbom/ensemble-objdet)\n  [4]: https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/68739#latest-411691",
    "415048": "Thanks for sharing.\nLots of team (including mine) use classification model to filter out the false positive samples, new to know that U-net segmentation is another approach.",
    "415079": "can you please share your code.",
    "415321": "I'll consider releasing my code if more people are interested.",
    "415772": "great work!",
    "416839": "Would love to go through and try out your code :)",
    "526808": "Please share your code.",
    "789731": "share your code plz"
  },
  "source": "meta"
}