{
  "id": 70632,
  "title": "3rd place solution",
  "url": "/competitions/rsna-pneumonia-detection-challenge/writeups/phillip-cheng-3rd-place-solution",
  "author_name": "",
  "post_date": "2018-11-06T04:29:04.103520400Z",
  "votes": 53,
  "comment_count": 24,
  "views": 0,
  "content": "<p>Hello!  This is an overview of my solution to the pneumonia detection challenge.  I'm an abdominal radiologist at the Keck School of Medicine of USC in Los Angeles, CA, USA.  Although I have played with convolutional neural networks for medical image classification, this is my first experience with object detection.  It is also my first experience with machine learning competitions.      </p>\n\n<h2>Summary</h2>\n\n<p>My models used <a href=\"https://github.com/fizyr/keras-retinanet\">keras-retinanet</a> by Hans Gaiser and collaborators, based on the <a href=\"https://arxiv.org/abs/1708.02002\">focal loss paper by Lin et al</a>.  This is a single-stage convolutional neural network detection architecture, which was appealing to me for training simplicity.  I optimized two RetinaNet models using Keras, with resnet-50 and resnet-101 backbones that were pretrained on ImageNet images.   I used non-maximum suppression to eliminate any overlapping bounding boxes from each network.  I then took weighted averages of overlapping bounding boxes from both trained neural networks.  I also applied a global fixed percentage size reduction to all final bounding boxes, which appeared to significantly improve Stage 1 test scores.</p>\n\n<p>Code is posted on <a href=\"https://github.com/pmcheng/rsna-pneumonia\">Github</a>.  Note that the two model .h5 files used in my solution are under the Releases tab of the repository, because these files are too large to be included within the repo.</p>\n\n<h2>Feature selection (or lack thereof)</h2>\n\n<p>I did not perform manual image feature selection or engineering.  I briefly played with using the DICOM view projection (AP vs PA), specifically assigning different score thresholds based on the view, but this did not improve my results.  </p>\n\n<p>I did not make use of the “No Lung Opacity / Not Normal” labels in the training set.</p>\n\n<p>I did not use any external training data for this competition.</p>\n\n<h2>Training</h2>\n\n<p>I decided early on that high image resolution was not necessary for pneumonia bounding box prediction.  I used the training images at a 224 x 224 resolution, which made training much more efficient on my hardware.  I used sklearn’s <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedKFold.html\">stratified K fold function</a> to divide the 25684 training images into a large training set (24399 images, 95%) and a small validation set (1285 images, 5%) .  I had originally used a much larger validation set but later found that I had significantly better training results when I shifted more images into training.  Images were augmented with rotation, translation, scaling, and horizontal flipping (shearing and vertical flipping were turned off).  I also added random constants to the images, an idea I found from a <a href=\"https://www.kaggle.com/c/google-ai-open-images-object-detection-track/discussion/64633\">post by @ZFTurbo</a> regarding RetinaNet in a previous competition.   </p>\n\n<p>The default training parameters were modified to match the focal loss paper (i.e., SGD with learning rate = 0.01, momentum =0.9, decay = 0.0001, nesterov = True), due to anecdotal reports about more generalizable models produced by SGD with momentum compared to Adam.  I used resnet-50 and resnet-101 backbones for the final models.  I did experiment with resnet-152, but the training took longer and the results were slightly worse.</p>\n\n<p>I trained for 25 epochs with 2500 steps per epoch and batch size of 8, saving the model snapshot after each epoch.  After each epoch I calculated bounding boxes on the validation set; I changed the score threshold in the filter_detections layer from keras-retinanet to 0.01 (from 0.05) so that I could evaluate lower score thresholds.  At each epoch I calculated the score threshold providing the maximum Youden index on the validation set (sensitivity + specificity – 1).  More specifically, I calculated sensitivity and specificity with respect to images, not bounding boxes, with the idea that from a scoring standpoint, it was important for the system to classify whether an image as a whole was positive or negative for pneumonia.  I also calculated the RSNA metric as implemented in <a href=\"https://www.kaggle.com/chenyc15/mean-average-precision-metric\">Yicheng Chen's excellent kernel</a>.  I found that snapshots that performed best on the leaderboard were ones with the highest Youden index, with the score threshold lowered to give a sensitivity close to 90%.  I suspect that the benefit of lowered score thresholds was due to higher prevalence of pneumonia in the test set relative to the training set. </p>\n\n<p>I aggressively used non-maximum suppression to eliminate any overlapping bounding boxes from each network’s output for a given image.  My idea was that physician annotators would most likely specify nonoverlapping ground truth bounding boxes.  Even though Tensorflow has an NMS function, I found it more useful and instructive to modify <a href=\"https://www.pyimagesearch.com/2015/02/16/faster-non-maximum-suppression-python\">code posted by Adrian Rosebrock on his blog</a>.</p>\n\n<p>I then took weighted averages of overlapping bounding boxes from both trained neural networks, using the scores for the boxes from each network as the weights; I experimented with box unions and intersections, which did not work nearly as well.  For bounding boxes that did not overlap between the two neural networks, I used a separate higher threshold value to decide whether the solitary box should be retained.  I also applied a global fixed percentage size reduction to all final bounding boxes, which significantly improved Stage 1 test scores.</p>\n\n<p>I did not train with Stage 1 test data for Stage 2, because I didn’t know this was an option (I did not upload automated training code).  However, even if I had been aware of this option, I doubt that I would have retrained in Stage 2.  I had used the Stage 1 test set scores extensively for validation, and was already worried about overfitting the Stage 1 test set. </p>\n\n<h2>Interesting findings</h2>\n\n<p>I think my single most important observation was that the bounding boxes from my models were systematically too large.  I found that by reducing all bounding boxes by a fixed percentage (17% in each dimension), I improved my Stage 1 leaderboard score substantially.  I had actually first observed this when I had a larger validation set and I manually reviewed the predicted bounding boxes superimposed on the internal validation set images, and saw that they were generally too large.  Shrinking the bounding boxes improved both my internal validation set score and my Stage 1 leaderboard score.  This led me to believe that perhaps the L1 loss used by RetinaNet may not be optimal for the mean average precision metric used in this competition. Alternatively, the image rotations for augmentation may have led to a slight increase in bounding box size for training, though I think this effect is small, as I limited the maximum possible rotation to about 0.05 radians.</p>\n\n<p>Interestingly, however, when I switched to the smaller validation set, shrinking the bounding boxes actually reduced my internal validation set score, but still improved my Stage 1 leaderboard score.  It was then that I carefully read <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723\">Dr. Anouk Stein’s post</a> describing the annotation method for test set cases; these cases were multiread and “the intersection was used if there was at least a 50% overlap by one of the boxes”.  This would suggest that bounding boxes of the test set images could quite easily be systematically smaller in the test set compared to the training set.  On the assumption that Stage 2 test set bounding boxes would have statistical properties similar to the Stage 1 test set bounding boxes, I decided to treat the Stage 1 leaderboard scores as more accurate validation set scores than my own internal validation set scores.</p>\n\n<p>At the end the only parameters I was tuning were the score thresholds for each model, the score threshold for solitary unmatched bounding boxes, and the bounding box shrinkage factor.</p>\n\n<p>Combining two RetinaNet models in an ensemble gave a mild boost to my Stage 1 leaderboard score.  My best trained resnet-101 model by itself (with the usual postprocessing steps including non-maximum suppression and bounding box shrinkage) gave a Stage 1 leaderboard score of 0.232.   The ensemble had a Stage 1 leaderboard score of 0.238.  I think the ensemble improved specificity by eliminating most bounding boxes proposed by only 1 of the 2 networks, and this may have helped in Stage 2.  </p>\n\n<p>My stage 2 test set score (0.239) was practically identical to my Stage 1 test set score (0.238).</p>\n\n<h2>Concluding thoughts</h2>\n\n<p>The competition was an exciting and educational experience.  I thank the RSNA/STR organizers for all their hard work organizing and annotating the data sets for competition; large medical image data sets of sufficient size and quality for this purpose are rare.  Thanks also to Kaggle and its staff for their support of this competition.  </p>\n\n<p>Given the close spacing between the scores of many of the top teams (and the virtual tie between 3rd and 4th place), I'm sure a different final test set would have produced different rankings, though Ian/Alex and Dmytro clearly set themselves apart at the top.  I was surprised and lucky to finish so high in the rankings.  I know that I have a lot to learn about object detection based on the varied and interesting forum posts; my solution is fairly minimalist by comparison.  My congratulations and respect to all the participants!    </p>",
  "messages": [
    {
      "id": "416043",
      "postDate": "11/06/2018 04:29:04",
      "content": "<p>Hello!  This is an overview of my solution to the pneumonia detection challenge.  I'm an abdominal radiologist at the Keck School of Medicine of USC in Los Angeles, CA, USA.  Although I have played with convolutional neural networks for medical image classification, this is my first experience with object detection.  It is also my first experience with machine learning competitions.      </p>\n\n<h2>Summary</h2>\n\n<p>My models used <a href=\"https://github.com/fizyr/keras-retinanet\">keras-retinanet</a> by Hans Gaiser and collaborators, based on the <a href=\"https://arxiv.org/abs/1708.02002\">focal loss paper by Lin et al</a>.  This is a single-stage convolutional neural network detection architecture, which was appealing to me for training simplicity.  I optimized two RetinaNet models using Keras, with resnet-50 and resnet-101 backbones that were pretrained on ImageNet images.   I used non-maximum suppression to eliminate any overlapping bounding boxes from each network.  I then took weighted averages of overlapping bounding boxes from both trained neural networks.  I also applied a global fixed percentage size reduction to all final bounding boxes, which appeared to significantly improve Stage 1 test scores.</p>\n\n<p>Code is posted on <a href=\"https://github.com/pmcheng/rsna-pneumonia\">Github</a>.  Note that the two model .h5 files used in my solution are under the Releases tab of the repository, because these files are too large to be included within the repo.</p>\n\n<h2>Feature selection (or lack thereof)</h2>\n\n<p>I did not perform manual image feature selection or engineering.  I briefly played with using the DICOM view projection (AP vs PA), specifically assigning different score thresholds based on the view, but this did not improve my results.  </p>\n\n<p>I did not make use of the “No Lung Opacity / Not Normal” labels in the training set.</p>\n\n<p>I did not use any external training data for this competition.</p>\n\n<h2>Training</h2>\n\n<p>I decided early on that high image resolution was not necessary for pneumonia bounding box prediction.  I used the training images at a 224 x 224 resolution, which made training much more efficient on my hardware.  I used sklearn’s <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedKFold.html\">stratified K fold function</a> to divide the 25684 training images into a large training set (24399 images, 95%) and a small validation set (1285 images, 5%) .  I had originally used a much larger validation set but later found that I had significantly better training results when I shifted more images into training.  Images were augmented with rotation, translation, scaling, and horizontal flipping (shearing and vertical flipping were turned off).  I also added random constants to the images, an idea I found from a <a href=\"https://www.kaggle.com/c/google-ai-open-images-object-detection-track/discussion/64633\">post by @ZFTurbo</a> regarding RetinaNet in a previous competition.   </p>\n\n<p>The default training parameters were modified to match the focal loss paper (i.e., SGD with learning rate = 0.01, momentum =0.9, decay = 0.0001, nesterov = True), due to anecdotal reports about more generalizable models produced by SGD with momentum compared to Adam.  I used resnet-50 and resnet-101 backbones for the final models.  I did experiment with resnet-152, but the training took longer and the results were slightly worse.</p>\n\n<p>I trained for 25 epochs with 2500 steps per epoch and batch size of 8, saving the model snapshot after each epoch.  After each epoch I calculated bounding boxes on the validation set; I changed the score threshold in the filter_detections layer from keras-retinanet to 0.01 (from 0.05) so that I could evaluate lower score thresholds.  At each epoch I calculated the score threshold providing the maximum Youden index on the validation set (sensitivity + specificity – 1).  More specifically, I calculated sensitivity and specificity with respect to images, not bounding boxes, with the idea that from a scoring standpoint, it was important for the system to classify whether an image as a whole was positive or negative for pneumonia.  I also calculated the RSNA metric as implemented in <a href=\"https://www.kaggle.com/chenyc15/mean-average-precision-metric\">Yicheng Chen's excellent kernel</a>.  I found that snapshots that performed best on the leaderboard were ones with the highest Youden index, with the score threshold lowered to give a sensitivity close to 90%.  I suspect that the benefit of lowered score thresholds was due to higher prevalence of pneumonia in the test set relative to the training set. </p>\n\n<p>I aggressively used non-maximum suppression to eliminate any overlapping bounding boxes from each network’s output for a given image.  My idea was that physician annotators would most likely specify nonoverlapping ground truth bounding boxes.  Even though Tensorflow has an NMS function, I found it more useful and instructive to modify <a href=\"https://www.pyimagesearch.com/2015/02/16/faster-non-maximum-suppression-python\">code posted by Adrian Rosebrock on his blog</a>.</p>\n\n<p>I then took weighted averages of overlapping bounding boxes from both trained neural networks, using the scores for the boxes from each network as the weights; I experimented with box unions and intersections, which did not work nearly as well.  For bounding boxes that did not overlap between the two neural networks, I used a separate higher threshold value to decide whether the solitary box should be retained.  I also applied a global fixed percentage size reduction to all final bounding boxes, which significantly improved Stage 1 test scores.</p>\n\n<p>I did not train with Stage 1 test data for Stage 2, because I didn’t know this was an option (I did not upload automated training code).  However, even if I had been aware of this option, I doubt that I would have retrained in Stage 2.  I had used the Stage 1 test set scores extensively for validation, and was already worried about overfitting the Stage 1 test set. </p>\n\n<h2>Interesting findings</h2>\n\n<p>I think my single most important observation was that the bounding boxes from my models were systematically too large.  I found that by reducing all bounding boxes by a fixed percentage (17% in each dimension), I improved my Stage 1 leaderboard score substantially.  I had actually first observed this when I had a larger validation set and I manually reviewed the predicted bounding boxes superimposed on the internal validation set images, and saw that they were generally too large.  Shrinking the bounding boxes improved both my internal validation set score and my Stage 1 leaderboard score.  This led me to believe that perhaps the L1 loss used by RetinaNet may not be optimal for the mean average precision metric used in this competition. Alternatively, the image rotations for augmentation may have led to a slight increase in bounding box size for training, though I think this effect is small, as I limited the maximum possible rotation to about 0.05 radians.</p>\n\n<p>Interestingly, however, when I switched to the smaller validation set, shrinking the bounding boxes actually reduced my internal validation set score, but still improved my Stage 1 leaderboard score.  It was then that I carefully read <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723\">Dr. Anouk Stein’s post</a> describing the annotation method for test set cases; these cases were multiread and “the intersection was used if there was at least a 50% overlap by one of the boxes”.  This would suggest that bounding boxes of the test set images could quite easily be systematically smaller in the test set compared to the training set.  On the assumption that Stage 2 test set bounding boxes would have statistical properties similar to the Stage 1 test set bounding boxes, I decided to treat the Stage 1 leaderboard scores as more accurate validation set scores than my own internal validation set scores.</p>\n\n<p>At the end the only parameters I was tuning were the score thresholds for each model, the score threshold for solitary unmatched bounding boxes, and the bounding box shrinkage factor.</p>\n\n<p>Combining two RetinaNet models in an ensemble gave a mild boost to my Stage 1 leaderboard score.  My best trained resnet-101 model by itself (with the usual postprocessing steps including non-maximum suppression and bounding box shrinkage) gave a Stage 1 leaderboard score of 0.232.   The ensemble had a Stage 1 leaderboard score of 0.238.  I think the ensemble improved specificity by eliminating most bounding boxes proposed by only 1 of the 2 networks, and this may have helped in Stage 2.  </p>\n\n<p>My stage 2 test set score (0.239) was practically identical to my Stage 1 test set score (0.238).</p>\n\n<h2>Concluding thoughts</h2>\n\n<p>The competition was an exciting and educational experience.  I thank the RSNA/STR organizers for all their hard work organizing and annotating the data sets for competition; large medical image data sets of sufficient size and quality for this purpose are rare.  Thanks also to Kaggle and its staff for their support of this competition.  </p>\n\n<p>Given the close spacing between the scores of many of the top teams (and the virtual tie between 3rd and 4th place), I'm sure a different final test set would have produced different rankings, though Ian/Alex and Dmytro clearly set themselves apart at the top.  I was surprised and lucky to finish so high in the rankings.  I know that I have a lot to learn about object detection based on the varied and interesting forum posts; my solution is fairly minimalist by comparison.  My congratulations and respect to all the participants!    </p>",
      "rawMarkdown": "Hello!  This is an overview of my solution to the pneumonia detection challenge.  I'm an abdominal radiologist at the Keck School of Medicine of USC in Los Angeles, CA, USA.  Although I have played with convolutional neural networks for medical image classification, this is my first experience with object detection.  It is also my first experience with machine learning competitions.      \n\n## Summary ##\n\nMy models used [keras-retinanet](https://github.com/fizyr/keras-retinanet) by Hans Gaiser and collaborators, based on the [focal loss paper by Lin et al](https://arxiv.org/abs/1708.02002).  This is a single-stage convolutional neural network detection architecture, which was appealing to me for training simplicity.  I optimized two RetinaNet models using Keras, with resnet-50 and resnet-101 backbones that were pretrained on ImageNet images.   I used non-maximum suppression to eliminate any overlapping bounding boxes from each network.  I then took weighted averages of overlapping bounding boxes from both trained neural networks.  I also applied a global fixed percentage size reduction to all final bounding boxes, which appeared to significantly improve Stage 1 test scores.\n\nCode is posted on [Github](https://github.com/pmcheng/rsna-pneumonia).  Note that the two model .h5 files used in my solution are under the Releases tab of the repository, because these files are too large to be included within the repo.\n\n## Feature selection (or lack thereof) ##\n\nI did not perform manual image feature selection or engineering.  I briefly played with using the DICOM view projection (AP vs PA), specifically assigning different score thresholds based on the view, but this did not improve my results.  \n\nI did not make use of the “No Lung Opacity / Not Normal” labels in the training set.\n  \nI did not use any external training data for this competition.\n\n## Training ##\nI decided early on that high image resolution was not necessary for pneumonia bounding box prediction.  I used the training images at a 224 x 224 resolution, which made training much more efficient on my hardware.  I used sklearn’s [stratified K fold function](http://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedKFold.html) to divide the 25684 training images into a large training set (24399 images, 95%) and a small validation set (1285 images, 5%) .  I had originally used a much larger validation set but later found that I had significantly better training results when I shifted more images into training.  Images were augmented with rotation, translation, scaling, and horizontal flipping (shearing and vertical flipping were turned off).  I also added random constants to the images, an idea I found from a [post by @ZFTurbo](https://www.kaggle.com/c/google-ai-open-images-object-detection-track/discussion/64633) regarding RetinaNet in a previous competition.   \n\n\nThe default training parameters were modified to match the focal loss paper (i.e., SGD with learning rate = 0.01, momentum =0.9, decay = 0.0001, nesterov = True), due to anecdotal reports about more generalizable models produced by SGD with momentum compared to Adam.  I used resnet-50 and resnet-101 backbones for the final models.  I did experiment with resnet-152, but the training took longer and the results were slightly worse.\n\nI trained for 25 epochs with 2500 steps per epoch and batch size of 8, saving the model snapshot after each epoch.  After each epoch I calculated bounding boxes on the validation set; I changed the score threshold in the filter_detections layer from keras-retinanet to 0.01 (from 0.05) so that I could evaluate lower score thresholds.  At each epoch I calculated the score threshold providing the maximum Youden index on the validation set (sensitivity + specificity – 1).  More specifically, I calculated sensitivity and specificity with respect to images, not bounding boxes, with the idea that from a scoring standpoint, it was important for the system to classify whether an image as a whole was positive or negative for pneumonia.  I also calculated the RSNA metric as implemented in [Yicheng Chen's excellent kernel](https://www.kaggle.com/chenyc15/mean-average-precision-metric).  I found that snapshots that performed best on the leaderboard were ones with the highest Youden index, with the score threshold lowered to give a sensitivity close to 90%.  I suspect that the benefit of lowered score thresholds was due to higher prevalence of pneumonia in the test set relative to the training set. \n\nI aggressively used non-maximum suppression to eliminate any overlapping bounding boxes from each network’s output for a given image.  My idea was that physician annotators would most likely specify nonoverlapping ground truth bounding boxes.  Even though Tensorflow has an NMS function, I found it more useful and instructive to modify [code posted by Adrian Rosebrock on his blog](https://www.pyimagesearch.com/2015/02/16/faster-non-maximum-suppression-python).\n\nI then took weighted averages of overlapping bounding boxes from both trained neural networks, using the scores for the boxes from each network as the weights; I experimented with box unions and intersections, which did not work nearly as well.  For bounding boxes that did not overlap between the two neural networks, I used a separate higher threshold value to decide whether the solitary box should be retained.  I also applied a global fixed percentage size reduction to all final bounding boxes, which significantly improved Stage 1 test scores.\n\nI did not train with Stage 1 test data for Stage 2, because I didn’t know this was an option (I did not upload automated training code).  However, even if I had been aware of this option, I doubt that I would have retrained in Stage 2.  I had used the Stage 1 test set scores extensively for validation, and was already worried about overfitting the Stage 1 test set. \n \n## Interesting findings ##\nI think my single most important observation was that the bounding boxes from my models were systematically too large.  I found that by reducing all bounding boxes by a fixed percentage (17% in each dimension), I improved my Stage 1 leaderboard score substantially.  I had actually first observed this when I had a larger validation set and I manually reviewed the predicted bounding boxes superimposed on the internal validation set images, and saw that they were generally too large.  Shrinking the bounding boxes improved both my internal validation set score and my Stage 1 leaderboard score.  This led me to believe that perhaps the L1 loss used by RetinaNet may not be optimal for the mean average precision metric used in this competition. Alternatively, the image rotations for augmentation may have led to a slight increase in bounding box size for training, though I think this effect is small, as I limited the maximum possible rotation to about 0.05 radians.\n\nInterestingly, however, when I switched to the smaller validation set, shrinking the bounding boxes actually reduced my internal validation set score, but still improved my Stage 1 leaderboard score.  It was then that I carefully read [Dr. Anouk Stein’s post](https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723) describing the annotation method for test set cases; these cases were multiread and “the intersection was used if there was at least a 50% overlap by one of the boxes”.  This would suggest that bounding boxes of the test set images could quite easily be systematically smaller in the test set compared to the training set.  On the assumption that Stage 2 test set bounding boxes would have statistical properties similar to the Stage 1 test set bounding boxes, I decided to treat the Stage 1 leaderboard scores as more accurate validation set scores than my own internal validation set scores.\n  \nAt the end the only parameters I was tuning were the score thresholds for each model, the score threshold for solitary unmatched bounding boxes, and the bounding box shrinkage factor.\n\nCombining two RetinaNet models in an ensemble gave a mild boost to my Stage 1 leaderboard score.  My best trained resnet-101 model by itself (with the usual postprocessing steps including non-maximum suppression and bounding box shrinkage) gave a Stage 1 leaderboard score of 0.232.   The ensemble had a Stage 1 leaderboard score of 0.238.  I think the ensemble improved specificity by eliminating most bounding boxes proposed by only 1 of the 2 networks, and this may have helped in Stage 2.  \n\nMy stage 2 test set score (0.239) was practically identical to my Stage 1 test set score (0.238).\n\n\n## Concluding thoughts ##\n\nThe competition was an exciting and educational experience.  I thank the RSNA/STR organizers for all their hard work organizing and annotating the data sets for competition; large medical image data sets of sufficient size and quality for this purpose are rare.  Thanks also to Kaggle and its staff for their support of this competition.  \n\nGiven the close spacing between the scores of many of the top teams (and the virtual tie between 3rd and 4th place), I'm sure a different final test set would have produced different rankings, though Ian/Alex and Dmytro clearly set themselves apart at the top.  I was surprised and lucky to finish so high in the rankings.  I know that I have a lot to learn about object detection based on the varied and interesting forum posts; my solution is fairly minimalist by comparison.  My congratulations and respect to all the participants!",
      "votes": null
    },
    {
      "id": "416046",
      "postDate": "11/06/2018 04:40:24",
      "content": "<p>Congratulations Phillip ! Interesting to see that domain expertise both oriented us to a similar post-processing technique. Anouk Stein's post was definitely critical.</p>",
      "rawMarkdown": "Congratulations Phillip ! Interesting to see that domain expertise both oriented us to a similar post-processing technique. Anouk Stein's post was definitely critical.",
      "votes": null
    },
    {
      "id": "416067",
      "postDate": "11/06/2018 05:58:52",
      "content": "<p>Congrats and thanks for sharing.</p>",
      "rawMarkdown": "Congrats and thanks for sharing.",
      "votes": null
    },
    {
      "id": "416103",
      "postDate": "11/06/2018 07:42:52",
      "content": "<p>Fantastic write up, thank you and congratulations! All 3 top solutions seem to share the use of Retinanet + scaling of the final boxes. Some clever thinking to understand the result of how annotations were created. </p>",
      "rawMarkdown": "Fantastic write up, thank you and congratulations! All 3 top solutions seem to share the use of Retinanet + scaling of the final boxes. Some clever thinking to understand the result of how annotations were created.",
      "votes": null
    },
    {
      "id": "416105",
      "postDate": "11/06/2018 07:44:31",
      "content": "<p>thanks for sharing your code </p>",
      "rawMarkdown": "thanks for sharing your code",
      "votes": null
    },
    {
      "id": "416114",
      "postDate": "11/06/2018 08:00:17",
      "content": "<p>Congrats and thanks for sharing.</p>",
      "rawMarkdown": "Congrats and thanks for sharing.",
      "votes": null
    },
    {
      "id": "416170",
      "postDate": "11/06/2018 10:05:16",
      "content": "<p>Congratulations.\nWhat an excellent write-up!\nThanks for all extra-ordinary thinking and sharing the code.</p>",
      "rawMarkdown": "Congratulations.\nWhat an excellent write-up!\nThanks for all extra-ordinary thinking and sharing the code.",
      "votes": null
    },
    {
      "id": "416227",
      "postDate": "11/06/2018 12:38:01",
      "content": "<p>Congratulations Phillip! Interesting finding about metrics. I had the similar concern about rotations to increase box sizes, to partially reduce it I rotated two points closer to center on each edge (8 in total) instead of corners and re-calculated bounding box.</p>",
      "rawMarkdown": "Congratulations Phillip! Interesting finding about metrics. I had the similar concern about rotations to increase box sizes, to partially reduce it I rotated two points closer to center on each edge (8 in total) instead of corners and re-calculated bounding box.",
      "votes": null
    },
    {
      "id": "416651",
      "postDate": "11/07/2018 03:23:07",
      "content": "<p>That's an interesting solution!  I'm looking forward to studying your code.  I haven't seen the problem of reduced accuracy of rotated bounding boxes addressed elsewhere.</p>",
      "rawMarkdown": "That's an interesting solution!  I'm looking forward to studying your code.  I haven't seen the problem of reduced accuracy of rotated bounding boxes addressed elsewhere.",
      "votes": null
    },
    {
      "id": "416816",
      "postDate": "11/07/2018 10:13:06",
      "content": "<p>Congrats on your solo gold medal!</p>",
      "rawMarkdown": "Congrats on your solo gold medal!",
      "votes": null
    },
    {
      "id": "417250",
      "postDate": "11/08/2018 01:39:29",
      "content": "<p>Thanks for your sharing.I forked your code and run it myself,I trained retinanet with resnet-50 backbone,but the score is always 0,is there something I did it wrong?</p>",
      "rawMarkdown": "Thanks for your sharing.I forked your code and run it myself,I trained retinanet with resnet-50 backbone,but the score is always 0,is there something I did it wrong?",
      "votes": null
    },
    {
      "id": "417280",
      "postDate": "11/08/2018 03:04:52",
      "content": "<p>Hm, that shouldn't be happening.  I'm assuming you ran <code>prepare_data.py</code> and are running the <code>train_50.sh</code> script?  Can you send me the output?</p>",
      "rawMarkdown": "Hm, that shouldn't be happening.  I'm assuming you ran `prepare_data.py` and are running the `train_50.sh` script?  Can you send me the output?",
      "votes": null
    },
    {
      "id": "417379",
      "postDate": "11/08/2018 07:02:30",
      "content": "<p>Yes.I trained with train_50.sh.During training the output seems normal,like this:\nEpoch 1/25\n 244/2500 [=&gt;............................] - ETA: 14:18 - loss: 3.4960 - regression_loss: 2.2909 - classification_loss: 1.2051</p>",
      "rawMarkdown": "Yes.I trained with train_50.sh.During training the output seems normal,like this:\nEpoch 1/25\n 244/2500 [=&gt;............................] - ETA: 14:18 - loss: 3.4960 - regression_loss: 2.2909 - classification_loss: 1.2051",
      "votes": null
    },
    {
      "id": "417382",
      "postDate": "11/08/2018 07:11:57",
      "content": "<p>And in prepare_data.py line 49,original code is \nconversion = sz/1000\nmaybe  it should be conversion = sz/1024.0,the original image size is 1024*1024</p>",
      "rawMarkdown": "And in prepare_data.py line 49,original code is \nconversion = sz/1000\nmaybe  it should be conversion = sz/1024.0,the original image size is 1024*1024",
      "votes": null
    },
    {
      "id": "417391",
      "postDate": "11/08/2018 07:20:31",
      "content": "<p>During evaluation,the output is\nThresh: 0  Score: 0  Se: 0 Sp: 0 Youden: 0</p>",
      "rawMarkdown": "During evaluation,the output is\nThresh: 0  Score: 0  Se: 0 Sp: 0 Youden: 0",
      "votes": null
    },
    {
      "id": "417611",
      "postDate": "11/08/2018 14:17:44",
      "content": "<p>Thanks.  The conversion in <code>prepare_data</code> should indeed probably be 1024, for some reason I converted to size 1000 jpg from the very start.  I'm going to leave it as 1000 in the code since that's the way I converted my data.  It probably doesn't make a big difference since the images are later used at size 224 resolution anyway.  I assume the converted jpg images look ok?</p>\n\n<p>I'm not sure why your outputs are all 0's, I get nonzero outputs even after the first epoch of training.    Are you using python 3 (I used python 3.6.6 from Anaconda distribution)?  Could you check whether you can predict bounding boxes using <code>predict.py</code> and my trained .h5 models (located on resources tab of Github repo)?</p>",
      "rawMarkdown": "Thanks.  The conversion in `prepare_data` should indeed probably be 1024, for some reason I converted to size 1000 jpg from the very start.  I'm going to leave it as 1000 in the code since that's the way I converted my data.  It probably doesn't make a big difference since the images are later used at size 224 resolution anyway.  I assume the converted jpg images look ok?\n\nI'm not sure why your outputs are all 0's, I get nonzero outputs even after the first epoch of training.    Are you using python 3 (I used python 3.6.6 from Anaconda distribution)?  Could you check whether you can predict bounding boxes using `predict.py` and my trained .h5 models (located on resources tab of Github repo)?",
      "votes": null
    },
    {
      "id": "417617",
      "postDate": "11/08/2018 14:26:47",
      "content": "<p>OK.I will do a further check and visulize the predicted box</p>",
      "rawMarkdown": "OK.I will do a further check and visulize the predicted box",
      "votes": null
    },
    {
      "id": "417930",
      "postDate": "11/09/2018 02:07:57",
      "content": "<p>I'm using python 2.7 without anaconda</p>",
      "rawMarkdown": "I'm using python 2.7 without anaconda",
      "votes": null
    },
    {
      "id": "417967",
      "postDate": "11/09/2018 04:09:08",
      "content": "<p>Python 2 is the problem, I know I used python 3 style division in my code, so an integer divided by an integer gives a float (in python 2 you get an integer).  I think you can import python 3 style division into python 2, but I don't have experience with this.  I used f-strings in the code as well, which I think is a python 3 feature.</p>",
      "rawMarkdown": "Python 2 is the problem, I know I used python 3 style division in my code, so an integer divided by an integer gives a float (in python 2 you get an integer).  I think you can import python 3 style division into python 2, but I don't have experience with this.  I used f-strings in the code as well, which I think is a python 3 feature.",
      "votes": null
    },
    {
      "id": "418101",
      "postDate": "11/09/2018 09:36:52",
      "content": "<p>I tried python3.6 with annaconda,this time the output is not zeros from the first epoch.I think it's normal now.I will train it with more epochs.\nThanks for your reply.</p>",
      "rawMarkdown": "I tried python3.6 with annaconda,this time the output is not zeros from the first epoch.I think it's normal now.I will train it with more epochs.\nThanks for your reply.",
      "votes": null
    },
    {
      "id": "429100",
      "postDate": "11/28/2018 10:42:41",
      "content": "<p>Congrats and thanks for sharing your solution! I have taken a look at your code and when I understand it correctly you just use the standard keras-retinanet preprocessing which consists in subtracting the imagenet means for three channels. Does this make sense for gray scale images when all three channels are identical and have you tried other preprocessing steps?</p>",
      "rawMarkdown": "Congrats and thanks for sharing your solution! I have taken a look at your code and when I understand it correctly you just use the standard keras-retinanet preprocessing which consists in subtracting the imagenet means for three channels. Does this make sense for gray scale images when all three channels are identical and have you tried other preprocessing steps?",
      "votes": null
    },
    {
      "id": "429631",
      "postDate": "11/29/2018 05:04:35",
      "content": "<p>Yes, I just used standard preprocessing.  I agree that it's not clear this would be optimal for grayscale images. However, when using pretrained models, each channel is treated differently by the network, and I imagine we might leverage a kind of ensemble effect across the channels.  I have had success in the past with classification models using similar Imagenet-style preprocessing of grayscale images.  It would be interesting to experiment with other preprocessing steps, but I didn't have time.</p>",
      "rawMarkdown": "Yes, I just used standard preprocessing.  I agree that it's not clear this would be optimal for grayscale images. However, when using pretrained models, each channel is treated differently by the network, and I imagine we might leverage a kind of ensemble effect across the channels.  I have had success in the past with classification models using similar Imagenet-style preprocessing of grayscale images.  It would be interesting to experiment with other preprocessing steps, but I didn't have time.",
      "votes": null
    },
    {
      "id": "430322",
      "postDate": "11/30/2018 07:51:12",
      "content": "<p>stage2-test    the predicted box   ：1c6cdd43-a38a-4464-b012-bc0eb9e73fd0    {scores[j]:.3f} {x1} {y1} {w} {h} {scores[j]:.3f} {x1} {y1} {w} {h} \nno  predicted   data  ,??????????????      why？   Phillip   ，I want to your helping,thank you</p>",
      "rawMarkdown": "stage2-test    the predicted box   ：1c6cdd43-a38a-4464-b012-bc0eb9e73fd0\t{scores[j]:.3f} {x1} {y1} {w} {h} {scores[j]:.3f} {x1} {y1} {w} {h} \nno  predicted   data  ,??????????????      why？   Phillip   ，I want to your helping,thank you",
      "votes": null
    },
    {
      "id": "440310",
      "postDate": "12/17/2018 11:37:15",
      "content": "<p>Congratulations Phillip! Why is the loss value great after the training is completed?</p>",
      "rawMarkdown": "Congratulations Phillip! Why is the loss value great after the training is completed?",
      "votes": null
    },
    {
      "id": "885973",
      "postDate": "06/14/2020 15:57:19",
      "content": "<p>Hello! I tried using your code to test on my own dataset (two-dimensional CT image), but my regression_loss is always zero.I know my question is very stupid, but I do not understand the reason, I have generated csv in the correct way.If you have time, I need your help.pls!!!</p>",
      "rawMarkdown": "Hello! I tried using your code to test on my own dataset (two-dimensional CT image), but my regression_loss is always zero.I know my question is very stupid, but I do not understand the reason, I have generated csv in the correct way.If you have time, I need your help.pls!!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 416046,
      "author_name": "alexandrecc",
      "author_url": "",
      "post_date": "11/06/2018 04:40:24",
      "content": "<p>Congratulations Phillip ! Interesting to see that domain expertise both oriented us to a similar post-processing technique. Anouk Stein's post was definitely critical.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 416067,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "11/06/2018 05:58:52",
      "content": "<p>Congrats and thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 416103,
      "author_name": "taindow",
      "author_url": "",
      "post_date": "11/06/2018 07:42:52",
      "content": "<p>Fantastic write up, thank you and congratulations! All 3 top solutions seem to share the use of Retinanet + scaling of the final boxes. Some clever thinking to understand the result of how annotations were created. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 416105,
      "author_name": "chenglong0313",
      "author_url": "",
      "post_date": "11/06/2018 07:44:31",
      "content": "<p>thanks for sharing your code </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 416114,
      "author_name": "raytroop",
      "author_url": "",
      "post_date": "11/06/2018 08:00:17",
      "content": "<p>Congrats and thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 416170,
      "author_name": "muhammedazamkhan",
      "author_url": "",
      "post_date": "11/06/2018 10:05:16",
      "content": "<p>Congratulations.\nWhat an excellent write-up!\nThanks for all extra-ordinary thinking and sharing the code.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 416227,
      "author_name": "dmytropoplavskiy",
      "author_url": "",
      "post_date": "11/06/2018 12:38:01",
      "content": "<p>Congratulations Phillip! Interesting finding about metrics. I had the similar concern about rotations to increase box sizes, to partially reduce it I rotated two points closer to center on each edge (8 in total) instead of corners and re-calculated bounding box.</p>",
      "votes": null,
      "replies": [
        {
          "id": 416651,
          "author_name": "pmcheng",
          "author_url": "",
          "post_date": "11/07/2018 03:23:07",
          "content": "<p>That's an interesting solution!  I'm looking forward to studying your code.  I haven't seen the problem of reduced accuracy of rotated bounding boxes addressed elsewhere.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 416816,
      "author_name": "felipekitamura",
      "author_url": "",
      "post_date": "11/07/2018 10:13:06",
      "content": "<p>Congrats on your solo gold medal!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 417250,
      "author_name": "pgb123",
      "author_url": "",
      "post_date": "11/08/2018 01:39:29",
      "content": "<p>Thanks for your sharing.I forked your code and run it myself,I trained retinanet with resnet-50 backbone,but the score is always 0,is there something I did it wrong?</p>",
      "votes": null,
      "replies": [
        {
          "id": 417280,
          "author_name": "pmcheng",
          "author_url": "",
          "post_date": "11/08/2018 03:04:52",
          "content": "<p>Hm, that shouldn't be happening.  I'm assuming you ran <code>prepare_data.py</code> and are running the <code>train_50.sh</code> script?  Can you send me the output?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417379,
          "author_name": "pgb123",
          "author_url": "",
          "post_date": "11/08/2018 07:02:30",
          "content": "<p>Yes.I trained with train_50.sh.During training the output seems normal,like this:\nEpoch 1/25\n 244/2500 [=&gt;............................] - ETA: 14:18 - loss: 3.4960 - regression_loss: 2.2909 - classification_loss: 1.2051</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417382,
          "author_name": "pgb123",
          "author_url": "",
          "post_date": "11/08/2018 07:11:57",
          "content": "<p>And in prepare_data.py line 49,original code is \nconversion = sz/1000\nmaybe  it should be conversion = sz/1024.0,the original image size is 1024*1024</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417391,
          "author_name": "pgb123",
          "author_url": "",
          "post_date": "11/08/2018 07:20:31",
          "content": "<p>During evaluation,the output is\nThresh: 0  Score: 0  Se: 0 Sp: 0 Youden: 0</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417611,
          "author_name": "pmcheng",
          "author_url": "",
          "post_date": "11/08/2018 14:17:44",
          "content": "<p>Thanks.  The conversion in <code>prepare_data</code> should indeed probably be 1024, for some reason I converted to size 1000 jpg from the very start.  I'm going to leave it as 1000 in the code since that's the way I converted my data.  It probably doesn't make a big difference since the images are later used at size 224 resolution anyway.  I assume the converted jpg images look ok?</p>\n\n<p>I'm not sure why your outputs are all 0's, I get nonzero outputs even after the first epoch of training.    Are you using python 3 (I used python 3.6.6 from Anaconda distribution)?  Could you check whether you can predict bounding boxes using <code>predict.py</code> and my trained .h5 models (located on resources tab of Github repo)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417617,
          "author_name": "pgb123",
          "author_url": "",
          "post_date": "11/08/2018 14:26:47",
          "content": "<p>OK.I will do a further check and visulize the predicted box</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417930,
          "author_name": "pgb123",
          "author_url": "",
          "post_date": "11/09/2018 02:07:57",
          "content": "<p>I'm using python 2.7 without anaconda</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417967,
          "author_name": "pmcheng",
          "author_url": "",
          "post_date": "11/09/2018 04:09:08",
          "content": "<p>Python 2 is the problem, I know I used python 3 style division in my code, so an integer divided by an integer gives a float (in python 2 you get an integer).  I think you can import python 3 style division into python 2, but I don't have experience with this.  I used f-strings in the code as well, which I think is a python 3 feature.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 418101,
          "author_name": "pgb123",
          "author_url": "",
          "post_date": "11/09/2018 09:36:52",
          "content": "<p>I tried python3.6 with annaconda,this time the output is not zeros from the first epoch.I think it's normal now.I will train it with more epochs.\nThanks for your reply.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 429100,
      "author_name": "makemate",
      "author_url": "",
      "post_date": "11/28/2018 10:42:41",
      "content": "<p>Congrats and thanks for sharing your solution! I have taken a look at your code and when I understand it correctly you just use the standard keras-retinanet preprocessing which consists in subtracting the imagenet means for three channels. Does this make sense for gray scale images when all three channels are identical and have you tried other preprocessing steps?</p>",
      "votes": null,
      "replies": [
        {
          "id": 429631,
          "author_name": "pmcheng",
          "author_url": "",
          "post_date": "11/29/2018 05:04:35",
          "content": "<p>Yes, I just used standard preprocessing.  I agree that it's not clear this would be optimal for grayscale images. However, when using pretrained models, each channel is treated differently by the network, and I imagine we might leverage a kind of ensemble effect across the channels.  I have had success in the past with classification models using similar Imagenet-style preprocessing of grayscale images.  It would be interesting to experiment with other preprocessing steps, but I didn't have time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 430322,
      "author_name": "amose520",
      "author_url": "",
      "post_date": "11/30/2018 07:51:12",
      "content": "<p>stage2-test    the predicted box   ：1c6cdd43-a38a-4464-b012-bc0eb9e73fd0    {scores[j]:.3f} {x1} {y1} {w} {h} {scores[j]:.3f} {x1} {y1} {w} {h} \nno  predicted   data  ,??????????????      why？   Phillip   ，I want to your helping,thank you</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 440310,
      "author_name": "liumao",
      "author_url": "",
      "post_date": "12/17/2018 11:37:15",
      "content": "<p>Congratulations Phillip! Why is the loss value great after the training is completed?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 885973,
      "author_name": "lpx2040",
      "author_url": "",
      "post_date": "06/14/2020 15:57:19",
      "content": "<p>Hello! I tried using your code to test on my own dataset (two-dimensional CT image), but my regression_loss is always zero.I know my question is very stupid, but I do not understand the reason, I have generated csv in the correct way.If you have time, I need your help.pls!!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "416043": "Hello!  This is an overview of my solution to the pneumonia detection challenge.  I'm an abdominal radiologist at the Keck School of Medicine of USC in Los Angeles, CA, USA.  Although I have played with convolutional neural networks for medical image classification, this is my first experience with object detection.  It is also my first experience with machine learning competitions.      \n\n## Summary ##\n\nMy models used [keras-retinanet](https://github.com/fizyr/keras-retinanet) by Hans Gaiser and collaborators, based on the [focal loss paper by Lin et al](https://arxiv.org/abs/1708.02002).  This is a single-stage convolutional neural network detection architecture, which was appealing to me for training simplicity.  I optimized two RetinaNet models using Keras, with resnet-50 and resnet-101 backbones that were pretrained on ImageNet images.   I used non-maximum suppression to eliminate any overlapping bounding boxes from each network.  I then took weighted averages of overlapping bounding boxes from both trained neural networks.  I also applied a global fixed percentage size reduction to all final bounding boxes, which appeared to significantly improve Stage 1 test scores.\n\nCode is posted on [Github](https://github.com/pmcheng/rsna-pneumonia).  Note that the two model .h5 files used in my solution are under the Releases tab of the repository, because these files are too large to be included within the repo.\n\n## Feature selection (or lack thereof) ##\n\nI did not perform manual image feature selection or engineering.  I briefly played with using the DICOM view projection (AP vs PA), specifically assigning different score thresholds based on the view, but this did not improve my results.  \n\nI did not make use of the “No Lung Opacity / Not Normal” labels in the training set.\n  \nI did not use any external training data for this competition.\n\n## Training ##\nI decided early on that high image resolution was not necessary for pneumonia bounding box prediction.  I used the training images at a 224 x 224 resolution, which made training much more efficient on my hardware.  I used sklearn’s [stratified K fold function](http://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedKFold.html) to divide the 25684 training images into a large training set (24399 images, 95%) and a small validation set (1285 images, 5%) .  I had originally used a much larger validation set but later found that I had significantly better training results when I shifted more images into training.  Images were augmented with rotation, translation, scaling, and horizontal flipping (shearing and vertical flipping were turned off).  I also added random constants to the images, an idea I found from a [post by @ZFTurbo](https://www.kaggle.com/c/google-ai-open-images-object-detection-track/discussion/64633) regarding RetinaNet in a previous competition.   \n\n\nThe default training parameters were modified to match the focal loss paper (i.e., SGD with learning rate = 0.01, momentum =0.9, decay = 0.0001, nesterov = True), due to anecdotal reports about more generalizable models produced by SGD with momentum compared to Adam.  I used resnet-50 and resnet-101 backbones for the final models.  I did experiment with resnet-152, but the training took longer and the results were slightly worse.\n\nI trained for 25 epochs with 2500 steps per epoch and batch size of 8, saving the model snapshot after each epoch.  After each epoch I calculated bounding boxes on the validation set; I changed the score threshold in the filter_detections layer from keras-retinanet to 0.01 (from 0.05) so that I could evaluate lower score thresholds.  At each epoch I calculated the score threshold providing the maximum Youden index on the validation set (sensitivity + specificity – 1).  More specifically, I calculated sensitivity and specificity with respect to images, not bounding boxes, with the idea that from a scoring standpoint, it was important for the system to classify whether an image as a whole was positive or negative for pneumonia.  I also calculated the RSNA metric as implemented in [Yicheng Chen's excellent kernel](https://www.kaggle.com/chenyc15/mean-average-precision-metric).  I found that snapshots that performed best on the leaderboard were ones with the highest Youden index, with the score threshold lowered to give a sensitivity close to 90%.  I suspect that the benefit of lowered score thresholds was due to higher prevalence of pneumonia in the test set relative to the training set. \n\nI aggressively used non-maximum suppression to eliminate any overlapping bounding boxes from each network’s output for a given image.  My idea was that physician annotators would most likely specify nonoverlapping ground truth bounding boxes.  Even though Tensorflow has an NMS function, I found it more useful and instructive to modify [code posted by Adrian Rosebrock on his blog](https://www.pyimagesearch.com/2015/02/16/faster-non-maximum-suppression-python).\n\nI then took weighted averages of overlapping bounding boxes from both trained neural networks, using the scores for the boxes from each network as the weights; I experimented with box unions and intersections, which did not work nearly as well.  For bounding boxes that did not overlap between the two neural networks, I used a separate higher threshold value to decide whether the solitary box should be retained.  I also applied a global fixed percentage size reduction to all final bounding boxes, which significantly improved Stage 1 test scores.\n\nI did not train with Stage 1 test data for Stage 2, because I didn’t know this was an option (I did not upload automated training code).  However, even if I had been aware of this option, I doubt that I would have retrained in Stage 2.  I had used the Stage 1 test set scores extensively for validation, and was already worried about overfitting the Stage 1 test set. \n \n## Interesting findings ##\nI think my single most important observation was that the bounding boxes from my models were systematically too large.  I found that by reducing all bounding boxes by a fixed percentage (17% in each dimension), I improved my Stage 1 leaderboard score substantially.  I had actually first observed this when I had a larger validation set and I manually reviewed the predicted bounding boxes superimposed on the internal validation set images, and saw that they were generally too large.  Shrinking the bounding boxes improved both my internal validation set score and my Stage 1 leaderboard score.  This led me to believe that perhaps the L1 loss used by RetinaNet may not be optimal for the mean average precision metric used in this competition. Alternatively, the image rotations for augmentation may have led to a slight increase in bounding box size for training, though I think this effect is small, as I limited the maximum possible rotation to about 0.05 radians.\n\nInterestingly, however, when I switched to the smaller validation set, shrinking the bounding boxes actually reduced my internal validation set score, but still improved my Stage 1 leaderboard score.  It was then that I carefully read [Dr. Anouk Stein’s post](https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723) describing the annotation method for test set cases; these cases were multiread and “the intersection was used if there was at least a 50% overlap by one of the boxes”.  This would suggest that bounding boxes of the test set images could quite easily be systematically smaller in the test set compared to the training set.  On the assumption that Stage 2 test set bounding boxes would have statistical properties similar to the Stage 1 test set bounding boxes, I decided to treat the Stage 1 leaderboard scores as more accurate validation set scores than my own internal validation set scores.\n  \nAt the end the only parameters I was tuning were the score thresholds for each model, the score threshold for solitary unmatched bounding boxes, and the bounding box shrinkage factor.\n\nCombining two RetinaNet models in an ensemble gave a mild boost to my Stage 1 leaderboard score.  My best trained resnet-101 model by itself (with the usual postprocessing steps including non-maximum suppression and bounding box shrinkage) gave a Stage 1 leaderboard score of 0.232.   The ensemble had a Stage 1 leaderboard score of 0.238.  I think the ensemble improved specificity by eliminating most bounding boxes proposed by only 1 of the 2 networks, and this may have helped in Stage 2.  \n\nMy stage 2 test set score (0.239) was practically identical to my Stage 1 test set score (0.238).\n\n\n## Concluding thoughts ##\n\nThe competition was an exciting and educational experience.  I thank the RSNA/STR organizers for all their hard work organizing and annotating the data sets for competition; large medical image data sets of sufficient size and quality for this purpose are rare.  Thanks also to Kaggle and its staff for their support of this competition.  \n\nGiven the close spacing between the scores of many of the top teams (and the virtual tie between 3rd and 4th place), I'm sure a different final test set would have produced different rankings, though Ian/Alex and Dmytro clearly set themselves apart at the top.  I was surprised and lucky to finish so high in the rankings.  I know that I have a lot to learn about object detection based on the varied and interesting forum posts; my solution is fairly minimalist by comparison.  My congratulations and respect to all the participants!",
    "416046": "Congratulations Phillip ! Interesting to see that domain expertise both oriented us to a similar post-processing technique. Anouk Stein's post was definitely critical.",
    "416067": "Congrats and thanks for sharing.",
    "416103": "Fantastic write up, thank you and congratulations! All 3 top solutions seem to share the use of Retinanet + scaling of the final boxes. Some clever thinking to understand the result of how annotations were created.",
    "416105": "thanks for sharing your code",
    "416114": "Congrats and thanks for sharing.",
    "416170": "Congratulations.\nWhat an excellent write-up!\nThanks for all extra-ordinary thinking and sharing the code.",
    "416227": "Congratulations Phillip! Interesting finding about metrics. I had the similar concern about rotations to increase box sizes, to partially reduce it I rotated two points closer to center on each edge (8 in total) instead of corners and re-calculated bounding box.",
    "416651": "That's an interesting solution!  I'm looking forward to studying your code.  I haven't seen the problem of reduced accuracy of rotated bounding boxes addressed elsewhere.",
    "416816": "Congrats on your solo gold medal!",
    "417250": "Thanks for your sharing.I forked your code and run it myself,I trained retinanet with resnet-50 backbone,but the score is always 0,is there something I did it wrong?",
    "417280": "Hm, that shouldn't be happening.  I'm assuming you ran `prepare_data.py` and are running the `train_50.sh` script?  Can you send me the output?",
    "417379": "Yes.I trained with train_50.sh.During training the output seems normal,like this:\nEpoch 1/25\n 244/2500 [=&gt;............................] - ETA: 14:18 - loss: 3.4960 - regression_loss: 2.2909 - classification_loss: 1.2051",
    "417382": "And in prepare_data.py line 49,original code is \nconversion = sz/1000\nmaybe  it should be conversion = sz/1024.0,the original image size is 1024*1024",
    "417391": "During evaluation,the output is\nThresh: 0  Score: 0  Se: 0 Sp: 0 Youden: 0",
    "417611": "Thanks.  The conversion in `prepare_data` should indeed probably be 1024, for some reason I converted to size 1000 jpg from the very start.  I'm going to leave it as 1000 in the code since that's the way I converted my data.  It probably doesn't make a big difference since the images are later used at size 224 resolution anyway.  I assume the converted jpg images look ok?\n\nI'm not sure why your outputs are all 0's, I get nonzero outputs even after the first epoch of training.    Are you using python 3 (I used python 3.6.6 from Anaconda distribution)?  Could you check whether you can predict bounding boxes using `predict.py` and my trained .h5 models (located on resources tab of Github repo)?",
    "417617": "OK.I will do a further check and visulize the predicted box",
    "417930": "I'm using python 2.7 without anaconda",
    "417967": "Python 2 is the problem, I know I used python 3 style division in my code, so an integer divided by an integer gives a float (in python 2 you get an integer).  I think you can import python 3 style division into python 2, but I don't have experience with this.  I used f-strings in the code as well, which I think is a python 3 feature.",
    "418101": "I tried python3.6 with annaconda,this time the output is not zeros from the first epoch.I think it's normal now.I will train it with more epochs.\nThanks for your reply.",
    "429100": "Congrats and thanks for sharing your solution! I have taken a look at your code and when I understand it correctly you just use the standard keras-retinanet preprocessing which consists in subtracting the imagenet means for three channels. Does this make sense for gray scale images when all three channels are identical and have you tried other preprocessing steps?",
    "429631": "Yes, I just used standard preprocessing.  I agree that it's not clear this would be optimal for grayscale images. However, when using pretrained models, each channel is treated differently by the network, and I imagine we might leverage a kind of ensemble effect across the channels.  I have had success in the past with classification models using similar Imagenet-style preprocessing of grayscale images.  It would be interesting to experiment with other preprocessing steps, but I didn't have time.",
    "430322": "stage2-test    the predicted box   ：1c6cdd43-a38a-4464-b012-bc0eb9e73fd0\t{scores[j]:.3f} {x1} {y1} {w} {h} {scores[j]:.3f} {x1} {y1} {w} {h} \nno  predicted   data  ,??????????????      why？   Phillip   ，I want to your helping,thank you",
    "440310": "Congratulations Phillip! Why is the loss value great after the training is completed?",
    "885973": "Hello! I tried using your code to test on my own dataset (two-dimensional CT image), but my regression_loss is always zero.I know my question is very stupid, but I do not understand the reason, I have generated csv in the correct way.If you have time, I need your help.pls!!!"
  },
  "source": "meta"
}