{
  "id": 70421,
  "title": "[1st place] Solution Overview & Code",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/70421",
  "author_name": "Ian Pan",
  "post_date": "2018-11-03T12:44:48.531000",
  "votes": 161,
  "comment_count": 39,
  "views": 0,
  "content": "<p>We are very excited to have been awarded 1st place in this competition. This was our first foray into the world of object detection. I'm Ian, currently a 3rd-year medical student at Brown University, Providence, RI, USA. My teammate Alexandre is a radiologist practicing in Lanaudière, Québec, Canada. We are happy to represent the medical community in this competition. Huge congratulations to Dmytro for becoming Grandmaster. His top score during the public LB really motivated us to make successive improvements to our model. Truly an honor to be in the same ranks as him for this challenge. Congratulations to all of the other top 10 winners as well. </p>\n\n<h1>tl;dr</h1>\n\n<ul>\n<li>We did NOT retrain when stage 1 test labels were released</li>\n<li>We used a classification-detection pipeline </li>\n<li>10-fold CV ensemble for classification, combination of 5 10-fold CV ensembles for detection (50 models) </li>\n<li>For detection, we used:\n<ol><li>RetinaNet: <a href=\"https://github.com/fizyr/keras-retinanet\"></a><a href=\"https://github.com/fizyr/keras-retinanet\">https://github.com/fizyr/keras-retinanet</a></li>\n<li>Deformable R-FCN: <a href=\"https://github.com/msracver/Deformable-ConvNets\"></a><a href=\"https://github.com/msracver/Deformable-ConvNets\">https://github.com/msracver/Deformable-ConvNets</a></li>\n<li>Deformable Relation Networks: <a href=\"https://github.com/msracver/Relation-Networks-for-Object-Detection\"></a><a href=\"https://github.com/msracver/Relation-Networks-for-Object-Detection\">https://github.com/msracver/Relation-Networks-for-Object-Detection</a></li></ol></li>\n<li>Boxes were ensembled using: <a href=\"https://github.com/ahrnbom/ensemble-objdet\">https://github.com/ahrnbom/ensemble-objdet</a></li>\n<li><em>We resized the lengths and widths of our final predictions by 87.5%</em></li>\n<li>Code available here: <a href=\"https://www.github.com/i-pan/kaggle-rsna18\">https://www.github.com/i-pan/kaggle-rsna18</a></li>\n<li>We found that a 6-model ensemble with 1 InceptionResNetV2 for classification and 5 Deformable Relation Networks achieved 0.253 on stage 2 private LB</li>\n</ul>\n\n<h1>Classification</h1>\n\n<p>We used Keras 2.2 for classification. Our classification model was composed of the following individual models [format: modelArchitecture (numClasses) (imgSize)]. Models were trained on either 2 classes (opacity vs. not) or 3 classes (opacity vs. not normal/no opacity vs. normal). Each model was trained on a different fold.\n- InceptionResNetV2 (2) (256), InceptionResNetV2 (2) (320) \n- InceptionResNetV2 (3) (256), InceptionResNetV2 (3) (320) \n- Xception (2) (384), Xception (2) (448) \n- Xception (3) (384), Xception (3) (448) \n- DenseNet169 (2) (512), DenseNet169 (3) (512)\nFirst, we trained ImageNet pre-trained networks on the NIH ChestX-ray14 dataset using 15 classes (14 findings + abnormal vs. normal). Then we fine-tuned those weights on the pneumonia dataset. This improved results from training using ImageNet weights only by about 1% locally. We were getting about 0.88-0.90 AUC across our folds. </p>\n\n<p>Models were trained with 50% probability of being color-inverted, 50% of being flipped, and 50% of being augmented with some other augmentation (e.g. contrast enhancement, crop, rotation). We used 15x TTA per model for our final predictions and averaged the 150 predictions for our final classification score. When the stage 1 test labels were released, our classification ensemble had an AUC of 0.93. </p>\n\n<p>As many other competitors noticed, the distribution of the training (single-read) data and test (triple-read) data were quite different. We think that thoracic fellowship-trained radiologists from the STR and the increased number of readers contributed to increased sensitivity or higher clinical suspicion of opacities that led to the increase in prevalence. We tuned our threshold based on the stage 1 public LB results.</p>\n\n<h1>Detection</h1>\n\n<p>Please see the tl;dr for the repos we used for detection. We used a combination of 5 10-fold CV ensembles for detection. We computed the metric provided by Yicheng Chen (<a href=\"https://www.kaggle.com/chenyc15/mean-average-precision-metric\">https://www.kaggle.com/chenyc15/mean-average-precision-metric</a>) for both positive images only and all images to inform our model selection. Unfortunately we did a poor job of keeping track of how well our experiments did on stage 1 public LB, and the results are no longer visible. </p>\n\n<h2>Detection Ensemble 1</h2>\n\n<p>Each fold was trained on a different resolution (224-512 by increments of 32). <em>Positive images only.</em> We found that lower resolutions did not lower performance locally and actually increased performance (0.005-0.008) on public LB, so we stuck with it. It also allowed us to spam models in our ensemble due to lower training/inference overhead. The first detection ensemble was a 10-fold CV ensemble of deformable R-FCN. We mainly used default parameters, except we unfroze the non-BN layers that are frozen in the default configs. These models use a ResNet101 ImageNet pre-trained backbone. We also changed the max number of detections per image to 5 but kept a threshold of 0.001. Only data augmentation used was flip. These models trained very quickly (~2 hours) depending on the image resolution. </p>\n\n<h2>Detection Ensemble 2</h2>\n\n<p>Basically the same as #1, except we used deformable relation networks (<a href=\"https://arxiv.org/abs/1711.11575\">https://arxiv.org/abs/1711.11575</a>). Honestly, I don't really know how these work, and I stumbled upon them late into the competition. Since it was basically the same as the deformable R-FCN repo, it was easy to train these models and they performed very well. We tried the version where you attempt to learn the NMS thresholds to use, but that performed very poorly. </p>\n\n<h2>Detection Ensemble 3</h2>\n\n<p>Exactly the same as #2, except we used the default config (i.e. kept the backbone layers frozen). Keeping the backbone layers frozen had a slight decrease in performance, but we threw in this ensemble to decrease model correlation with the other ensembles, and it had a slight improvement in our stage 1 public LB score. </p>\n\n<h2>Detection Ensemble 4</h2>\n\n<p>This is a RetinaNet ensemble, also 10-fold CV, but trained only at 384 x 384 resolution because we had issues with changing anchor sizes. I believe a recent update allows you to specify a config.ini file where these changes can be made. This was probably our favorite ensemble. We trained on <em>concatenated</em> images where each negative image was randomly concatenated with a positive image from the same fold on the left or right (so final image sizes were 384 x 768). We trained for 8000 steps/epoch, batch size 1, for 8 epochs dividing learning rate by 10 after epoch 4 and 6. Validation was performed after each epoch. We selected the model with the best mAP metric over the 8 epochs. Using the <code>--random-transform</code> argument didn't really help/hurt but seemed to make training more unstable so we didn't use it (without specifying this, only data augmentation is flip). Half of the 10-fold CV ensemble was trained using ResNet101 backbone, the other half with ResNet152. ResNet101 was clearly better than ResNet50, but ResNet152 was the same as ResNet101. </p>\n\n<p>Training on concatenated images allowed RetinaNet to better balance precision and recall. You can achieve the same results by picking the right proportion of positives/negatives as well, but this seemed more \"clean\" to us. Inference was performed on single images. We looked at the AUC of these models using the max box score as class prediction and it was on par with our classification networks. The mAP on positive images only went down, but the mAP on all images went up, so this trade-off was beneficial to include in our ensemble. If we wanted to use a single model, it would be this one as it does not need a classifier to achieve good performance. Interestingly, training on concatenated images did not work well for detection ensembles 1-3. This may be because they are 2-stage detectors, but we didn't have time to look at this more closely. </p>\n\n<h2>Detection Ensemble 5</h2>\n\n<p>Same as #4, except trained on <em>positive images only.</em> </p>\n\n<h1>Ensemble</h1>\n\n<h2>Detection Ensemble 1+2+3</h2>\n\n<p>For detection ensembles 1-3, we applied 6x TTA (original, flipped, 80%/120% for both original and flipped). For each of the 30 models, we combined the TTA predictions using (<a href=\"https://github.com/ahrnbom/ensemble-objdet\">https://github.com/ahrnbom/ensemble-objdet</a>) with IoU threshold 0.4. Box score predictions were then adjusted by multiplying by the fraction of TTAs that contained that box (i.e. if 5/6 TTAs predicted a box with average score 0.5, it was multiplied by 5/6). This code expects <em>center coordinates</em> when computing the IoU overlap between boxes in the <code>getCoords</code> function. We didn't realize this until late in the competition (we were using top left), and when we changed this to take in top left coordinates, there was a drop in performance of about 0.01 in public LB. If anyone can help us figure out why, we haven't solved this yet. </p>\n\n<p>Now we have 30 models worth of predictions, so we combined those again with IoU threshold 0.4. No weighting was used. Box score predictions were then adjusted by multiplying by the fraction of models that contained that box (i.e. if 24/30 models predicted a box with average score 0.5, it was multiplied by 0.8 and the score was adjusted to 0.4). </p>\n\n<p>To incorporate the classification network, we multiplied the ensemble-averaged box score predictions by the classification score for that image. We eliminated boxes with an adjusted score of &lt;0.225. </p>\n\n<h2>Detection Ensemble 4+5</h2>\n\n<p>We applied 10x TTA (resolutions 320, 352, 384, 416, 448 for both original/flipped), performing the same kind of ensembling as above, including score adjustment based on fraction of TTAs predicting that box. Top 10 (or fewer) detections per TTA were selected. However we did not combine the predictions across 20 models in both ensembles at this stage. Instead, we used a classification score threshold of 0.2 and box score threshold of 0.3 for detection ensemble 4. The multiplication method we used above did not work well for RetinaNet. Though we said detection ensemble 4 didn't necessarily need a paired classifier, it did improve the stage 1 public LB score by ~0.003 so we used it since it was readily available. For detection ensemble 5, we used a classification score threshold of 0.325 and box score threshold of 0.35. </p>\n\n<p>After applying these thresholds, we combined boxes from ensembles 4 and 5 using the same code and IoU threshold 0.4. In this case, detection ensemble 4 was given 1.2 weight versus 0.8 for ensemble 5. This was because ensemble 4 performed slightly better (~0.004) than ensemble 5. Then we adjusted the score using the same strategy as above (in this case, if a box was only present in one model, score was divided in half). </p>\n\n<p>For tuning these thresholds, we aimed for a prevalence of about 35-37% and experimented with what worked best on stage 1 public LB. We didn't have any real local validation because we realized early on that LB score would be the best indicator of model performance on the final test data. </p>\n\n<h1>Final Ensemble</h1>\n\n<p>We are now left with 2 ensembles. We combined them using equal weight and IoU threshold as described above, using the same adjustment strategy. Final box threshold used was 0.15.</p>\n\n<h1>Post-processing</h1>\n\n<p>We realized that the triple-read boxes were smaller than the single-read boxes. This makes sense because in the annotation process described here (<a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723\">https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723</a>) the <em>intersection</em> of boxes was used as opposed to the average. We tried to mimic this intersection process by taking the average of multiple intersections of boxes across models. This gave ~10-15% (!) improvement on stage 1 public LB. As an alternative to this, we simply resized the boxes by multiplying length/width by a fraction. We found 87.5% for each was a good reduction and worked a bit better than doing the intersection. It was also much easier to implement. We discovered this early on, and didn't submit anything without resizing after that. Early on in the competition we had an improvement from 0.181-&gt;0.209 with resizing. Towards the end, I wanted to see the effect of resizing again and it was an improvement of 0.218-&gt;0.252, which is huge. For our final stage 1 submission, resize improved our score from 0.222-&gt;0.260. It seemed like other people were seeing better success with higher resolution models, so we wonder if the resizing was more complementary with lower resolution models. Maybe if we didn't resize, higher resolution models would perform better.</p>\n\n<h1>Statistics for final submission</h1>\n\n<pre><code>Stage 1: \nNUM POSITIVES: 351 \n% POSITIVES: 0.351 \nNUM BOXES: 582 \nAVG # BOXES PER CASE: 1.65811965812 \nBOX SIZE [MEDIAN]: 60084.5 \nBOX SIZE [MEAN]: 66052.2405498 \nBOX SIZE [MIN]: 16065 \nBOX SIZE [MAX]: 151580\n\nStage 2: \nNUM POSITIVES: 1106 \n% POSITIVES: 0.368666666667 \nNUM BOXES: 1830 \nAVG # BOXES PER CASE: 1.65461121157 \nBOX SIZE [MEDIAN]: 60877.5\nBOX SIZE [MEAN]: 66454.8060109 \nBOX SIZE [MIN]: 13431 \nBOX SIZE [MAX]: 192648\n</code></pre>\n\n<h1>Miscellaneous thoughts</h1>\n\n<p>Basically we had a good model that we turned into a \"great\" model by resizing the output. It will be interesting to see if other top teams achieved their score by improving more upon classification versus box precision and how their scores would be affected by resize if they did not apply any post-processing to their predictions. We went overboard with ensembles and probably could have achieved the same performance with &lt;20 models, but it became so easy to train them that we just included a bunch. We could have experimented with more hyperparameters in our detection models as well. I really dislike hyperparameter tuning (never really developed a good strategy for it) and often try and compensate by ensembling different models together. </p>\n\n<h1>Things that didn't work</h1>\n\n<ul>\n<li>We tried training another classifier on out-of-fold bounding box predictions produced by our detection models to classify into IoU &gt;0.4 and &lt;0.4</li>\n<li>NMS of overlapping bounding boxes in our final submission: a number of images in our final submission for both stage 1 and stage 2 LBs had overlapping boxes (usually a smaller one contained in a larger one). Suppressing these actually reduced our score. </li>\n<li>We tried various experiments that treated AP/PA images differently (e.g. different thresholds, different resizes), but in the end it was easier and better to be view-agnostic</li>\n<li>Getting all bounding box predictions from all TTAs and models and combining them at once. This didn't work as well as the stepwise approach we described above. This may have something to do with non-standardized prediction scores (though we tried standardizing and it didn't help much). </li>\n<li>SoftNMS -- this makes sense because you would not expect overlapping objects in this challenge as opposed to others like COCO. </li>\n<li>For RetinaNet: other backbones. Only the ResNet backbones worked well for us. </li>\n<li>For RetinaNet: pre-training detector heads. It was a lot easier to use ImageNet pre-trained backbones and then just start training the whole network from the beginning.</li>\n<li>For RetinaNet: pre-training the backbone on the pneumonia dataset. No improvement. </li>\n</ul>",
  "messages": [
    {
      "id": 414711,
      "postDate": "2018-11-03T12:44:48.530Z",
      "content": "<p>We are very excited to have been awarded 1st place in this competition. This was our first foray into the world of object detection. I'm Ian, currently a 3rd-year medical student at Brown University, Providence, RI, USA. My teammate Alexandre is a radiologist practicing in Lanaudière, Québec, Canada. We are happy to represent the medical community in this competition. Huge congratulations to Dmytro for becoming Grandmaster. His top score during the public LB really motivated us to make successive improvements to our model. Truly an honor to be in the same ranks as him for this challenge. Congratulations to all of the other top 10 winners as well. </p>\n\n<h1>tl;dr</h1>\n\n<ul>\n<li>We did NOT retrain when stage 1 test labels were released</li>\n<li>We used a classification-detection pipeline </li>\n<li>10-fold CV ensemble for classification, combination of 5 10-fold CV ensembles for detection (50 models) </li>\n<li>For detection, we used:\n<ol><li>RetinaNet: <a href=\"https://github.com/fizyr/keras-retinanet\"></a><a href=\"https://github.com/fizyr/keras-retinanet\">https://github.com/fizyr/keras-retinanet</a></li>\n<li>Deformable R-FCN: <a href=\"https://github.com/msracver/Deformable-ConvNets\"></a><a href=\"https://github.com/msracver/Deformable-ConvNets\">https://github.com/msracver/Deformable-ConvNets</a></li>\n<li>Deformable Relation Networks: <a href=\"https://github.com/msracver/Relation-Networks-for-Object-Detection\"></a><a href=\"https://github.com/msracver/Relation-Networks-for-Object-Detection\">https://github.com/msracver/Relation-Networks-for-Object-Detection</a></li></ol></li>\n<li>Boxes were ensembled using: <a href=\"https://github.com/ahrnbom/ensemble-objdet\">https://github.com/ahrnbom/ensemble-objdet</a></li>\n<li><em>We resized the lengths and widths of our final predictions by 87.5%</em></li>\n<li>Code available here: <a href=\"https://www.github.com/i-pan/kaggle-rsna18\">https://www.github.com/i-pan/kaggle-rsna18</a></li>\n<li>We found that a 6-model ensemble with 1 InceptionResNetV2 for classification and 5 Deformable Relation Networks achieved 0.253 on stage 2 private LB</li>\n</ul>\n\n<h1>Classification</h1>\n\n<p>We used Keras 2.2 for classification. Our classification model was composed of the following individual models [format: modelArchitecture (numClasses) (imgSize)]. Models were trained on either 2 classes (opacity vs. not) or 3 classes (opacity vs. not normal/no opacity vs. normal). Each model was trained on a different fold.\n- InceptionResNetV2 (2) (256), InceptionResNetV2 (2) (320) \n- InceptionResNetV2 (3) (256), InceptionResNetV2 (3) (320) \n- Xception (2) (384), Xception (2) (448) \n- Xception (3) (384), Xception (3) (448) \n- DenseNet169 (2) (512), DenseNet169 (3) (512)\nFirst, we trained ImageNet pre-trained networks on the NIH ChestX-ray14 dataset using 15 classes (14 findings + abnormal vs. normal). Then we fine-tuned those weights on the pneumonia dataset. This improved results from training using ImageNet weights only by about 1% locally. We were getting about 0.88-0.90 AUC across our folds. </p>\n\n<p>Models were trained with 50% probability of being color-inverted, 50% of being flipped, and 50% of being augmented with some other augmentation (e.g. contrast enhancement, crop, rotation). We used 15x TTA per model for our final predictions and averaged the 150 predictions for our final classification score. When the stage 1 test labels were released, our classification ensemble had an AUC of 0.93. </p>\n\n<p>As many other competitors noticed, the distribution of the training (single-read) data and test (triple-read) data were quite different. We think that thoracic fellowship-trained radiologists from the STR and the increased number of readers contributed to increased sensitivity or higher clinical suspicion of opacities that led to the increase in prevalence. We tuned our threshold based on the stage 1 public LB results.</p>\n\n<h1>Detection</h1>\n\n<p>Please see the tl;dr for the repos we used for detection. We used a combination of 5 10-fold CV ensembles for detection. We computed the metric provided by Yicheng Chen (<a href=\"https://www.kaggle.com/chenyc15/mean-average-precision-metric\">https://www.kaggle.com/chenyc15/mean-average-precision-metric</a>) for both positive images only and all images to inform our model selection. Unfortunately we did a poor job of keeping track of how well our experiments did on stage 1 public LB, and the results are no longer visible. </p>\n\n<h2>Detection Ensemble 1</h2>\n\n<p>Each fold was trained on a different resolution (224-512 by increments of 32). <em>Positive images only.</em> We found that lower resolutions did not lower performance locally and actually increased performance (0.005-0.008) on public LB, so we stuck with it. It also allowed us to spam models in our ensemble due to lower training/inference overhead. The first detection ensemble was a 10-fold CV ensemble of deformable R-FCN. We mainly used default parameters, except we unfroze the non-BN layers that are frozen in the default configs. These models use a ResNet101 ImageNet pre-trained backbone. We also changed the max number of detections per image to 5 but kept a threshold of 0.001. Only data augmentation used was flip. These models trained very quickly (~2 hours) depending on the image resolution. </p>\n\n<h2>Detection Ensemble 2</h2>\n\n<p>Basically the same as #1, except we used deformable relation networks (<a href=\"https://arxiv.org/abs/1711.11575\">https://arxiv.org/abs/1711.11575</a>). Honestly, I don't really know how these work, and I stumbled upon them late into the competition. Since it was basically the same as the deformable R-FCN repo, it was easy to train these models and they performed very well. We tried the version where you attempt to learn the NMS thresholds to use, but that performed very poorly. </p>\n\n<h2>Detection Ensemble 3</h2>\n\n<p>Exactly the same as #2, except we used the default config (i.e. kept the backbone layers frozen). Keeping the backbone layers frozen had a slight decrease in performance, but we threw in this ensemble to decrease model correlation with the other ensembles, and it had a slight improvement in our stage 1 public LB score. </p>\n\n<h2>Detection Ensemble 4</h2>\n\n<p>This is a RetinaNet ensemble, also 10-fold CV, but trained only at 384 x 384 resolution because we had issues with changing anchor sizes. I believe a recent update allows you to specify a config.ini file where these changes can be made. This was probably our favorite ensemble. We trained on <em>concatenated</em> images where each negative image was randomly concatenated with a positive image from the same fold on the left or right (so final image sizes were 384 x 768). We trained for 8000 steps/epoch, batch size 1, for 8 epochs dividing learning rate by 10 after epoch 4 and 6. Validation was performed after each epoch. We selected the model with the best mAP metric over the 8 epochs. Using the <code>--random-transform</code> argument didn't really help/hurt but seemed to make training more unstable so we didn't use it (without specifying this, only data augmentation is flip). Half of the 10-fold CV ensemble was trained using ResNet101 backbone, the other half with ResNet152. ResNet101 was clearly better than ResNet50, but ResNet152 was the same as ResNet101. </p>\n\n<p>Training on concatenated images allowed RetinaNet to better balance precision and recall. You can achieve the same results by picking the right proportion of positives/negatives as well, but this seemed more \"clean\" to us. Inference was performed on single images. We looked at the AUC of these models using the max box score as class prediction and it was on par with our classification networks. The mAP on positive images only went down, but the mAP on all images went up, so this trade-off was beneficial to include in our ensemble. If we wanted to use a single model, it would be this one as it does not need a classifier to achieve good performance. Interestingly, training on concatenated images did not work well for detection ensembles 1-3. This may be because they are 2-stage detectors, but we didn't have time to look at this more closely. </p>\n\n<h2>Detection Ensemble 5</h2>\n\n<p>Same as #4, except trained on <em>positive images only.</em> </p>\n\n<h1>Ensemble</h1>\n\n<h2>Detection Ensemble 1+2+3</h2>\n\n<p>For detection ensembles 1-3, we applied 6x TTA (original, flipped, 80%/120% for both original and flipped). For each of the 30 models, we combined the TTA predictions using (<a href=\"https://github.com/ahrnbom/ensemble-objdet\">https://github.com/ahrnbom/ensemble-objdet</a>) with IoU threshold 0.4. Box score predictions were then adjusted by multiplying by the fraction of TTAs that contained that box (i.e. if 5/6 TTAs predicted a box with average score 0.5, it was multiplied by 5/6). This code expects <em>center coordinates</em> when computing the IoU overlap between boxes in the <code>getCoords</code> function. We didn't realize this until late in the competition (we were using top left), and when we changed this to take in top left coordinates, there was a drop in performance of about 0.01 in public LB. If anyone can help us figure out why, we haven't solved this yet. </p>\n\n<p>Now we have 30 models worth of predictions, so we combined those again with IoU threshold 0.4. No weighting was used. Box score predictions were then adjusted by multiplying by the fraction of models that contained that box (i.e. if 24/30 models predicted a box with average score 0.5, it was multiplied by 0.8 and the score was adjusted to 0.4). </p>\n\n<p>To incorporate the classification network, we multiplied the ensemble-averaged box score predictions by the classification score for that image. We eliminated boxes with an adjusted score of &lt;0.225. </p>\n\n<h2>Detection Ensemble 4+5</h2>\n\n<p>We applied 10x TTA (resolutions 320, 352, 384, 416, 448 for both original/flipped), performing the same kind of ensembling as above, including score adjustment based on fraction of TTAs predicting that box. Top 10 (or fewer) detections per TTA were selected. However we did not combine the predictions across 20 models in both ensembles at this stage. Instead, we used a classification score threshold of 0.2 and box score threshold of 0.3 for detection ensemble 4. The multiplication method we used above did not work well for RetinaNet. Though we said detection ensemble 4 didn't necessarily need a paired classifier, it did improve the stage 1 public LB score by ~0.003 so we used it since it was readily available. For detection ensemble 5, we used a classification score threshold of 0.325 and box score threshold of 0.35. </p>\n\n<p>After applying these thresholds, we combined boxes from ensembles 4 and 5 using the same code and IoU threshold 0.4. In this case, detection ensemble 4 was given 1.2 weight versus 0.8 for ensemble 5. This was because ensemble 4 performed slightly better (~0.004) than ensemble 5. Then we adjusted the score using the same strategy as above (in this case, if a box was only present in one model, score was divided in half). </p>\n\n<p>For tuning these thresholds, we aimed for a prevalence of about 35-37% and experimented with what worked best on stage 1 public LB. We didn't have any real local validation because we realized early on that LB score would be the best indicator of model performance on the final test data. </p>\n\n<h1>Final Ensemble</h1>\n\n<p>We are now left with 2 ensembles. We combined them using equal weight and IoU threshold as described above, using the same adjustment strategy. Final box threshold used was 0.15.</p>\n\n<h1>Post-processing</h1>\n\n<p>We realized that the triple-read boxes were smaller than the single-read boxes. This makes sense because in the annotation process described here (<a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723\">https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723</a>) the <em>intersection</em> of boxes was used as opposed to the average. We tried to mimic this intersection process by taking the average of multiple intersections of boxes across models. This gave ~10-15% (!) improvement on stage 1 public LB. As an alternative to this, we simply resized the boxes by multiplying length/width by a fraction. We found 87.5% for each was a good reduction and worked a bit better than doing the intersection. It was also much easier to implement. We discovered this early on, and didn't submit anything without resizing after that. Early on in the competition we had an improvement from 0.181-&gt;0.209 with resizing. Towards the end, I wanted to see the effect of resizing again and it was an improvement of 0.218-&gt;0.252, which is huge. For our final stage 1 submission, resize improved our score from 0.222-&gt;0.260. It seemed like other people were seeing better success with higher resolution models, so we wonder if the resizing was more complementary with lower resolution models. Maybe if we didn't resize, higher resolution models would perform better.</p>\n\n<h1>Statistics for final submission</h1>\n\n<pre><code>Stage 1: \nNUM POSITIVES: 351 \n% POSITIVES: 0.351 \nNUM BOXES: 582 \nAVG # BOXES PER CASE: 1.65811965812 \nBOX SIZE [MEDIAN]: 60084.5 \nBOX SIZE [MEAN]: 66052.2405498 \nBOX SIZE [MIN]: 16065 \nBOX SIZE [MAX]: 151580\n\nStage 2: \nNUM POSITIVES: 1106 \n% POSITIVES: 0.368666666667 \nNUM BOXES: 1830 \nAVG # BOXES PER CASE: 1.65461121157 \nBOX SIZE [MEDIAN]: 60877.5\nBOX SIZE [MEAN]: 66454.8060109 \nBOX SIZE [MIN]: 13431 \nBOX SIZE [MAX]: 192648\n</code></pre>\n\n<h1>Miscellaneous thoughts</h1>\n\n<p>Basically we had a good model that we turned into a \"great\" model by resizing the output. It will be interesting to see if other top teams achieved their score by improving more upon classification versus box precision and how their scores would be affected by resize if they did not apply any post-processing to their predictions. We went overboard with ensembles and probably could have achieved the same performance with &lt;20 models, but it became so easy to train them that we just included a bunch. We could have experimented with more hyperparameters in our detection models as well. I really dislike hyperparameter tuning (never really developed a good strategy for it) and often try and compensate by ensembling different models together. </p>\n\n<h1>Things that didn't work</h1>\n\n<ul>\n<li>We tried training another classifier on out-of-fold bounding box predictions produced by our detection models to classify into IoU &gt;0.4 and &lt;0.4</li>\n<li>NMS of overlapping bounding boxes in our final submission: a number of images in our final submission for both stage 1 and stage 2 LBs had overlapping boxes (usually a smaller one contained in a larger one). Suppressing these actually reduced our score. </li>\n<li>We tried various experiments that treated AP/PA images differently (e.g. different thresholds, different resizes), but in the end it was easier and better to be view-agnostic</li>\n<li>Getting all bounding box predictions from all TTAs and models and combining them at once. This didn't work as well as the stepwise approach we described above. This may have something to do with non-standardized prediction scores (though we tried standardizing and it didn't help much). </li>\n<li>SoftNMS -- this makes sense because you would not expect overlapping objects in this challenge as opposed to others like COCO. </li>\n<li>For RetinaNet: other backbones. Only the ResNet backbones worked well for us. </li>\n<li>For RetinaNet: pre-training detector heads. It was a lot easier to use ImageNet pre-trained backbones and then just start training the whole network from the beginning.</li>\n<li>For RetinaNet: pre-training the backbone on the pneumonia dataset. No improvement. </li>\n</ul>",
      "rawMarkdown": "We are very excited to have been awarded 1st place in this competition. This was our first foray into the world of object detection. I'm Ian, currently a 3rd-year medical student at Brown University, Providence, RI, USA. My teammate Alexandre is a radiologist practicing in Lanaudière, Québec, Canada. We are happy to represent the medical community in this competition. Huge congratulations to Dmytro for becoming Grandmaster. His top score during the public LB really motivated us to make successive improvements to our model. Truly an honor to be in the same ranks as him for this challenge. Congratulations to all of the other top 10 winners as well. \n\n# tl;dr \n- We did NOT retrain when stage 1 test labels were released\n- We used a classification-detection pipeline \n- 10-fold CV ensemble for classification, combination of 5 10-fold CV ensembles for detection (50 models) \n- For detection, we used:\n1. RetinaNet: https://github.com/fizyr/keras-retinanet\n2. Deformable R-FCN: https://github.com/msracver/Deformable-ConvNets\n3. Deformable Relation Networks: https://github.com/msracver/Relation-Networks-for-Object-Detection\n- Boxes were ensembled using: https://github.com/ahrnbom/ensemble-objdet\n- *We resized the lengths and widths of our final predictions by 87.5%*\n- Code available here: https://www.github.com/i-pan/kaggle-rsna18\n- We found that a 6-model ensemble with 1 InceptionResNetV2 for classification and 5 Deformable Relation Networks achieved 0.253 on stage 2 private LB\n\n# Classification \nWe used Keras 2.2 for classification. Our classification model was composed of the following individual models [format: modelArchitecture (numClasses) (imgSize)]. Models were trained on either 2 classes (opacity vs. not) or 3 classes (opacity vs. not normal/no opacity vs. normal). Each model was trained on a different fold.\n- InceptionResNetV2 (2) (256), InceptionResNetV2 (2) (320) \n- InceptionResNetV2 (3) (256), InceptionResNetV2 (3) (320) \n- Xception (2) (384), Xception (2) (448) \n- Xception (3) (384), Xception (3) (448) \n- DenseNet169 (2) (512), DenseNet169 (3) (512)\nFirst, we trained ImageNet pre-trained networks on the NIH ChestX-ray14 dataset using 15 classes (14 findings + abnormal vs. normal). Then we fine-tuned those weights on the pneumonia dataset. This improved results from training using ImageNet weights only by about 1% locally. We were getting about 0.88-0.90 AUC across our folds. \n\nModels were trained with 50% probability of being color-inverted, 50% of being flipped, and 50% of being augmented with some other augmentation (e.g. contrast enhancement, crop, rotation). We used 15x TTA per model for our final predictions and averaged the 150 predictions for our final classification score. When the stage 1 test labels were released, our classification ensemble had an AUC of 0.93. \n\nAs many other competitors noticed, the distribution of the training (single-read) data and test (triple-read) data were quite different. We think that thoracic fellowship-trained radiologists from the STR and the increased number of readers contributed to increased sensitivity or higher clinical suspicion of opacities that led to the increase in prevalence. We tuned our threshold based on the stage 1 public LB results.\n\n# Detection\nPlease see the tl;dr for the repos we used for detection. We used a combination of 5 10-fold CV ensembles for detection. We computed the metric provided by Yicheng Chen (https://www.kaggle.com/chenyc15/mean-average-precision-metric) for both positive images only and all images to inform our model selection. Unfortunately we did a poor job of keeping track of how well our experiments did on stage 1 public LB, and the results are no longer visible. \n\n## Detection Ensemble 1\nEach fold was trained on a different resolution (224-512 by increments of 32). *Positive images only.* We found that lower resolutions did not lower performance locally and actually increased performance (0.005-0.008) on public LB, so we stuck with it. It also allowed us to spam models in our ensemble due to lower training/inference overhead. The first detection ensemble was a 10-fold CV ensemble of deformable R-FCN. We mainly used default parameters, except we unfroze the non-BN layers that are frozen in the default configs. These models use a ResNet101 ImageNet pre-trained backbone. We also changed the max number of detections per image to 5 but kept a threshold of 0.001. Only data augmentation used was flip. These models trained very quickly (~2 hours) depending on the image resolution. \n\n## Detection Ensemble 2\nBasically the same as #1, except we used deformable relation networks (https://arxiv.org/abs/1711.11575). Honestly, I don't really know how these work, and I stumbled upon them late into the competition. Since it was basically the same as the deformable R-FCN repo, it was easy to train these models and they performed very well. We tried the version where you attempt to learn the NMS thresholds to use, but that performed very poorly. \n\n## Detection Ensemble 3\nExactly the same as #2, except we used the default config (i.e. kept the backbone layers frozen). Keeping the backbone layers frozen had a slight decrease in performance, but we threw in this ensemble to decrease model correlation with the other ensembles, and it had a slight improvement in our stage 1 public LB score. \n\n## Detection Ensemble 4\nThis is a RetinaNet ensemble, also 10-fold CV, but trained only at 384 x 384 resolution because we had issues with changing anchor sizes. I believe a recent update allows you to specify a config.ini file where these changes can be made. This was probably our favorite ensemble. We trained on *concatenated* images where each negative image was randomly concatenated with a positive image from the same fold on the left or right (so final image sizes were 384 x 768). We trained for 8000 steps/epoch, batch size 1, for 8 epochs dividing learning rate by 10 after epoch 4 and 6. Validation was performed after each epoch. We selected the model with the best mAP metric over the 8 epochs. Using the `--random-transform` argument didn't really help/hurt but seemed to make training more unstable so we didn't use it (without specifying this, only data augmentation is flip). Half of the 10-fold CV ensemble was trained using ResNet101 backbone, the other half with ResNet152. ResNet101 was clearly better than ResNet50, but ResNet152 was the same as ResNet101. \n\nTraining on concatenated images allowed RetinaNet to better balance precision and recall. You can achieve the same results by picking the right proportion of positives/negatives as well, but this seemed more \"clean\" to us. Inference was performed on single images. We looked at the AUC of these models using the max box score as class prediction and it was on par with our classification networks. The mAP on positive images only went down, but the mAP on all images went up, so this trade-off was beneficial to include in our ensemble. If we wanted to use a single model, it would be this one as it does not need a classifier to achieve good performance. Interestingly, training on concatenated images did not work well for detection ensembles 1-3. This may be because they are 2-stage detectors, but we didn't have time to look at this more closely. \n\n## Detection Ensemble 5\nSame as #4, except trained on *positive images only.* \n\n# Ensemble\n\n## Detection Ensemble 1+2+3\nFor detection ensembles 1-3, we applied 6x TTA (original, flipped, 80%/120% for both original and flipped). For each of the 30 models, we combined the TTA predictions using (https://github.com/ahrnbom/ensemble-objdet) with IoU threshold 0.4. Box score predictions were then adjusted by multiplying by the fraction of TTAs that contained that box (i.e. if 5/6 TTAs predicted a box with average score 0.5, it was multiplied by 5/6). This code expects *center coordinates* when computing the IoU overlap between boxes in the `getCoords` function. We didn't realize this until late in the competition (we were using top left), and when we changed this to take in top left coordinates, there was a drop in performance of about 0.01 in public LB. If anyone can help us figure out why, we haven't solved this yet. \n\nNow we have 30 models worth of predictions, so we combined those again with IoU threshold 0.4. No weighting was used. Box score predictions were then adjusted by multiplying by the fraction of models that contained that box (i.e. if 24/30 models predicted a box with average score 0.5, it was multiplied by 0.8 and the score was adjusted to 0.4). \n\nTo incorporate the classification network, we multiplied the ensemble-averaged box score predictions by the classification score for that image. We eliminated boxes with an adjusted score of &lt;0.225. \n\n## Detection Ensemble 4+5\nWe applied 10x TTA (resolutions 320, 352, 384, 416, 448 for both original/flipped), performing the same kind of ensembling as above, including score adjustment based on fraction of TTAs predicting that box. Top 10 (or fewer) detections per TTA were selected. However we did not combine the predictions across 20 models in both ensembles at this stage. Instead, we used a classification score threshold of 0.2 and box score threshold of 0.3 for detection ensemble 4. The multiplication method we used above did not work well for RetinaNet. Though we said detection ensemble 4 didn't necessarily need a paired classifier, it did improve the stage 1 public LB score by ~0.003 so we used it since it was readily available. For detection ensemble 5, we used a classification score threshold of 0.325 and box score threshold of 0.35. \n\nAfter applying these thresholds, we combined boxes from ensembles 4 and 5 using the same code and IoU threshold 0.4. In this case, detection ensemble 4 was given 1.2 weight versus 0.8 for ensemble 5. This was because ensemble 4 performed slightly better (~0.004) than ensemble 5. Then we adjusted the score using the same strategy as above (in this case, if a box was only present in one model, score was divided in half). \n\nFor tuning these thresholds, we aimed for a prevalence of about 35-37% and experimented with what worked best on stage 1 public LB. We didn't have any real local validation because we realized early on that LB score would be the best indicator of model performance on the final test data. \n\n# Final Ensemble\nWe are now left with 2 ensembles. We combined them using equal weight and IoU threshold as described above, using the same adjustment strategy. Final box threshold used was 0.15.\n\n# Post-processing \nWe realized that the triple-read boxes were smaller than the single-read boxes. This makes sense because in the annotation process described here (https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723) the *intersection* of boxes was used as opposed to the average. We tried to mimic this intersection process by taking the average of multiple intersections of boxes across models. This gave ~10-15% (!) improvement on stage 1 public LB. As an alternative to this, we simply resized the boxes by multiplying length/width by a fraction. We found 87.5% for each was a good reduction and worked a bit better than doing the intersection. It was also much easier to implement. We discovered this early on, and didn't submit anything without resizing after that. Early on in the competition we had an improvement from 0.181-&gt;0.209 with resizing. Towards the end, I wanted to see the effect of resizing again and it was an improvement of 0.218-&gt;0.252, which is huge. For our final stage 1 submission, resize improved our score from 0.222-&gt;0.260. It seemed like other people were seeing better success with higher resolution models, so we wonder if the resizing was more complementary with lower resolution models. Maybe if we didn't resize, higher resolution models would perform better.\n\n# Statistics for final submission\n\n    Stage 1: \n    NUM POSITIVES: 351 \n    % POSITIVES: 0.351 \n    NUM BOXES: 582 \n    AVG # BOXES PER CASE: 1.65811965812 \n    BOX SIZE [MEDIAN]: 60084.5 \n    BOX SIZE [MEAN]: 66052.2405498 \n    BOX SIZE [MIN]: 16065 \n    BOX SIZE [MAX]: 151580\n    \n    Stage 2: \n    NUM POSITIVES: 1106 \n    % POSITIVES: 0.368666666667 \n    NUM BOXES: 1830 \n    AVG # BOXES PER CASE: 1.65461121157 \n    BOX SIZE [MEDIAN]: 60877.5\n    BOX SIZE [MEAN]: 66454.8060109 \n    BOX SIZE [MIN]: 13431 \n    BOX SIZE [MAX]: 192648\n\n# Miscellaneous thoughts\nBasically we had a good model that we turned into a \"great\" model by resizing the output. It will be interesting to see if other top teams achieved their score by improving more upon classification versus box precision and how their scores would be affected by resize if they did not apply any post-processing to their predictions. We went overboard with ensembles and probably could have achieved the same performance with &lt;20 models, but it became so easy to train them that we just included a bunch. We could have experimented with more hyperparameters in our detection models as well. I really dislike hyperparameter tuning (never really developed a good strategy for it) and often try and compensate by ensembling different models together. \n\n# Things that didn't work \n- We tried training another classifier on out-of-fold bounding box predictions produced by our detection models to classify into IoU &gt;0.4 and &lt;0.4\n- NMS of overlapping bounding boxes in our final submission: a number of images in our final submission for both stage 1 and stage 2 LBs had overlapping boxes (usually a smaller one contained in a larger one). Suppressing these actually reduced our score. \n- We tried various experiments that treated AP/PA images differently (e.g. different thresholds, different resizes), but in the end it was easier and better to be view-agnostic\n- Getting all bounding box predictions from all TTAs and models and combining them at once. This didn't work as well as the stepwise approach we described above. This may have something to do with non-standardized prediction scores (though we tried standardizing and it didn't help much). \n- SoftNMS -- this makes sense because you would not expect overlapping objects in this challenge as opposed to others like COCO. \n- For RetinaNet: other backbones. Only the ResNet backbones worked well for us. \n- For RetinaNet: pre-training detector heads. It was a lot easier to use ImageNet pre-trained backbones and then just start training the whole network from the beginning.\n- For RetinaNet: pre-training the backbone on the pneumonia dataset. No improvement. \n",
      "votes": 161
    },
    {
      "id": 417614,
      "postDate": "2018-11-08T14:23:30.810Z",
      "content": "<p>We have uploaded our solution code to GitHub. Please find it here: <a href=\"https://www.github.com/i-pan/kaggle-rsna18\">https://www.github.com/i-pan/kaggle-rsna18</a></p>\n\n<p>If you have any issues, please let us know! </p>",
      "rawMarkdown": "We have uploaded our solution code to GitHub. Please find it here: https://www.github.com/i-pan/kaggle-rsna18\n\nIf you have any issues, please let us know! ",
      "votes": 10
    },
    {
      "id": 414754,
      "postDate": "2018-11-03T14:35:58.087Z",
      "content": "<p>Congratulations with the 1st place!</p>\n\n<p>I scaled the size of \"harder\" boxes as well, but slightly less and proportionally to variance (actually difference between 20 and 80 percentile) between folds predictions. The intuition behind - to simulate the labeling process of test samples when intersection used from different radiologists labels.</p>\n\n<p>my calculations for adjusted size, h_perc20 is 20 percentile of anchor h between folds and checkpoints:\nh = h_perc20 - 1.6 * (h_perc80 - h_perc20)\nw = w_perc20 - 1.6 * (w_perc80 - w_perc20)</p>\n\n<p>box size mean 67632.9\nbox size median 65335.1</p>",
      "rawMarkdown": "Congratulations with the 1st place!\n\nI scaled the size of \"harder\" boxes as well, but slightly less and proportionally to variance (actually difference between 20 and 80 percentile) between folds predictions. The intuition behind - to simulate the labeling process of test samples when intersection used from different radiologists labels.\n\nmy calculations for adjusted size, h_perc20 is 20 percentile of anchor h between folds and checkpoints:\nh = h_perc20 - 1.6 * (h_perc80 - h_perc20)\nw = w_perc20 - 1.6 * (w_perc80 - w_perc20)\n\nbox size mean 67632.9\nbox size median 65335.1",
      "votes": 1,
      "replies": [
        {
          "id": 414904,
          "postDate": "2018-11-03T21:32:00.563Z",
          "content": "<p>Interesting - did you ever try doing a fixed resizing of the boxes as a comparison? We also wanted to account for variance which we tried to simulate using our intersection method, but it actually performed worse. Your method might be better than that though. </p>",
          "rawMarkdown": "Interesting - did you ever try doing a fixed resizing of the boxes as a comparison? We also wanted to account for variance which we tried to simulate using our intersection method, but it actually performed worse. Your method might be better than that though. ",
          "votes": 1
        },
        {
          "id": 415002,
          "postDate": "2018-11-04T05:08:56.900Z",
          "content": "<p>The intuition when we tried to mimick the intersection from N different annotators was to catch a different variance for each border. Practically, the annotators variance for each border can potentially be asymetric if a lung opacity is visually well defined on a specific border but undefined on other borders. But as Ian specified, simple linear resize just worked better because variance was probably, on average, relatively symetric for all borders on the stage 1 and stage 2 test datasets</p>",
          "rawMarkdown": "The intuition when we tried to mimick the intersection from N different annotators was to catch a different variance for each border. Practically, the annotators variance for each border can potentially be asymetric if a lung opacity is visually well defined on a specific border but undefined on other borders. But as Ian specified, simple linear resize just worked better because variance was probably, on average, relatively symetric for all borders on the stage 1 and stage 2 test datasets"
        },
        {
          "id": 415035,
          "postDate": "2018-11-04T08:24:37.513Z",
          "content": "<p>I planned to test the similar scale reduction as you have done (I even started to implement it first), but run out of time. I have just tried to do a submission with boxes size reduced by 87.5% and received a close but better score (0.247 -&gt; 0.249 on stage 2)</p>",
          "rawMarkdown": "I planned to test the similar scale reduction as you have done (I even started to implement it first), but run out of time. I have just tried to do a submission with boxes size reduced by 87.5% and received a close but better score (0.247 -&gt; 0.249 on stage 2)"
        }
      ]
    },
    {
      "id": 416643,
      "postDate": "2018-11-07T02:59:54.667Z",
      "content": "<p>Congratulations Ian and Alex on your first place achievement!  Your solution is an optimization tour de force, I'm thoroughly impressed and look forward to exploring your solution in depth.  I came up with only a small subset of your insights, though it looks like we did similar bounding box resizing and box coordinate averaging.  I found that a slightly larger bounding box size reduction improved my scores, perhaps because I included rotations in my image augmentation.</p>",
      "rawMarkdown": "Congratulations Ian and Alex on your first place achievement!  Your solution is an optimization tour de force, I'm thoroughly impressed and look forward to exploring your solution in depth.  I came up with only a small subset of your insights, though it looks like we did similar bounding box resizing and box coordinate averaging.  I found that a slightly larger bounding box size reduction improved my scores, perhaps because I included rotations in my image augmentation.",
      "votes": 2
    },
    {
      "id": 3304997,
      "postDate": "2025-10-21T19:42:56.813Z",
      "content": "<p>Many congratulations 🎉🎉</p>",
      "rawMarkdown": "Many congratulations 🎉🎉"
    },
    {
      "id": 440150,
      "postDate": "2018-12-17T06:15:19.417Z",
      "content": "<p>Hi Ian!</p>\n\n<p>Thanks for sharing. I'm getting the following error when trying to replicate your work. I'm unable to transform data into the COCO format:</p>\n\n<p>Transforming data into COCO format ...</p>\n\n<pre><code>usage: 6_COCOify.py [-h]\n                    subset [TRAIN_LABELS_PATH] [TRAIN_IMAGES_DIR]\n                    [TEST_IMAGES_DIR] [FOLDS_DF_PATH]\n6_COCOify.py: error: too few arguments\n</code></pre>\n\n<p>I've tried resolving it but nothings working. I'll appreciate if you can help me out.</p>",
      "rawMarkdown": "Hi Ian!\n\nThanks for sharing. I'm getting the following error when trying to replicate your work. I'm unable to transform data into the COCO format:\n\nTransforming data into COCO format ...\n\n    usage: 6_COCOify.py [-h]\n                        subset [TRAIN_LABELS_PATH] [TRAIN_IMAGES_DIR]\n                        [TEST_IMAGES_DIR] [FOLDS_DF_PATH]\n    6_COCOify.py: error: too few arguments\n\nI've tried resolving it but nothings working. I'll appreciate if you can help me out."
    },
    {
      "id": 436156,
      "postDate": "2018-12-09T17:45:43.430Z",
      "content": "<p>Congratulations :) did you alter the image resolution for classification?</p>",
      "rawMarkdown": "Congratulations :) did you alter the image resolution for classification?"
    },
    {
      "id": 428140,
      "postDate": "2018-11-26T20:32:12.127Z",
      "content": "<p>Anyone know the venue and time when these solutions are presented at RSNA? I've purchased the virtual meeting but couldn't find any session pertaining to this challenge on the program.</p>",
      "rawMarkdown": "Anyone know the venue and time when these solutions are presented at RSNA? I've purchased the virtual meeting but couldn't find any session pertaining to this challenge on the program."
    },
    {
      "id": 427629,
      "postDate": "2018-11-25T22:08:45.547Z",
      "content": "<p>Congratulations and Many thanks for sharing.</p>",
      "rawMarkdown": "Congratulations and Many thanks for sharing."
    },
    {
      "id": 425105,
      "postDate": "2018-11-21T06:28:15.727Z",
      "content": "<p>Nice work!\nBecause you have used classification model first and then use detection model, I wonder how you can make sure that your classification model had trained enough that you can use detection to train the abnormal pics.</p>",
      "rawMarkdown": "Nice work!\nBecause you have used classification model first and then use detection model, I wonder how you can make sure that your classification model had trained enough that you can use detection to train the abnormal pics."
    },
    {
      "id": 425103,
      "postDate": "2018-11-21T06:27:42.753Z",
      "content": "<p>Nice work!Because you have used classification model first and then use detection model, I wonder how you can make sure that your classification model had trained enough that you can use detection to train the abnormal pics.</p>",
      "rawMarkdown": "Nice work!Because you have used classification model first and then use detection model, I wonder how you can make sure that your classification model had trained enough that you can use detection to train the abnormal pics."
    },
    {
      "id": 425102,
      "postDate": "2018-11-21T06:27:34.360Z",
      "content": "<p>Nice work!\nBecause you have used classification model first and then use detection model, I wonder how you can make sure that your classification model had trained enough that you can use detection to train the abnormal pics.</p>",
      "rawMarkdown": "Nice work!\nBecause you have used classification model first and then use detection model, I wonder how you can make sure that your classification model had trained enough that you can use detection to train the abnormal pics."
    },
    {
      "id": 425100,
      "postDate": "2018-11-21T06:27:10.033Z",
      "content": "<p>Nice work!Because you have used classification model first and then use detection model, I wonder how you can make sure that your classification model had trained enough that you can use detection to train the abnormal pics.</p>",
      "rawMarkdown": "Nice work!Because you have used classification model first and then use detection model, I wonder how you can make sure that your classification model had trained enough that you can use detection to train the abnormal pics."
    },
    {
      "id": 422512,
      "postDate": "2018-11-16T10:30:52.973Z",
      "content": "<p>Congrats and thanks for sharing! Very interesting read. </p>",
      "rawMarkdown": "Congrats and thanks for sharing! Very interesting read. "
    },
    {
      "id": 419117,
      "postDate": "2018-11-11T10:05:31.583Z",
      "content": "<p>This is great stuff!! Thanks for sharing.</p>",
      "rawMarkdown": "This is great stuff!! Thanks for sharing."
    },
    {
      "id": 416978,
      "postDate": "2018-11-07T15:01:37.187Z",
      "content": "<p>Congratulations and Thanks for sharing! Amazing achievement and awesome solution. </p>",
      "rawMarkdown": "Congratulations and Thanks for sharing! Amazing achievement and awesome solution. "
    },
    {
      "id": 416661,
      "postDate": "2018-11-07T03:41:33.680Z",
      "content": "<p>Congratulations! I didn't participate in the competition but it was very interesting to read your solution. I would like to try some of your approaches for other problems. Have you released your code yet? Thanks!</p>",
      "rawMarkdown": "Congratulations! I didn't participate in the competition but it was very interesting to read your solution. I would like to try some of your approaches for other problems. Have you released your code yet? Thanks!"
    },
    {
      "id": 416578,
      "postDate": "2018-11-06T22:44:16.390Z",
      "content": "<p>Congrats!Great work.</p>",
      "rawMarkdown": "Congrats!Great work."
    },
    {
      "id": 416100,
      "postDate": "2018-11-06T07:36:57.657Z",
      "content": "<p>congratulations and thanks for sharing.  Looking forward to your code.</p>",
      "rawMarkdown": "congratulations and thanks for sharing.  Looking forward to your code."
    },
    {
      "id": 415007,
      "postDate": "2018-11-04T05:23:55.063Z",
      "content": "<p>I noticed that you used the same <a href=\"https://github.com/ahrnbom/ensemble-objdet/\">ensemble code</a> as I did. Did you address this <a href=\"https://github.com/ahrnbom/ensemble-objdet/issues/2\">bug</a> at all? I did fix it in my code, but I just finally got around to submit a <a href=\"https://github.com/ahrnbom/ensemble-objdet/pull/3\">PR</a> for it. I wonder the accuracy drop you mentioned is partially due to this bug.</p>",
      "rawMarkdown": "I noticed that you used the same [ensemble code](https://github.com/ahrnbom/ensemble-objdet/) as I did. Did you address this [bug](https://github.com/ahrnbom/ensemble-objdet/issues/2) at all? I did fix it in my code, but I just finally got around to submit a [PR](https://github.com/ahrnbom/ensemble-objdet/pull/3) for it. I wonder the accuracy drop you mentioned is partially due to this bug.",
      "replies": [
        {
          "id": 415243,
          "postDate": "2018-11-04T18:27:24.597Z",
          "content": "<p>Thanks for the tip, we didn't address this bug and used the code as is. We will try fixing this in the code and seeing what happens.</p>",
          "rawMarkdown": "Thanks for the tip, we didn't address this bug and used the code as is. We will try fixing this in the code and seeing what happens."
        }
      ]
    },
    {
      "id": 414965,
      "postDate": "2018-11-04T01:41:14.447Z",
      "content": "<p>Well done! I can only imagine what the process of bundling all that up in a zip file for submitting to Kaggle in an intelligible way must have been like!</p>",
      "rawMarkdown": "Well done! I can only imagine what the process of bundling all that up in a zip file for submitting to Kaggle in an intelligible way must have been like!"
    },
    {
      "id": 414820,
      "postDate": "2018-11-03T16:58:24.263Z",
      "content": "<p>Congratulations Ian and Team! \nIbelieve resizing  lengths and widths of final predictions by 87.5% was one of the key differentiator. \nI have two questions:\nHow you arrived at final box thresold as 0.15?\nDid you use cloud or GT10XX machine?</p>",
      "rawMarkdown": "Congratulations Ian and Team! \nIbelieve resizing  lengths and widths of final predictions by 87.5% was one of the key differentiator. \nI have two questions:\nHow you arrived at final box thresold as 0.15?\nDid you use cloud or GT10XX machine?\n",
      "replies": [
        {
          "id": 414902,
          "postDate": "2018-11-03T21:29:51.087Z",
          "content": "<p>We tried 0.125, 0.15, 0.175 for final box thresholds because they gave us a final prevalence of 35-37% positive cases. 0.15 performed best on stage 1 public LB so we stuck with that.</p>\n\n<p>I used 2x 1080 Ti and Alexandre has 2x 1070. I also had access to a couple of Titan Vs that I used to train the classification models. </p>",
          "rawMarkdown": "We tried 0.125, 0.15, 0.175 for final box thresholds because they gave us a final prevalence of 35-37% positive cases. 0.15 performed best on stage 1 public LB so we stuck with that.\n\nI used 2x 1080 Ti and Alexandre has 2x 1070. I also had access to a couple of Titan Vs that I used to train the classification models. "
        }
      ]
    },
    {
      "id": 414750,
      "postDate": "2018-11-03T14:27:00.607Z",
      "content": "<p>Congratulations! The postprocessing mimicking annotation process is really awesome and amazing.</p>",
      "rawMarkdown": "Congratulations! The postprocessing mimicking annotation process is really awesome and amazing."
    },
    {
      "id": 414736,
      "postDate": "2018-11-03T13:37:46.870Z",
      "content": "<p>Congratulations guys, very impressive! I will take some time to digest this. </p>\n\n<p>Quick question: what lead you to realise that the triple checked boxes were smaller? IIRC, we didn't know which of the training examples had been triple-checked, but knew the positive test examples had. Did it just kinda click when you read that intersection of annotations was used to label? Wish I had noticed this.</p>\n\n<p>Also love the idea of concatenating images !</p>",
      "rawMarkdown": "Congratulations guys, very impressive! I will take some time to digest this. \n\nQuick question: what lead you to realise that the triple checked boxes were smaller? IIRC, we didn't know which of the training examples had been triple-checked, but knew the positive test examples had. Did it just kinda click when you read that intersection of annotations was used to label? Wish I had noticed this.\n\nAlso love the idea of concatenating images !\n",
      "replies": [
        {
          "id": 414737,
          "postDate": "2018-11-03T13:43:08.723Z",
          "content": "<p>Thanks! Alexandre noticed this first and experimented with resizing boxes and found a significant improvement with his early models. He anticipated that inter-rater agreement would be fairly low, so taking the intersection of boxes with 50% overlap would likely lead to a significant reduction in box size for triple-read studies. </p>\n\n<p>We had no idea which images in the training set were triple-read, but the vast majority were single-read, so we knew that the models would learn to predict larger boxes.  </p>",
          "rawMarkdown": "Thanks! Alexandre noticed this first and experimented with resizing boxes and found a significant improvement with his early models. He anticipated that inter-rater agreement would be fairly low, so taking the intersection of boxes with 50% overlap would likely lead to a significant reduction in box size for triple-read studies. \n\nWe had no idea which images in the training set were triple-read, but the vast majority were single-read, so we knew that the models would learn to predict larger boxes.  ",
          "votes": 1
        }
      ]
    },
    {
      "id": 414716,
      "postDate": "2018-11-03T12:58:54.870Z",
      "content": "<p>Congratulations and thanks for sharing. Amazing achievement for a 1st time in the object detection world.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing. Amazing achievement for a 1st time in the object detection world."
    },
    {
      "id": 2432513,
      "postDate": "2023-09-10T23:19:26.927Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 415006,
      "postDate": "2018-11-04T05:23:07.950Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 415005,
      "postDate": "2018-11-04T05:22:58.920Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 418103,
      "postDate": "2018-11-09T09:37:59.970Z",
      "content": "<p>Nice Work! Thanks for Sharing.</p>",
      "rawMarkdown": "Nice Work! Thanks for Sharing.",
      "votes": 4
    },
    {
      "id": 496413,
      "postDate": "2019-03-22T06:59:04.707Z",
      "content": "<p>Nice Work! Thanks for Sharing</p>",
      "rawMarkdown": "Nice Work! Thanks for Sharing"
    },
    {
      "id": 487206,
      "postDate": "2019-03-10T11:02:29.283Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 422761,
      "postDate": "2018-11-16T18:42:24.023Z",
      "content": "<p>Congratulations! Thanks for sharing!</p>",
      "rawMarkdown": "Congratulations! Thanks for sharing!"
    },
    {
      "id": 420367,
      "postDate": "2018-11-13T14:01:14.047Z",
      "content": "<p>Congrats and thanks for sharing.</p>",
      "rawMarkdown": "Congrats and thanks for sharing."
    },
    {
      "id": 417208,
      "postDate": "2018-11-08T00:02:56.843Z",
      "content": "<p>Great work and thanks for sharing!</p>",
      "rawMarkdown": "Great work and thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 417614,
      "author_name": "Ian Pan",
      "author_url": "",
      "post_date": "2018-11-08T14:23:30.810000",
      "content": "<p>We have uploaded our solution code to GitHub. Please find it here: <a href=\"https://www.github.com/i-pan/kaggle-rsna18\">https://www.github.com/i-pan/kaggle-rsna18</a></p>\n\n<p>If you have any issues, please let us know! </p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 414754,
      "author_name": "Dmytro Poplavskiy",
      "author_url": "",
      "post_date": "2018-11-03T14:35:58.087000",
      "content": "<p>Congratulations with the 1st place!</p>\n\n<p>I scaled the size of \"harder\" boxes as well, but slightly less and proportionally to variance (actually difference between 20 and 80 percentile) between folds predictions. The intuition behind - to simulate the labeling process of test samples when intersection used from different radiologists labels.</p>\n\n<p>my calculations for adjusted size, h_perc20 is 20 percentile of anchor h between folds and checkpoints:\nh = h_perc20 - 1.6 * (h_perc80 - h_perc20)\nw = w_perc20 - 1.6 * (w_perc80 - w_perc20)</p>\n\n<p>box size mean 67632.9\nbox size median 65335.1</p>",
      "votes": 1,
      "replies": [
        {
          "id": 414904,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2018-11-03T21:32:00.563000",
          "content": "<p>Interesting - did you ever try doing a fixed resizing of the boxes as a comparison? We also wanted to account for variance which we tried to simulate using our intersection method, but it actually performed worse. Your method might be better than that though. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 415002,
          "author_name": "Alexandre Cadrin-Chênevert",
          "author_url": "",
          "post_date": "2018-11-04T05:08:56.900000",
          "content": "<p>The intuition when we tried to mimick the intersection from N different annotators was to catch a different variance for each border. Practically, the annotators variance for each border can potentially be asymetric if a lung opacity is visually well defined on a specific border but undefined on other borders. But as Ian specified, simple linear resize just worked better because variance was probably, on average, relatively symetric for all borders on the stage 1 and stage 2 test datasets</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415035,
          "author_name": "Dmytro Poplavskiy",
          "author_url": "",
          "post_date": "2018-11-04T08:24:37.513000",
          "content": "<p>I planned to test the similar scale reduction as you have done (I even started to implement it first), but run out of time. I have just tried to do a submission with boxes size reduced by 87.5% and received a close but better score (0.247 -&gt; 0.249 on stage 2)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 416643,
      "author_name": "Phillip Cheng",
      "author_url": "",
      "post_date": "2018-11-07T02:59:54.667000",
      "content": "<p>Congratulations Ian and Alex on your first place achievement!  Your solution is an optimization tour de force, I'm thoroughly impressed and look forward to exploring your solution in depth.  I came up with only a small subset of your insights, though it looks like we did similar bounding box resizing and box coordinate averaging.  I found that a slightly larger bounding box size reduction improved my scores, perhaps because I included rotations in my image augmentation.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3304997,
      "author_name": "Muhammad Ehsan",
      "author_url": "",
      "post_date": "2025-10-21T19:42:56.813000",
      "content": "<p>Many congratulations 🎉🎉</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 440150,
      "author_name": "nbajwa",
      "author_url": "",
      "post_date": "2018-12-17T06:15:19.417000",
      "content": "<p>Hi Ian!</p>\n\n<p>Thanks for sharing. I'm getting the following error when trying to replicate your work. I'm unable to transform data into the COCO format:</p>\n\n<p>Transforming data into COCO format ...</p>\n\n<pre><code>usage: 6_COCOify.py [-h]\n                    subset [TRAIN_LABELS_PATH] [TRAIN_IMAGES_DIR]\n                    [TEST_IMAGES_DIR] [FOLDS_DF_PATH]\n6_COCOify.py: error: too few arguments\n</code></pre>\n\n<p>I've tried resolving it but nothings working. I'll appreciate if you can help me out.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 436156,
      "author_name": "Flo Wallny",
      "author_url": "",
      "post_date": "2018-12-09T17:45:43.430000",
      "content": "<p>Congratulations :) did you alter the image resolution for classification?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 428140,
      "author_name": "Yee Ng",
      "author_url": "",
      "post_date": "2018-11-26T20:32:12.127000",
      "content": "<p>Anyone know the venue and time when these solutions are presented at RSNA? I've purchased the virtual meeting but couldn't find any session pertaining to this challenge on the program.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 427629,
      "author_name": "Jilja Joy",
      "author_url": "",
      "post_date": "2018-11-25T22:08:45.547000",
      "content": "<p>Congratulations and Many thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 425105,
      "author_name": "jacken312",
      "author_url": "",
      "post_date": "2018-11-21T06:28:15.727000",
      "content": "<p>Nice work!\nBecause you have used classification model first and then use detection model, I wonder how you can make sure that your classification model had trained enough that you can use detection to train the abnormal pics.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 425103,
      "author_name": "jacken312",
      "author_url": "",
      "post_date": "2018-11-21T06:27:42.753000",
      "content": "<p>Nice work!Because you have used classification model first and then use detection model, I wonder how you can make sure that your classification model had trained enough that you can use detection to train the abnormal pics.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 425102,
      "author_name": "jacken312",
      "author_url": "",
      "post_date": "2018-11-21T06:27:34.360000",
      "content": "<p>Nice work!\nBecause you have used classification model first and then use detection model, I wonder how you can make sure that your classification model had trained enough that you can use detection to train the abnormal pics.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 425100,
      "author_name": "jacken312",
      "author_url": "",
      "post_date": "2018-11-21T06:27:10.033000",
      "content": "<p>Nice work!Because you have used classification model first and then use detection model, I wonder how you can make sure that your classification model had trained enough that you can use detection to train the abnormal pics.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 422512,
      "author_name": "Christian Bluethgen",
      "author_url": "",
      "post_date": "2018-11-16T10:30:52.973000",
      "content": "<p>Congrats and thanks for sharing! Very interesting read. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 419117,
      "author_name": "Keshav",
      "author_url": "",
      "post_date": "2018-11-11T10:05:31.583000",
      "content": "<p>This is great stuff!! Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 416978,
      "author_name": "Fatih",
      "author_url": "",
      "post_date": "2018-11-07T15:01:37.187000",
      "content": "<p>Congratulations and Thanks for sharing! Amazing achievement and awesome solution. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 416661,
      "author_name": "Gopsy",
      "author_url": "",
      "post_date": "2018-11-07T03:41:33.680000",
      "content": "<p>Congratulations! I didn't participate in the competition but it was very interesting to read your solution. I would like to try some of your approaches for other problems. Have you released your code yet? Thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 416578,
      "author_name": "Muhammed Gulaydin",
      "author_url": "",
      "post_date": "2018-11-06T22:44:16.390000",
      "content": "<p>Congrats!Great work.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 416100,
      "author_name": "chenglong",
      "author_url": "",
      "post_date": "2018-11-06T07:36:57.657000",
      "content": "<p>congratulations and thanks for sharing.  Looking forward to your code.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 415007,
      "author_name": "Peter Yu",
      "author_url": "",
      "post_date": "2018-11-04T05:23:55.063000",
      "content": "<p>I noticed that you used the same <a href=\"https://github.com/ahrnbom/ensemble-objdet/\">ensemble code</a> as I did. Did you address this <a href=\"https://github.com/ahrnbom/ensemble-objdet/issues/2\">bug</a> at all? I did fix it in my code, but I just finally got around to submit a <a href=\"https://github.com/ahrnbom/ensemble-objdet/pull/3\">PR</a> for it. I wonder the accuracy drop you mentioned is partially due to this bug.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 415243,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2018-11-04T18:27:24.597000",
          "content": "<p>Thanks for the tip, we didn't address this bug and used the code as is. We will try fixing this in the code and seeing what happens.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 414965,
      "author_name": "Tim H",
      "author_url": "",
      "post_date": "2018-11-04T01:41:14.447000",
      "content": "<p>Well done! I can only imagine what the process of bundling all that up in a zip file for submitting to Kaggle in an intelligible way must have been like!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 414820,
      "author_name": "PrasunMishra",
      "author_url": "",
      "post_date": "2018-11-03T16:58:24.263000",
      "content": "<p>Congratulations Ian and Team! \nIbelieve resizing  lengths and widths of final predictions by 87.5% was one of the key differentiator. \nI have two questions:\nHow you arrived at final box thresold as 0.15?\nDid you use cloud or GT10XX machine?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 414902,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2018-11-03T21:29:51.087000",
          "content": "<p>We tried 0.125, 0.15, 0.175 for final box thresholds because they gave us a final prevalence of 35-37% positive cases. 0.15 performed best on stage 1 public LB so we stuck with that.</p>\n\n<p>I used 2x 1080 Ti and Alexandre has 2x 1070. I also had access to a couple of Titan Vs that I used to train the classification models. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 414750,
      "author_name": "OsciiArt",
      "author_url": "",
      "post_date": "2018-11-03T14:27:00.607000",
      "content": "<p>Congratulations! The postprocessing mimicking annotation process is really awesome and amazing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 414736,
      "author_name": "Tom Aindow",
      "author_url": "",
      "post_date": "2018-11-03T13:37:46.870000",
      "content": "<p>Congratulations guys, very impressive! I will take some time to digest this. </p>\n\n<p>Quick question: what lead you to realise that the triple checked boxes were smaller? IIRC, we didn't know which of the training examples had been triple-checked, but knew the positive test examples had. Did it just kinda click when you read that intersection of annotations was used to label? Wish I had noticed this.</p>\n\n<p>Also love the idea of concatenating images !</p>",
      "votes": 0,
      "replies": [
        {
          "id": 414737,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2018-11-03T13:43:08.723000",
          "content": "<p>Thanks! Alexandre noticed this first and experimented with resizing boxes and found a significant improvement with his early models. He anticipated that inter-rater agreement would be fairly low, so taking the intersection of boxes with 50% overlap would likely lead to a significant reduction in box size for triple-read studies. </p>\n\n<p>We had no idea which images in the training set were triple-read, but the vast majority were single-read, so we knew that the models would learn to predict larger boxes.  </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 414716,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2018-11-03T12:58:54.870000",
      "content": "<p>Congratulations and thanks for sharing. Amazing achievement for a 1st time in the object detection world.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2432513,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-10T23:19:26.927000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 415006,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-04T05:23:07.950000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 415005,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-04T05:22:58.920000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 418103,
      "author_name": "Arunkumar Venkataramanan",
      "author_url": "",
      "post_date": "2018-11-09T09:37:59.970000",
      "content": "<p>Nice Work! Thanks for Sharing.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 496413,
      "author_name": "kevin wu",
      "author_url": "",
      "post_date": "2019-03-22T06:59:04.707000",
      "content": "<p>Nice Work! Thanks for Sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 487206,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-03-10T11:02:29.283000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 422761,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-16T18:42:24.023000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 420367,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-13T14:01:14.047000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 417208,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-08T00:02:56.843000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "414711": "We are very excited to have been awarded 1st place in this competition. This was our first foray into the world of object detection. I'm Ian, currently a 3rd-year medical student at Brown University, Providence, RI, USA. My teammate Alexandre is a radiologist practicing in Lanaudière, Québec, Canada. We are happy to represent the medical community in this competition. Huge congratulations to Dmytro for becoming Grandmaster. His top score during the public LB really motivated us to make successive improvements to our model. Truly an honor to be in the same ranks as him for this challenge. Congratulations to all of the other top 10 winners as well. \n\n# tl;dr \n- We did NOT retrain when stage 1 test labels were released\n- We used a classification-detection pipeline \n- 10-fold CV ensemble for classification, combination of 5 10-fold CV ensembles for detection (50 models) \n- For detection, we used:\n1. RetinaNet: https://github.com/fizyr/keras-retinanet\n2. Deformable R-FCN: https://github.com/msracver/Deformable-ConvNets\n3. Deformable Relation Networks: https://github.com/msracver/Relation-Networks-for-Object-Detection\n- Boxes were ensembled using: https://github.com/ahrnbom/ensemble-objdet\n- *We resized the lengths and widths of our final predictions by 87.5%*\n- Code available here: https://www.github.com/i-pan/kaggle-rsna18\n- We found that a 6-model ensemble with 1 InceptionResNetV2 for classification and 5 Deformable Relation Networks achieved 0.253 on stage 2 private LB\n\n# Classification \nWe used Keras 2.2 for classification. Our classification model was composed of the following individual models [format: modelArchitecture (numClasses) (imgSize)]. Models were trained on either 2 classes (opacity vs. not) or 3 classes (opacity vs. not normal/no opacity vs. normal). Each model was trained on a different fold.\n- InceptionResNetV2 (2) (256), InceptionResNetV2 (2) (320) \n- InceptionResNetV2 (3) (256), InceptionResNetV2 (3) (320) \n- Xception (2) (384), Xception (2) (448) \n- Xception (3) (384), Xception (3) (448) \n- DenseNet169 (2) (512), DenseNet169 (3) (512)\nFirst, we trained ImageNet pre-trained networks on the NIH ChestX-ray14 dataset using 15 classes (14 findings + abnormal vs. normal). Then we fine-tuned those weights on the pneumonia dataset. This improved results from training using ImageNet weights only by about 1% locally. We were getting about 0.88-0.90 AUC across our folds. \n\nModels were trained with 50% probability of being color-inverted, 50% of being flipped, and 50% of being augmented with some other augmentation (e.g. contrast enhancement, crop, rotation). We used 15x TTA per model for our final predictions and averaged the 150 predictions for our final classification score. When the stage 1 test labels were released, our classification ensemble had an AUC of 0.93. \n\nAs many other competitors noticed, the distribution of the training (single-read) data and test (triple-read) data were quite different. We think that thoracic fellowship-trained radiologists from the STR and the increased number of readers contributed to increased sensitivity or higher clinical suspicion of opacities that led to the increase in prevalence. We tuned our threshold based on the stage 1 public LB results.\n\n# Detection\nPlease see the tl;dr for the repos we used for detection. We used a combination of 5 10-fold CV ensembles for detection. We computed the metric provided by Yicheng Chen (https://www.kaggle.com/chenyc15/mean-average-precision-metric) for both positive images only and all images to inform our model selection. Unfortunately we did a poor job of keeping track of how well our experiments did on stage 1 public LB, and the results are no longer visible. \n\n## Detection Ensemble 1\nEach fold was trained on a different resolution (224-512 by increments of 32). *Positive images only.* We found that lower resolutions did not lower performance locally and actually increased performance (0.005-0.008) on public LB, so we stuck with it. It also allowed us to spam models in our ensemble due to lower training/inference overhead. The first detection ensemble was a 10-fold CV ensemble of deformable R-FCN. We mainly used default parameters, except we unfroze the non-BN layers that are frozen in the default configs. These models use a ResNet101 ImageNet pre-trained backbone. We also changed the max number of detections per image to 5 but kept a threshold of 0.001. Only data augmentation used was flip. These models trained very quickly (~2 hours) depending on the image resolution. \n\n## Detection Ensemble 2\nBasically the same as #1, except we used deformable relation networks (https://arxiv.org/abs/1711.11575). Honestly, I don't really know how these work, and I stumbled upon them late into the competition. Since it was basically the same as the deformable R-FCN repo, it was easy to train these models and they performed very well. We tried the version where you attempt to learn the NMS thresholds to use, but that performed very poorly. \n\n## Detection Ensemble 3\nExactly the same as #2, except we used the default config (i.e. kept the backbone layers frozen). Keeping the backbone layers frozen had a slight decrease in performance, but we threw in this ensemble to decrease model correlation with the other ensembles, and it had a slight improvement in our stage 1 public LB score. \n\n## Detection Ensemble 4\nThis is a RetinaNet ensemble, also 10-fold CV, but trained only at 384 x 384 resolution because we had issues with changing anchor sizes. I believe a recent update allows you to specify a config.ini file where these changes can be made. This was probably our favorite ensemble. We trained on *concatenated* images where each negative image was randomly concatenated with a positive image from the same fold on the left or right (so final image sizes were 384 x 768). We trained for 8000 steps/epoch, batch size 1, for 8 epochs dividing learning rate by 10 after epoch 4 and 6. Validation was performed after each epoch. We selected the model with the best mAP metric over the 8 epochs. Using the `--random-transform` argument didn't really help/hurt but seemed to make training more unstable so we didn't use it (without specifying this, only data augmentation is flip). Half of the 10-fold CV ensemble was trained using ResNet101 backbone, the other half with ResNet152. ResNet101 was clearly better than ResNet50, but ResNet152 was the same as ResNet101. \n\nTraining on concatenated images allowed RetinaNet to better balance precision and recall. You can achieve the same results by picking the right proportion of positives/negatives as well, but this seemed more \"clean\" to us. Inference was performed on single images. We looked at the AUC of these models using the max box score as class prediction and it was on par with our classification networks. The mAP on positive images only went down, but the mAP on all images went up, so this trade-off was beneficial to include in our ensemble. If we wanted to use a single model, it would be this one as it does not need a classifier to achieve good performance. Interestingly, training on concatenated images did not work well for detection ensembles 1-3. This may be because they are 2-stage detectors, but we didn't have time to look at this more closely. \n\n## Detection Ensemble 5\nSame as #4, except trained on *positive images only.* \n\n# Ensemble\n\n## Detection Ensemble 1+2+3\nFor detection ensembles 1-3, we applied 6x TTA (original, flipped, 80%/120% for both original and flipped). For each of the 30 models, we combined the TTA predictions using (https://github.com/ahrnbom/ensemble-objdet) with IoU threshold 0.4. Box score predictions were then adjusted by multiplying by the fraction of TTAs that contained that box (i.e. if 5/6 TTAs predicted a box with average score 0.5, it was multiplied by 5/6). This code expects *center coordinates* when computing the IoU overlap between boxes in the `getCoords` function. We didn't realize this until late in the competition (we were using top left), and when we changed this to take in top left coordinates, there was a drop in performance of about 0.01 in public LB. If anyone can help us figure out why, we haven't solved this yet. \n\nNow we have 30 models worth of predictions, so we combined those again with IoU threshold 0.4. No weighting was used. Box score predictions were then adjusted by multiplying by the fraction of models that contained that box (i.e. if 24/30 models predicted a box with average score 0.5, it was multiplied by 0.8 and the score was adjusted to 0.4). \n\nTo incorporate the classification network, we multiplied the ensemble-averaged box score predictions by the classification score for that image. We eliminated boxes with an adjusted score of &lt;0.225. \n\n## Detection Ensemble 4+5\nWe applied 10x TTA (resolutions 320, 352, 384, 416, 448 for both original/flipped), performing the same kind of ensembling as above, including score adjustment based on fraction of TTAs predicting that box. Top 10 (or fewer) detections per TTA were selected. However we did not combine the predictions across 20 models in both ensembles at this stage. Instead, we used a classification score threshold of 0.2 and box score threshold of 0.3 for detection ensemble 4. The multiplication method we used above did not work well for RetinaNet. Though we said detection ensemble 4 didn't necessarily need a paired classifier, it did improve the stage 1 public LB score by ~0.003 so we used it since it was readily available. For detection ensemble 5, we used a classification score threshold of 0.325 and box score threshold of 0.35. \n\nAfter applying these thresholds, we combined boxes from ensembles 4 and 5 using the same code and IoU threshold 0.4. In this case, detection ensemble 4 was given 1.2 weight versus 0.8 for ensemble 5. This was because ensemble 4 performed slightly better (~0.004) than ensemble 5. Then we adjusted the score using the same strategy as above (in this case, if a box was only present in one model, score was divided in half). \n\nFor tuning these thresholds, we aimed for a prevalence of about 35-37% and experimented with what worked best on stage 1 public LB. We didn't have any real local validation because we realized early on that LB score would be the best indicator of model performance on the final test data. \n\n# Final Ensemble\nWe are now left with 2 ensembles. We combined them using equal weight and IoU threshold as described above, using the same adjustment strategy. Final box threshold used was 0.15.\n\n# Post-processing \nWe realized that the triple-read boxes were smaller than the single-read boxes. This makes sense because in the annotation process described here (https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723) the *intersection* of boxes was used as opposed to the average. We tried to mimic this intersection process by taking the average of multiple intersections of boxes across models. This gave ~10-15% (!) improvement on stage 1 public LB. As an alternative to this, we simply resized the boxes by multiplying length/width by a fraction. We found 87.5% for each was a good reduction and worked a bit better than doing the intersection. It was also much easier to implement. We discovered this early on, and didn't submit anything without resizing after that. Early on in the competition we had an improvement from 0.181-&gt;0.209 with resizing. Towards the end, I wanted to see the effect of resizing again and it was an improvement of 0.218-&gt;0.252, which is huge. For our final stage 1 submission, resize improved our score from 0.222-&gt;0.260. It seemed like other people were seeing better success with higher resolution models, so we wonder if the resizing was more complementary with lower resolution models. Maybe if we didn't resize, higher resolution models would perform better.\n\n# Statistics for final submission\n\n    Stage 1: \n    NUM POSITIVES: 351 \n    % POSITIVES: 0.351 \n    NUM BOXES: 582 \n    AVG # BOXES PER CASE: 1.65811965812 \n    BOX SIZE [MEDIAN]: 60084.5 \n    BOX SIZE [MEAN]: 66052.2405498 \n    BOX SIZE [MIN]: 16065 \n    BOX SIZE [MAX]: 151580\n    \n    Stage 2: \n    NUM POSITIVES: 1106 \n    % POSITIVES: 0.368666666667 \n    NUM BOXES: 1830 \n    AVG # BOXES PER CASE: 1.65461121157 \n    BOX SIZE [MEDIAN]: 60877.5\n    BOX SIZE [MEAN]: 66454.8060109 \n    BOX SIZE [MIN]: 13431 \n    BOX SIZE [MAX]: 192648\n\n# Miscellaneous thoughts\nBasically we had a good model that we turned into a \"great\" model by resizing the output. It will be interesting to see if other top teams achieved their score by improving more upon classification versus box precision and how their scores would be affected by resize if they did not apply any post-processing to their predictions. We went overboard with ensembles and probably could have achieved the same performance with &lt;20 models, but it became so easy to train them that we just included a bunch. We could have experimented with more hyperparameters in our detection models as well. I really dislike hyperparameter tuning (never really developed a good strategy for it) and often try and compensate by ensembling different models together. \n\n# Things that didn't work \n- We tried training another classifier on out-of-fold bounding box predictions produced by our detection models to classify into IoU &gt;0.4 and &lt;0.4\n- NMS of overlapping bounding boxes in our final submission: a number of images in our final submission for both stage 1 and stage 2 LBs had overlapping boxes (usually a smaller one contained in a larger one). Suppressing these actually reduced our score. \n- We tried various experiments that treated AP/PA images differently (e.g. different thresholds, different resizes), but in the end it was easier and better to be view-agnostic\n- Getting all bounding box predictions from all TTAs and models and combining them at once. This didn't work as well as the stepwise approach we described above. This may have something to do with non-standardized prediction scores (though we tried standardizing and it didn't help much). \n- SoftNMS -- this makes sense because you would not expect overlapping objects in this challenge as opposed to others like COCO. \n- For RetinaNet: other backbones. Only the ResNet backbones worked well for us. \n- For RetinaNet: pre-training detector heads. It was a lot easier to use ImageNet pre-trained backbones and then just start training the whole network from the beginning.\n- For RetinaNet: pre-training the backbone on the pneumonia dataset. No improvement. \n",
    "417614": "We have uploaded our solution code to GitHub. Please find it here: https://www.github.com/i-pan/kaggle-rsna18\n\nIf you have any issues, please let us know! ",
    "414754": "Congratulations with the 1st place!\n\nI scaled the size of \"harder\" boxes as well, but slightly less and proportionally to variance (actually difference between 20 and 80 percentile) between folds predictions. The intuition behind - to simulate the labeling process of test samples when intersection used from different radiologists labels.\n\nmy calculations for adjusted size, h_perc20 is 20 percentile of anchor h between folds and checkpoints:\nh = h_perc20 - 1.6 * (h_perc80 - h_perc20)\nw = w_perc20 - 1.6 * (w_perc80 - w_perc20)\n\nbox size mean 67632.9\nbox size median 65335.1",
    "416643": "Congratulations Ian and Alex on your first place achievement!  Your solution is an optimization tour de force, I'm thoroughly impressed and look forward to exploring your solution in depth.  I came up with only a small subset of your insights, though it looks like we did similar bounding box resizing and box coordinate averaging.  I found that a slightly larger bounding box size reduction improved my scores, perhaps because I included rotations in my image augmentation.",
    "3304997": "Many congratulations 🎉🎉",
    "440150": "Hi Ian!\n\nThanks for sharing. I'm getting the following error when trying to replicate your work. I'm unable to transform data into the COCO format:\n\nTransforming data into COCO format ...\n\n    usage: 6_COCOify.py [-h]\n                        subset [TRAIN_LABELS_PATH] [TRAIN_IMAGES_DIR]\n                        [TEST_IMAGES_DIR] [FOLDS_DF_PATH]\n    6_COCOify.py: error: too few arguments\n\nI've tried resolving it but nothings working. I'll appreciate if you can help me out.",
    "436156": "Congratulations :) did you alter the image resolution for classification?",
    "428140": "Anyone know the venue and time when these solutions are presented at RSNA? I've purchased the virtual meeting but couldn't find any session pertaining to this challenge on the program.",
    "427629": "Congratulations and Many thanks for sharing.",
    "425105": "Nice work!\nBecause you have used classification model first and then use detection model, I wonder how you can make sure that your classification model had trained enough that you can use detection to train the abnormal pics.",
    "425103": "Nice work!Because you have used classification model first and then use detection model, I wonder how you can make sure that your classification model had trained enough that you can use detection to train the abnormal pics.",
    "425102": "Nice work!\nBecause you have used classification model first and then use detection model, I wonder how you can make sure that your classification model had trained enough that you can use detection to train the abnormal pics.",
    "425100": "Nice work!Because you have used classification model first and then use detection model, I wonder how you can make sure that your classification model had trained enough that you can use detection to train the abnormal pics.",
    "422512": "Congrats and thanks for sharing! Very interesting read. ",
    "419117": "This is great stuff!! Thanks for sharing.",
    "416978": "Congratulations and Thanks for sharing! Amazing achievement and awesome solution. ",
    "416661": "Congratulations! I didn't participate in the competition but it was very interesting to read your solution. I would like to try some of your approaches for other problems. Have you released your code yet? Thanks!",
    "416578": "Congrats!Great work.",
    "416100": "congratulations and thanks for sharing.  Looking forward to your code.",
    "415007": "I noticed that you used the same [ensemble code](https://github.com/ahrnbom/ensemble-objdet/) as I did. Did you address this [bug](https://github.com/ahrnbom/ensemble-objdet/issues/2) at all? I did fix it in my code, but I just finally got around to submit a [PR](https://github.com/ahrnbom/ensemble-objdet/pull/3) for it. I wonder the accuracy drop you mentioned is partially due to this bug.",
    "414965": "Well done! I can only imagine what the process of bundling all that up in a zip file for submitting to Kaggle in an intelligible way must have been like!",
    "414820": "Congratulations Ian and Team! \nIbelieve resizing  lengths and widths of final predictions by 87.5% was one of the key differentiator. \nI have two questions:\nHow you arrived at final box thresold as 0.15?\nDid you use cloud or GT10XX machine?\n",
    "414750": "Congratulations! The postprocessing mimicking annotation process is really awesome and amazing.",
    "414736": "Congratulations guys, very impressive! I will take some time to digest this. \n\nQuick question: what lead you to realise that the triple checked boxes were smaller? IIRC, we didn't know which of the training examples had been triple-checked, but knew the positive test examples had. Did it just kinda click when you read that intersection of annotations was used to label? Wish I had noticed this.\n\nAlso love the idea of concatenating images !\n",
    "414716": "Congratulations and thanks for sharing. Amazing achievement for a 1st time in the object detection world.",
    "2432513": "",
    "415006": "",
    "415005": "",
    "418103": "Nice Work! Thanks for Sharing.",
    "496413": "Nice Work! Thanks for Sharing",
    "487206": "Thanks for sharing!",
    "422761": "Congratulations! Thanks for sharing!",
    "420367": "Congrats and thanks for sharing.",
    "417208": "Great work and thanks for sharing!"
  }
}