{
  "id": 71022,
  "title": "[10th place] Solution Overview and Code",
  "url": "/competitions/rsna-pneumonia-detection-challenge/writeups/dasa-fidi-iara-10th-place-solution-overview-and-co",
  "author_name": "",
  "post_date": "2018-11-09T11:37:43.863Z",
  "votes": 18,
  "comment_count": 5,
  "views": 0,
  "content": "<p>This is an overview of our team solution to the pneumonia detection challenge. Our team, DASA-FIDI-IARA, is composed of computer scientists, a computer engineer,  a radiologist informatician, and a neuroradiologist. We had much fun participating on this challenge, and we want to thank organizers and all people that worked hard to make it possible. Code and trained models are available on GitHub (link in the end).</p>\n\n<h1>Summary</h1>\n\n<p>Our solution is based on <a href=\"https://github.com/fizyr/keras-retinanet\">retinanet</a> with resnet101 as backbone. We trained 5 folds of the training set from stage 1 and used the best single one. We experimented many ‘out-of-the-box’ ideas instead of implementing ensemble, making the stage 1 submission to be a single model. We have heavily used data augmentation, both on images and bounding boxes. After a careful analysis of width and height ratios of bounding boxes, we have tried many anchor ratios setups. Also, we played with inference hyperparameters, such as non-maximum suppression and score thresholds.</p>\n\n<h1>Training</h1>\n\n<p>We resized the original images to 800x800, as we got worse results for smaller sizes. We have modified the original retinanet implementation, including the following features:</p>\n\n<ol>\n<li>Adding more anchor ratios, in a total of 9, ranging from 0.25 to 4;</li>\n<li>Hyperparameter tuning: learning rate on training, non-maximum suppression and score threshold on inference;</li>\n<li>Data augmentation\n<ul><li>On images and bounding boxes: rotations, horizontal flips, translations, shearing and scaling</li>\n<li>On images: contrast and brightness increase/decrease, gaussian noise</li>\n<li>On bounding box: shifts on points (x1,y1) and (x2,y2); scaling up/down bounding boxes (this feature was not actually used in our model submission)</li></ul></li>\n<li>Dropouts both for regression and classifier heads;</li>\n<li>We had information from the discussion in Kaggle that the annotation was performed in a different way for the training and test sets. So, we guessed how different it could be and trained our best model with the same training set, except by progressively reducing bounding boxes. The best model was trained with 15% reduction in height and width.</li>\n</ol>\n\n<h1>Ideas that did not work</h1>\n\n<p>We tried a lot of different things that did not work so well for us:</p>\n\n<ol>\n<li>Resnet52, resnet152, vgg, mobilenet, and even custom architectures as backbones;</li>\n<li>Adding more anchor boxes;</li>\n<li>Test Time Augmentation;</li>\n<li>YOLO and Light-Head RCNN;</li>\n<li>Freezing first layers of backbone;</li>\n<li>Histogram equalization as pre-processing;</li>\n<li>Two step approach: classifier + detector (pneumonia vs others)</li>\n<li>Adding gender and view position as features.</li>\n<li>Training specified networks for each gender (M/F) and each view position (AP/PA)</li>\n<li>Manually re-annotate the training set (not-normal cases) to remove evident false positives/negatives.</li>\n</ol>\n\n<h1>Final Thoughts</h1>\n\n<p>As mentioned, our Stage 1 submission was a single model, and we believe that ensemble could improve our score. For instance, after the final submission, we implemented an ensemble of 2 of our best models (just for curiosity), achieving 0.254 score in the Stage 1 testset. </p>\n\n<p>The most important realization was finding that the annotation of the test set was different than the training set. Specifically, we found that targeting smaller bounding boxes yielded better results. The use of additional anchor-ratios also improved our mAP significantly. Our code for training and inference, as well as the trained models, are available on GitHub. There is also a demo where you can test our model (you need to request an account to try it). All the links can be found below.</p>\n\n<h1>Links</h1>\n\n<ul>\n<li>Documentation: <a href=\"https://rsna-challenge-2018.readthedocs.io/en/latest/\">https://rsna-challenge-2018.readthedocs.io/en/latest/</a></li>\n<li>Training/inference code: <a href=\"https://github.com/alessonscap/rsna-challenge-2018.git\">https://github.com/alessonscap/rsna-challenge-2018.git</a></li>\n<li>Demo page: <a href=\"https://iarahealth.com/pneumonia/\">https://iarahealth.com/pneumonia/</a></li>\n</ul>",
  "messages": [
    {
      "id": "418134",
      "postDate": "11/09/2018 11:01:04",
      "content": "<p>This is an overview of our team solution to the pneumonia detection challenge. Our team, DASA-FIDI-IARA, is composed of computer scientists, a computer engineer,  a radiologist informatician, and a neuroradiologist. We had much fun participating on this challenge, and we want to thank organizers and all people that worked hard to make it possible. Code and trained models are available on GitHub (link in the end).</p>\n\n<h1>Summary</h1>\n\n<p>Our solution is based on <a href=\"https://github.com/fizyr/keras-retinanet\">retinanet</a> with resnet101 as backbone. We trained 5 folds of the training set from stage 1 and used the best single one. We experimented many ‘out-of-the-box’ ideas instead of implementing ensemble, making the stage 1 submission to be a single model. We have heavily used data augmentation, both on images and bounding boxes. After a careful analysis of width and height ratios of bounding boxes, we have tried many anchor ratios setups. Also, we played with inference hyperparameters, such as non-maximum suppression and score thresholds.</p>\n\n<h1>Training</h1>\n\n<p>We resized the original images to 800x800, as we got worse results for smaller sizes. We have modified the original retinanet implementation, including the following features:</p>\n\n<ol>\n<li>Adding more anchor ratios, in a total of 9, ranging from 0.25 to 4;</li>\n<li>Hyperparameter tuning: learning rate on training, non-maximum suppression and score threshold on inference;</li>\n<li>Data augmentation\n<ul><li>On images and bounding boxes: rotations, horizontal flips, translations, shearing and scaling</li>\n<li>On images: contrast and brightness increase/decrease, gaussian noise</li>\n<li>On bounding box: shifts on points (x1,y1) and (x2,y2); scaling up/down bounding boxes (this feature was not actually used in our model submission)</li></ul></li>\n<li>Dropouts both for regression and classifier heads;</li>\n<li>We had information from the discussion in Kaggle that the annotation was performed in a different way for the training and test sets. So, we guessed how different it could be and trained our best model with the same training set, except by progressively reducing bounding boxes. The best model was trained with 15% reduction in height and width.</li>\n</ol>\n\n<h1>Ideas that did not work</h1>\n\n<p>We tried a lot of different things that did not work so well for us:</p>\n\n<ol>\n<li>Resnet52, resnet152, vgg, mobilenet, and even custom architectures as backbones;</li>\n<li>Adding more anchor boxes;</li>\n<li>Test Time Augmentation;</li>\n<li>YOLO and Light-Head RCNN;</li>\n<li>Freezing first layers of backbone;</li>\n<li>Histogram equalization as pre-processing;</li>\n<li>Two step approach: classifier + detector (pneumonia vs others)</li>\n<li>Adding gender and view position as features.</li>\n<li>Training specified networks for each gender (M/F) and each view position (AP/PA)</li>\n<li>Manually re-annotate the training set (not-normal cases) to remove evident false positives/negatives.</li>\n</ol>\n\n<h1>Final Thoughts</h1>\n\n<p>As mentioned, our Stage 1 submission was a single model, and we believe that ensemble could improve our score. For instance, after the final submission, we implemented an ensemble of 2 of our best models (just for curiosity), achieving 0.254 score in the Stage 1 testset. </p>\n\n<p>The most important realization was finding that the annotation of the test set was different than the training set. Specifically, we found that targeting smaller bounding boxes yielded better results. The use of additional anchor-ratios also improved our mAP significantly. Our code for training and inference, as well as the trained models, are available on GitHub. There is also a demo where you can test our model (you need to request an account to try it). All the links can be found below.</p>\n\n<h1>Links</h1>\n\n<ul>\n<li>Documentation: <a href=\"https://rsna-challenge-2018.readthedocs.io/en/latest/\">https://rsna-challenge-2018.readthedocs.io/en/latest/</a></li>\n<li>Training/inference code: <a href=\"https://github.com/alessonscap/rsna-challenge-2018.git\">https://github.com/alessonscap/rsna-challenge-2018.git</a></li>\n<li>Demo page: <a href=\"https://iarahealth.com/pneumonia/\">https://iarahealth.com/pneumonia/</a></li>\n</ul>",
      "rawMarkdown": "This is an overview of our team solution to the pneumonia detection challenge. Our team, DASA-FIDI-IARA, is composed of computer scientists, a computer engineer,  a radiologist informatician, and a neuroradiologist. We had much fun participating on this challenge, and we want to thank organizers and all people that worked hard to make it possible. Code and trained models are available on GitHub (link in the end).\n\n# Summary\n\nOur solution is based on [retinanet](https://github.com/fizyr/keras-retinanet) with resnet101 as backbone. We trained 5 folds of the training set from stage 1 and used the best single one. We experimented many ‘out-of-the-box’ ideas instead of implementing ensemble, making the stage 1 submission to be a single model. We have heavily used data augmentation, both on images and bounding boxes. After a careful analysis of width and height ratios of bounding boxes, we have tried many anchor ratios setups. Also, we played with inference hyperparameters, such as non-maximum suppression and score thresholds.\n\n# Training \nWe resized the original images to 800x800, as we got worse results for smaller sizes. We have modified the original retinanet implementation, including the following features:\n\n1.  Adding more anchor ratios, in a total of 9, ranging from 0.25 to 4;\n2.  Hyperparameter tuning: learning rate on training, non-maximum suppression and score threshold on inference;\n3.  Data augmentation\n   *  On images and bounding boxes: rotations, horizontal flips, translations, shearing and scaling\n   *  On images: contrast and brightness increase/decrease, gaussian noise\n   *  On bounding box: shifts on points (x1,y1) and (x2,y2); scaling up/down bounding boxes (this feature was not actually used in our model submission)\n4.  Dropouts both for regression and classifier heads;\n5.  We had information from the discussion in Kaggle that the annotation was performed in a different way for the training and test sets. So, we guessed how different it could be and trained our best model with the same training set, except by progressively reducing bounding boxes. The best model was trained with 15% reduction in height and width.\n\n# Ideas that did not work\nWe tried a lot of different things that did not work so well for us:\n\n1. Resnet52, resnet152, vgg, mobilenet, and even custom architectures as backbones;\n2. Adding more anchor boxes;\n3. Test Time Augmentation;\n4. YOLO and Light-Head RCNN;\n5. Freezing first layers of backbone;\n6. Histogram equalization as pre-processing;\n7. Two step approach: classifier + detector (pneumonia vs others)\n8. Adding gender and view position as features.\n9. Training specified networks for each gender (M/F) and each view position (AP/PA)\n10. Manually re-annotate the training set (not-normal cases) to remove evident false positives/negatives.\n\n# Final Thoughts\nAs mentioned, our Stage 1 submission was a single model, and we believe that ensemble could improve our score. For instance, after the final submission, we implemented an ensemble of 2 of our best models (just for curiosity), achieving 0.254 score in the Stage 1 testset. \n\nThe most important realization was finding that the annotation of the test set was different than the training set. Specifically, we found that targeting smaller bounding boxes yielded better results. The use of additional anchor-ratios also improved our mAP significantly. Our code for training and inference, as well as the trained models, are available on GitHub. There is also a demo where you can test our model (you need to request an account to try it). All the links can be found below.\n\n# Links\n* Documentation: [https://rsna-challenge-2018.readthedocs.io/en/latest/](https://rsna-challenge-2018.readthedocs.io/en/latest/)\n* Training/inference code: [https://github.com/alessonscap/rsna-challenge-2018.git](https://github.com/alessonscap/rsna-challenge-2018.git)\n* Demo page: [https://iarahealth.com/pneumonia/](https://iarahealth.com/pneumonia/)",
      "votes": null
    },
    {
      "id": "418135",
      "postDate": "11/09/2018 11:03:39",
      "content": "<p>Thank you, @Bernardo, @Alesson, @Daniel, @José and @Igor.</p>",
      "rawMarkdown": "Thank you, @Bernardo, @Alesson, @Daniel, @José and @Igor.",
      "votes": null
    },
    {
      "id": "418216",
      "postDate": "11/09/2018 13:45:53",
      "content": "<p>Thanks for the solution @Bernardo ! 0.254 on stage 1 LB with 2 models is excellent. Adding more anchor ratios is a nice adaptation to this problem.</p>",
      "rawMarkdown": "Thanks for the solution @Bernardo ! 0.254 on stage 1 LB with 2 models is excellent. Adding more anchor ratios is a nice adaptation to this problem.",
      "votes": null
    },
    {
      "id": "418245",
      "postDate": "11/09/2018 14:31:23",
      "content": "<p>Thank you too! Congrats on the great results, everyone :)</p>",
      "rawMarkdown": "Thank you too! Congrats on the great results, everyone :)",
      "votes": null
    },
    {
      "id": "419085",
      "postDate": "11/11/2018 08:44:52",
      "content": "<p>Great idea to add more anchor ratios! Impressive result for the single model, I received very significant improvement from the single averaging of 4 fold results alone.</p>",
      "rawMarkdown": "Great idea to add more anchor ratios! Impressive result for the single model, I received very significant improvement from the single averaging of 4 fold results alone.",
      "votes": null
    },
    {
      "id": "419873",
      "postDate": "11/12/2018 17:18:14",
      "content": "<p>Interesting to know that. Thanks for sharing!</p>",
      "rawMarkdown": "Interesting to know that. Thanks for sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 418135,
      "author_name": "felipekitamura",
      "author_url": "",
      "post_date": "11/09/2018 11:03:39",
      "content": "<p>Thank you, @Bernardo, @Alesson, @Daniel, @José and @Igor.</p>",
      "votes": null,
      "replies": [
        {
          "id": 418245,
          "author_name": "dsouza",
          "author_url": "",
          "post_date": "11/09/2018 14:31:23",
          "content": "<p>Thank you too! Congrats on the great results, everyone :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 418216,
      "author_name": "alexandrecc",
      "author_url": "",
      "post_date": "11/09/2018 13:45:53",
      "content": "<p>Thanks for the solution @Bernardo ! 0.254 on stage 1 LB with 2 models is excellent. Adding more anchor ratios is a nice adaptation to this problem.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 419085,
      "author_name": "dmytropoplavskiy",
      "author_url": "",
      "post_date": "11/11/2018 08:44:52",
      "content": "<p>Great idea to add more anchor ratios! Impressive result for the single model, I received very significant improvement from the single averaging of 4 fold results alone.</p>",
      "votes": null,
      "replies": [
        {
          "id": 419873,
          "author_name": "felipekitamura",
          "author_url": "",
          "post_date": "11/12/2018 17:18:14",
          "content": "<p>Interesting to know that. Thanks for sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "418134": "This is an overview of our team solution to the pneumonia detection challenge. Our team, DASA-FIDI-IARA, is composed of computer scientists, a computer engineer,  a radiologist informatician, and a neuroradiologist. We had much fun participating on this challenge, and we want to thank organizers and all people that worked hard to make it possible. Code and trained models are available on GitHub (link in the end).\n\n# Summary\n\nOur solution is based on [retinanet](https://github.com/fizyr/keras-retinanet) with resnet101 as backbone. We trained 5 folds of the training set from stage 1 and used the best single one. We experimented many ‘out-of-the-box’ ideas instead of implementing ensemble, making the stage 1 submission to be a single model. We have heavily used data augmentation, both on images and bounding boxes. After a careful analysis of width and height ratios of bounding boxes, we have tried many anchor ratios setups. Also, we played with inference hyperparameters, such as non-maximum suppression and score thresholds.\n\n# Training \nWe resized the original images to 800x800, as we got worse results for smaller sizes. We have modified the original retinanet implementation, including the following features:\n\n1.  Adding more anchor ratios, in a total of 9, ranging from 0.25 to 4;\n2.  Hyperparameter tuning: learning rate on training, non-maximum suppression and score threshold on inference;\n3.  Data augmentation\n   *  On images and bounding boxes: rotations, horizontal flips, translations, shearing and scaling\n   *  On images: contrast and brightness increase/decrease, gaussian noise\n   *  On bounding box: shifts on points (x1,y1) and (x2,y2); scaling up/down bounding boxes (this feature was not actually used in our model submission)\n4.  Dropouts both for regression and classifier heads;\n5.  We had information from the discussion in Kaggle that the annotation was performed in a different way for the training and test sets. So, we guessed how different it could be and trained our best model with the same training set, except by progressively reducing bounding boxes. The best model was trained with 15% reduction in height and width.\n\n# Ideas that did not work\nWe tried a lot of different things that did not work so well for us:\n\n1. Resnet52, resnet152, vgg, mobilenet, and even custom architectures as backbones;\n2. Adding more anchor boxes;\n3. Test Time Augmentation;\n4. YOLO and Light-Head RCNN;\n5. Freezing first layers of backbone;\n6. Histogram equalization as pre-processing;\n7. Two step approach: classifier + detector (pneumonia vs others)\n8. Adding gender and view position as features.\n9. Training specified networks for each gender (M/F) and each view position (AP/PA)\n10. Manually re-annotate the training set (not-normal cases) to remove evident false positives/negatives.\n\n# Final Thoughts\nAs mentioned, our Stage 1 submission was a single model, and we believe that ensemble could improve our score. For instance, after the final submission, we implemented an ensemble of 2 of our best models (just for curiosity), achieving 0.254 score in the Stage 1 testset. \n\nThe most important realization was finding that the annotation of the test set was different than the training set. Specifically, we found that targeting smaller bounding boxes yielded better results. The use of additional anchor-ratios also improved our mAP significantly. Our code for training and inference, as well as the trained models, are available on GitHub. There is also a demo where you can test our model (you need to request an account to try it). All the links can be found below.\n\n# Links\n* Documentation: [https://rsna-challenge-2018.readthedocs.io/en/latest/](https://rsna-challenge-2018.readthedocs.io/en/latest/)\n* Training/inference code: [https://github.com/alessonscap/rsna-challenge-2018.git](https://github.com/alessonscap/rsna-challenge-2018.git)\n* Demo page: [https://iarahealth.com/pneumonia/](https://iarahealth.com/pneumonia/)",
    "418135": "Thank you, @Bernardo, @Alesson, @Daniel, @José and @Igor.",
    "418216": "Thanks for the solution @Bernardo ! 0.254 on stage 1 LB with 2 models is excellent. Adding more anchor ratios is a nice adaptation to this problem.",
    "418245": "Thank you too! Congrats on the great results, everyone :)",
    "419085": "Great idea to add more anchor ratios! Impressive result for the single model, I received very significant improvement from the single averaging of 4 fold results alone.",
    "419873": "Interesting to know that. Thanks for sharing!"
  },
  "source": "meta"
}