{
  "id": 228775,
  "title": "3rd Place Solution [Preferred CLiP]",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/writeups/preferred-clip-3rd-place-solution-preferred-clip",
  "author_name": "",
  "post_date": "2021-03-26T11:42:18.190568300Z",
  "votes": 43,
  "comment_count": 3,
  "views": 0,
  "content": "<p><strong>Congratulations to all the winners and thanks to the organizers for making this competition possible!</strong></p>\n<p>Here is the solution of our team(Preferred CLiP: <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a>, <a href=\"https://www.kaggle.com/la4laaa\" target=\"_blank\">@la4laaa</a>, <a href=\"https://www.kaggle.com/yhirano\" target=\"_blank\">@yhirano</a>, <a href=\"https://www.kaggle.com/suga93\" target=\"_blank\">@suga93</a> )</p>\n<h1>Short Summary (TL; DR)</h1>\n<p>Our strategy is based on the 3-stage training proposed by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> and initially implemented by <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> in a public notebook. We also trained a segmenter to give pseudo annotations to unannotated or external data in stage 1 and 2. We did not use the segmenter in the final inference.</p>\n<p>Our training pipeline in a nutshell:</p>\n<ul>\n<li>Train a segmenter that predicts by which type of catheter (or none) each pixel of the input image is occupied.</li>\n<li>Make predictions on unannotated RANZCR data and external data (NIH Chest X-rays and MIMIC-CXR) with the segmenter. We excluded samples without any catheters from the external data.</li>\n<li>Perform multi-stage classification.<ul>\n<li>Stage 1: Superimpose ground-truth annotation (if exists) or the output of the segmenter (otherwise) on the original images and train a <strong>teacher</strong> model with them.</li>\n<li>Stage 2: Train a <strong>student</strong> model using both classification loss and consistency loss calculated by comparing its features with that of the teacher model.</li>\n<li>Stage 3: Fine-tune the student model.</li></ul></li>\n</ul>\n<p><img src=\"https://media.discordapp.net/attachments/824969420731318335/824969475153199104/ranzcr-clip-overview.png?width=1674&amp;height=936\" alt=\"https://media.discordapp.net/attachments/824969420731318335/824969475153199104/ranzcr-clip-overview.png?width=1674&amp;height=936\"></p>\n<p><strong>[CV / Public LB / Private LB]</strong><br>\nVanilla multi-stage training, TTA (resnet200d): 0.96606/0.97020/0.97328<br>\n+Segmenter: 0.96742/0.97323/0.97389<br>\n(+Ensemble of 4 different architectures: 0.97101/0.97430/0.97515)<br>\n+Pseudo labels(NIH): 0.96918/0.97335/0.97455<br>\n+Ensemble of 8 different training setups, wo TTA: 0.97215/0.97386/0.97611<br>\n+Determining ensemble coefficients by logistic regression: <strong>0.97228/0.97434/0.97624</strong></p>\n<h2>Data augmentation</h2>\n<p>We used the data augmentation proposed by sin in the following notebook for both segmentation and classification. We used <strong>720</strong> px for image_size.<br>\n<a href=\"https://www.kaggle.com/underwearfitting/resnet200d-public-benchmark-2xtta-lb0-965\" target=\"_blank\">https://www.kaggle.com/underwearfitting/resnet200d-public-benchmark-2xtta-lb0-965</a></p>\n<h2>Segmentation</h2>\n<p><strong>Problem setting</strong></p>\n<ul>\n<li>Pixel-wise multi-label classification</li>\n<li>4 classes (ETT, NGT, CVC, Swan ganz)</li>\n</ul>\n<p>Architecture: DeepLabV3+ <br>\nEncoder backbones:</p>\n<ul>\n<li>Resnet152</li>\n<li>Regnety160</li>\n</ul>\n<p>Image size: 720x720<br>\nLoss: BCEWithLogitsLoss + DiceLoss + RecallLoss</p>\n<ul>\n<li>Since we found that recall is important for this task, we adopted recall loss which is introduced in <a href=\"https://openreview.net/forum?id=SlprFTIQP3\" target=\"_blank\">Recall Loss for Imbalanced Image Classification and Semantic Segmentation</a>.</li>\n</ul>\n<p><strong>Training steps</strong></p>\n<ol>\n<li>Trained the first segmenter only with RANZCR annotated data. </li>\n<li>Fine-tuned the segmenter using the NIH Chest X-rays as an external dataset in addition to RANZCR annotated data. Images of the NIH Chest X-rays were selected if any of the predicted values from the previous best classifier were over 0.95. Pseudo labels of them were generated from the previous best ensemble segmenter.</li>\n</ol>\n<p>Segmentation masks used in the classification task are generated by the ensemble segmenter, which is an ensemble of 10 pretrained segmenters (5 fold x 2 backbone). Although we found that there are some noises in annotations, they are still more informative than our segmenter outputs. So we only used the predicted segmentation masks for unannotated data and external data, and used the ground-truth annotations for annotated data.</p>\n<h2>Classification</h2>\n<p>For classification, we adopted almost the same strategy proposed in the following notebooks:</p>\n<ul>\n<li>3 stage training (by <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>)<ul>\n<li><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577</a> </li></ul></li>\n<li>use annotation without segmentation (by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>)<ul>\n<li><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243</a></li></ul></li>\n</ul>\n<p>Training procedure:</p>\n<ul>\n<li>batch size: 64</li>\n<li>optimizer: RAdam</li>\n<li>scheduler: CosineAnnealingWarmRestarts (stage1, 2) or ExponentialLR (stage 3)<ul>\n<li>lr_decay_rate=0.8 for stage3</li></ul></li>\n<li>epoch: 15 (stage1, 2) or 10 (stage 3)</li>\n</ul>\n<h2>Pseudo-labeling</h2>\n<p>To leverage external data, we gave pseudo labels to both MIMIC-CXR and NIH Chest X-rays. We first trained preliminary models only with the official RANZCR dataset. Then, we filtered out images without any catheters using their prediction values, and iteratively trained both segmenters and classifiers by using the last models' outputs as additional pseudo labels. As for the NIH dataset, we also omitted duplicated images. To avoid leakage, we used different pseudo labels for each fold.</p>\n<h2>Ensemble</h2>\n<p>For classification, we finally trained 4 architectures: resnet200d, seresnet152d, resnest50d, and efficientnet-b5, with either pseudo-labeled MIMIC-CXR or NIH Chest X-rays as an additional dataset. We performed 5-fold CV for each training setup, yielding 40 models (=4x2x5) in total. Our final submission was an ensemble of these 40 models. We determined the ensemble coefficients of length 8 (corresponding to each training setup) by performing logistic regression on the local validation set.</p>\n<p>The execution time in Kaggle notebook was about 7~8 hours. </p>\n<p>We used 8 NVIDIA V100 GPUs with 32GB memory for all the training. A training of 15 epochs took roughly 4 hours. Whole 5 fold multi-stage training with resnet200d took roughly 15 GPU days.</p>",
  "messages": [
    {
      "id": "1253126",
      "postDate": "03/26/2021 11:42:18",
      "content": "<p><strong>Congratulations to all the winners and thanks to the organizers for making this competition possible!</strong></p>\n<p>Here is the solution of our team(Preferred CLiP: <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a>, <a href=\"https://www.kaggle.com/la4laaa\" target=\"_blank\">@la4laaa</a>, <a href=\"https://www.kaggle.com/yhirano\" target=\"_blank\">@yhirano</a>, <a href=\"https://www.kaggle.com/suga93\" target=\"_blank\">@suga93</a> )</p>\n<h1>Short Summary (TL; DR)</h1>\n<p>Our strategy is based on the 3-stage training proposed by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> and initially implemented by <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> in a public notebook. We also trained a segmenter to give pseudo annotations to unannotated or external data in stage 1 and 2. We did not use the segmenter in the final inference.</p>\n<p>Our training pipeline in a nutshell:</p>\n<ul>\n<li>Train a segmenter that predicts by which type of catheter (or none) each pixel of the input image is occupied.</li>\n<li>Make predictions on unannotated RANZCR data and external data (NIH Chest X-rays and MIMIC-CXR) with the segmenter. We excluded samples without any catheters from the external data.</li>\n<li>Perform multi-stage classification.<ul>\n<li>Stage 1: Superimpose ground-truth annotation (if exists) or the output of the segmenter (otherwise) on the original images and train a <strong>teacher</strong> model with them.</li>\n<li>Stage 2: Train a <strong>student</strong> model using both classification loss and consistency loss calculated by comparing its features with that of the teacher model.</li>\n<li>Stage 3: Fine-tune the student model.</li></ul></li>\n</ul>\n<p><img src=\"https://media.discordapp.net/attachments/824969420731318335/824969475153199104/ranzcr-clip-overview.png?width=1674&amp;height=936\" alt=\"https://media.discordapp.net/attachments/824969420731318335/824969475153199104/ranzcr-clip-overview.png?width=1674&amp;height=936\"></p>\n<p><strong>[CV / Public LB / Private LB]</strong><br>\nVanilla multi-stage training, TTA (resnet200d): 0.96606/0.97020/0.97328<br>\n+Segmenter: 0.96742/0.97323/0.97389<br>\n(+Ensemble of 4 different architectures: 0.97101/0.97430/0.97515)<br>\n+Pseudo labels(NIH): 0.96918/0.97335/0.97455<br>\n+Ensemble of 8 different training setups, wo TTA: 0.97215/0.97386/0.97611<br>\n+Determining ensemble coefficients by logistic regression: <strong>0.97228/0.97434/0.97624</strong></p>\n<h2>Data augmentation</h2>\n<p>We used the data augmentation proposed by sin in the following notebook for both segmentation and classification. We used <strong>720</strong> px for image_size.<br>\n<a href=\"https://www.kaggle.com/underwearfitting/resnet200d-public-benchmark-2xtta-lb0-965\" target=\"_blank\">https://www.kaggle.com/underwearfitting/resnet200d-public-benchmark-2xtta-lb0-965</a></p>\n<h2>Segmentation</h2>\n<p><strong>Problem setting</strong></p>\n<ul>\n<li>Pixel-wise multi-label classification</li>\n<li>4 classes (ETT, NGT, CVC, Swan ganz)</li>\n</ul>\n<p>Architecture: DeepLabV3+ <br>\nEncoder backbones:</p>\n<ul>\n<li>Resnet152</li>\n<li>Regnety160</li>\n</ul>\n<p>Image size: 720x720<br>\nLoss: BCEWithLogitsLoss + DiceLoss + RecallLoss</p>\n<ul>\n<li>Since we found that recall is important for this task, we adopted recall loss which is introduced in <a href=\"https://openreview.net/forum?id=SlprFTIQP3\" target=\"_blank\">Recall Loss for Imbalanced Image Classification and Semantic Segmentation</a>.</li>\n</ul>\n<p><strong>Training steps</strong></p>\n<ol>\n<li>Trained the first segmenter only with RANZCR annotated data. </li>\n<li>Fine-tuned the segmenter using the NIH Chest X-rays as an external dataset in addition to RANZCR annotated data. Images of the NIH Chest X-rays were selected if any of the predicted values from the previous best classifier were over 0.95. Pseudo labels of them were generated from the previous best ensemble segmenter.</li>\n</ol>\n<p>Segmentation masks used in the classification task are generated by the ensemble segmenter, which is an ensemble of 10 pretrained segmenters (5 fold x 2 backbone). Although we found that there are some noises in annotations, they are still more informative than our segmenter outputs. So we only used the predicted segmentation masks for unannotated data and external data, and used the ground-truth annotations for annotated data.</p>\n<h2>Classification</h2>\n<p>For classification, we adopted almost the same strategy proposed in the following notebooks:</p>\n<ul>\n<li>3 stage training (by <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>)<ul>\n<li><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577</a> </li></ul></li>\n<li>use annotation without segmentation (by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>)<ul>\n<li><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243</a></li></ul></li>\n</ul>\n<p>Training procedure:</p>\n<ul>\n<li>batch size: 64</li>\n<li>optimizer: RAdam</li>\n<li>scheduler: CosineAnnealingWarmRestarts (stage1, 2) or ExponentialLR (stage 3)<ul>\n<li>lr_decay_rate=0.8 for stage3</li></ul></li>\n<li>epoch: 15 (stage1, 2) or 10 (stage 3)</li>\n</ul>\n<h2>Pseudo-labeling</h2>\n<p>To leverage external data, we gave pseudo labels to both MIMIC-CXR and NIH Chest X-rays. We first trained preliminary models only with the official RANZCR dataset. Then, we filtered out images without any catheters using their prediction values, and iteratively trained both segmenters and classifiers by using the last models' outputs as additional pseudo labels. As for the NIH dataset, we also omitted duplicated images. To avoid leakage, we used different pseudo labels for each fold.</p>\n<h2>Ensemble</h2>\n<p>For classification, we finally trained 4 architectures: resnet200d, seresnet152d, resnest50d, and efficientnet-b5, with either pseudo-labeled MIMIC-CXR or NIH Chest X-rays as an additional dataset. We performed 5-fold CV for each training setup, yielding 40 models (=4x2x5) in total. Our final submission was an ensemble of these 40 models. We determined the ensemble coefficients of length 8 (corresponding to each training setup) by performing logistic regression on the local validation set.</p>\n<p>The execution time in Kaggle notebook was about 7~8 hours. </p>\n<p>We used 8 NVIDIA V100 GPUs with 32GB memory for all the training. A training of 15 epochs took roughly 4 hours. Whole 5 fold multi-stage training with resnet200d took roughly 15 GPU days.</p>",
      "rawMarkdown": "**Congratulations to all the winners and thanks to the organizers for making this competition possible!**\n\nHere is the solution of our team(Preferred CLiP: @charmq, @la4laaa, @yhirano, @suga93 )\n\n\n# Short Summary (TL; DR)\nOur strategy is based on the 3-stage training proposed by @hengck23 and initially implemented by @yasufuminakama in a public notebook. We also trained a segmenter to give pseudo annotations to unannotated or external data in stage 1 and 2. We did not use the segmenter in the final inference.\n\nOur training pipeline in a nutshell:\n- Train a segmenter that predicts by which type of catheter (or none) each pixel of the input image is occupied.\n- Make predictions on unannotated RANZCR data and external data (NIH Chest X-rays and MIMIC-CXR) with the segmenter. We excluded samples without any catheters from the external data.\n- Perform multi-stage classification.\n   - Stage 1: Superimpose ground-truth annotation (if exists) or the output of the segmenter (otherwise) on the original images and train a **teacher** model with them.\n   - Stage 2: Train a **student** model using both classification loss and consistency loss calculated by comparing its features with that of the teacher model.\n   - Stage 3: Fine-tune the student model.\n\n![https://media.discordapp.net/attachments/824969420731318335/824969475153199104/ranzcr-clip-overview.png?width=1674&height=936](https://media.discordapp.net/attachments/824969420731318335/824969475153199104/ranzcr-clip-overview.png?width=1674&height=936)\n\n**[CV / Public LB / Private LB]**\nVanilla multi-stage training, TTA (resnet200d): 0.96606/0.97020/0.97328\n+Segmenter: 0.96742/0.97323/0.97389\n(+Ensemble of 4 different architectures: 0.97101/0.97430/0.97515)\n+Pseudo labels(NIH): 0.96918/0.97335/0.97455\n+Ensemble of 8 different training setups, wo TTA: 0.97215/0.97386/0.97611\n+Determining ensemble coefficients by logistic regression: **0.97228/0.97434/0.97624**\n\n\n## Data augmentation\nWe used the data augmentation proposed by sin in the following notebook for both segmentation and classification. We used **720** px for image_size.\nhttps://www.kaggle.com/underwearfitting/resnet200d-public-benchmark-2xtta-lb0-965\n\n## Segmentation\n**Problem setting**\n- Pixel-wise multi-label classification\n- 4 classes (ETT, NGT, CVC, Swan ganz)\n\nArchitecture: DeepLabV3+ \nEncoder backbones:\n- Resnet152\n- Regnety160\n\nImage size: 720x720\nLoss: BCEWithLogitsLoss + DiceLoss + RecallLoss\n- Since we found that recall is important for this task, we adopted recall loss which is introduced in [Recall Loss for Imbalanced Image Classification and Semantic Segmentation](https://openreview.net/forum?id=SlprFTIQP3).\n\n**Training steps**\n1. Trained the first segmenter only with RANZCR annotated data. \n2. Fine-tuned the segmenter using the NIH Chest X-rays as an external dataset in addition to RANZCR annotated data. Images of the NIH Chest X-rays were selected if any of the predicted values from the previous best classifier were over 0.95. Pseudo labels of them were generated from the previous best ensemble segmenter.\n\nSegmentation masks used in the classification task are generated by the ensemble segmenter, which is an ensemble of 10 pretrained segmenters (5 fold x 2 backbone). Although we found that there are some noises in annotations, they are still more informative than our segmenter outputs. So we only used the predicted segmentation masks for unannotated data and external data, and used the ground-truth annotations for annotated data.\n\n## Classification\nFor classification, we adopted almost the same strategy proposed in the following notebooks:\n- 3 stage training (by @yasufuminakama)\n   - https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577 \n- use annotation without segmentation (by @hengck23)\n   - https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243\n\nTraining procedure:\n- batch size: 64\n- optimizer: RAdam\n- scheduler: CosineAnnealingWarmRestarts (stage1, 2) or ExponentialLR (stage 3)\n   - lr_decay_rate=0.8 for stage3\n- epoch: 15 (stage1, 2) or 10 (stage 3)\n\n## Pseudo-labeling\nTo leverage external data, we gave pseudo labels to both MIMIC-CXR and NIH Chest X-rays. We first trained preliminary models only with the official RANZCR dataset. Then, we filtered out images without any catheters using their prediction values, and iteratively trained both segmenters and classifiers by using the last models' outputs as additional pseudo labels. As for the NIH dataset, we also omitted duplicated images. To avoid leakage, we used different pseudo labels for each fold.\n\n## Ensemble\nFor classification, we finally trained 4 architectures: resnet200d, seresnet152d, resnest50d, and efficientnet-b5, with either pseudo-labeled MIMIC-CXR or NIH Chest X-rays as an additional dataset. We performed 5-fold CV for each training setup, yielding 40 models (=4x2x5) in total. Our final submission was an ensemble of these 40 models. We determined the ensemble coefficients of length 8 (corresponding to each training setup) by performing logistic regression on the local validation set.\n\nThe execution time in Kaggle notebook was about 7~8 hours. \n\nWe used 8 NVIDIA V100 GPUs with 32GB memory for all the training. A training of 15 epochs took roughly 4 hours. Whole 5 fold multi-stage training with resnet200d took roughly 15 GPU days.",
      "votes": null
    },
    {
      "id": "1253173",
      "postDate": "03/26/2021 12:38:40",
      "content": "<p>Congrats on 3rd place <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> and team. Thanks for sharing your team solution</p>",
      "rawMarkdown": "Congrats on 3rd place @charmq and team. Thanks for sharing your team solution",
      "votes": null
    },
    {
      "id": "1253797",
      "postDate": "03/27/2021 03:34:53",
      "content": "<p>You achieved that high score with only <strong>720 px</strong> images ! Amazing :) </p>",
      "rawMarkdown": "You achieved that high score with only **720 px** images ! Amazing :)",
      "votes": null
    },
    {
      "id": "1365178",
      "postDate": "06/25/2021 13:45:56",
      "content": "<blockquote>\n  <p>We determined the ensemble coefficients of length 8 (corresponding to each training setup) by performing logistic regression on the local validation set.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> if it's okay, would you like to elaborate more on this?</p>",
      "rawMarkdown": "> We determined the ensemble coefficients of length 8 (corresponding to each training setup) by performing logistic regression on the local validation set.\n\n@charmq if it's okay, would you like to elaborate more on this?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1253173,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "03/26/2021 12:38:40",
      "content": "<p>Congrats on 3rd place <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> and team. Thanks for sharing your team solution</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1253797,
      "author_name": "analokamus",
      "author_url": "",
      "post_date": "03/27/2021 03:34:53",
      "content": "<p>You achieved that high score with only <strong>720 px</strong> images ! Amazing :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1365178,
      "author_name": "rhtsingh",
      "author_url": "",
      "post_date": "06/25/2021 13:45:56",
      "content": "<blockquote>\n  <p>We determined the ensemble coefficients of length 8 (corresponding to each training setup) by performing logistic regression on the local validation set.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> if it's okay, would you like to elaborate more on this?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1253126": "**Congratulations to all the winners and thanks to the organizers for making this competition possible!**\n\nHere is the solution of our team(Preferred CLiP: @charmq, @la4laaa, @yhirano, @suga93 )\n\n\n# Short Summary (TL; DR)\nOur strategy is based on the 3-stage training proposed by @hengck23 and initially implemented by @yasufuminakama in a public notebook. We also trained a segmenter to give pseudo annotations to unannotated or external data in stage 1 and 2. We did not use the segmenter in the final inference.\n\nOur training pipeline in a nutshell:\n- Train a segmenter that predicts by which type of catheter (or none) each pixel of the input image is occupied.\n- Make predictions on unannotated RANZCR data and external data (NIH Chest X-rays and MIMIC-CXR) with the segmenter. We excluded samples without any catheters from the external data.\n- Perform multi-stage classification.\n   - Stage 1: Superimpose ground-truth annotation (if exists) or the output of the segmenter (otherwise) on the original images and train a **teacher** model with them.\n   - Stage 2: Train a **student** model using both classification loss and consistency loss calculated by comparing its features with that of the teacher model.\n   - Stage 3: Fine-tune the student model.\n\n![https://media.discordapp.net/attachments/824969420731318335/824969475153199104/ranzcr-clip-overview.png?width=1674&height=936](https://media.discordapp.net/attachments/824969420731318335/824969475153199104/ranzcr-clip-overview.png?width=1674&height=936)\n\n**[CV / Public LB / Private LB]**\nVanilla multi-stage training, TTA (resnet200d): 0.96606/0.97020/0.97328\n+Segmenter: 0.96742/0.97323/0.97389\n(+Ensemble of 4 different architectures: 0.97101/0.97430/0.97515)\n+Pseudo labels(NIH): 0.96918/0.97335/0.97455\n+Ensemble of 8 different training setups, wo TTA: 0.97215/0.97386/0.97611\n+Determining ensemble coefficients by logistic regression: **0.97228/0.97434/0.97624**\n\n\n## Data augmentation\nWe used the data augmentation proposed by sin in the following notebook for both segmentation and classification. We used **720** px for image_size.\nhttps://www.kaggle.com/underwearfitting/resnet200d-public-benchmark-2xtta-lb0-965\n\n## Segmentation\n**Problem setting**\n- Pixel-wise multi-label classification\n- 4 classes (ETT, NGT, CVC, Swan ganz)\n\nArchitecture: DeepLabV3+ \nEncoder backbones:\n- Resnet152\n- Regnety160\n\nImage size: 720x720\nLoss: BCEWithLogitsLoss + DiceLoss + RecallLoss\n- Since we found that recall is important for this task, we adopted recall loss which is introduced in [Recall Loss for Imbalanced Image Classification and Semantic Segmentation](https://openreview.net/forum?id=SlprFTIQP3).\n\n**Training steps**\n1. Trained the first segmenter only with RANZCR annotated data. \n2. Fine-tuned the segmenter using the NIH Chest X-rays as an external dataset in addition to RANZCR annotated data. Images of the NIH Chest X-rays were selected if any of the predicted values from the previous best classifier were over 0.95. Pseudo labels of them were generated from the previous best ensemble segmenter.\n\nSegmentation masks used in the classification task are generated by the ensemble segmenter, which is an ensemble of 10 pretrained segmenters (5 fold x 2 backbone). Although we found that there are some noises in annotations, they are still more informative than our segmenter outputs. So we only used the predicted segmentation masks for unannotated data and external data, and used the ground-truth annotations for annotated data.\n\n## Classification\nFor classification, we adopted almost the same strategy proposed in the following notebooks:\n- 3 stage training (by @yasufuminakama)\n   - https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577 \n- use annotation without segmentation (by @hengck23)\n   - https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243\n\nTraining procedure:\n- batch size: 64\n- optimizer: RAdam\n- scheduler: CosineAnnealingWarmRestarts (stage1, 2) or ExponentialLR (stage 3)\n   - lr_decay_rate=0.8 for stage3\n- epoch: 15 (stage1, 2) or 10 (stage 3)\n\n## Pseudo-labeling\nTo leverage external data, we gave pseudo labels to both MIMIC-CXR and NIH Chest X-rays. We first trained preliminary models only with the official RANZCR dataset. Then, we filtered out images without any catheters using their prediction values, and iteratively trained both segmenters and classifiers by using the last models' outputs as additional pseudo labels. As for the NIH dataset, we also omitted duplicated images. To avoid leakage, we used different pseudo labels for each fold.\n\n## Ensemble\nFor classification, we finally trained 4 architectures: resnet200d, seresnet152d, resnest50d, and efficientnet-b5, with either pseudo-labeled MIMIC-CXR or NIH Chest X-rays as an additional dataset. We performed 5-fold CV for each training setup, yielding 40 models (=4x2x5) in total. Our final submission was an ensemble of these 40 models. We determined the ensemble coefficients of length 8 (corresponding to each training setup) by performing logistic regression on the local validation set.\n\nThe execution time in Kaggle notebook was about 7~8 hours. \n\nWe used 8 NVIDIA V100 GPUs with 32GB memory for all the training. A training of 15 epochs took roughly 4 hours. Whole 5 fold multi-stage training with resnet200d took roughly 15 GPU days.",
    "1253173": "Congrats on 3rd place @charmq and team. Thanks for sharing your team solution",
    "1253797": "You achieved that high score with only **720 px** images ! Amazing :)",
    "1365178": "> We determined the ensemble coefficients of length 8 (corresponding to each training setup) by performing logistic regression on the local validation set.\n\n@charmq if it's okay, would you like to elaborate more on this?"
  },
  "source": "meta"
}