{
  "id": 226633,
  "title": "1st Place Solution",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/writeups/all-data-are-ext-1st-place-solution",
  "author_name": "",
  "post_date": "2021-03-19T13:36:36.713Z",
  "votes": 159,
  "comment_count": 45,
  "views": 0,
  "content": "<p>Congrats to all the winners. Thank you for great collaboration, my long time teammates <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> and <a href=\"https://www.kaggle.com/garybios\" target=\"_blank\">@garybios</a>. I had a blast once again working with such talented teammates.</p>\n<h2>TL;DR</h2>\n<p>2-stage segmentation and 2-stage classification pipeline. Using pseudo labels in both segmentation and classification.</p>\n<h2>Dataset definitions</h2>\n<p>Different subsets of the NIH ChestX dataset (112k images) are illustrated in this Venn diagram. </p>\n<ul>\n<li>Official dataset (30k) contains 9k images with tube segmentation ground truth. </li>\n<li>We also used Dr. Konya’s <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> trachea bifurcation annotation <a href=\"https://www.kaggle.com/sandorkonya/5k-trachea-bifurcation-on-chest-xray\" target=\"_blank\">dataset</a> which has 5k images. Thank you Doctor!</li>\n<li>For pseudo labeling, we identified 28k external images outside the Train set which contain tubes: (1) first run a segmentation model inference on all the 112k images to filter out all the images without tubes; (2) using imagehash to de-duplicate the 30k images that are already in 30k train set; (3) link patient IDs to make sure same patient in external data and original data falls into the same fold.<br>\n<img src=\"https://i.imgur.com/KJQumVO.png\" alt=\"\"></li>\n</ul>\n<h2>Pre-processing</h2>\n<p>For images, we applied <a href=\"https://www.kaggle.com/ratthachat/aptos-eye-preprocessing-in-diabetic-retinopathy#2.-Try-Ben-Graham's-preprocessing-method.\" target=\"_blank\">“Ben’s pre-processing”</a> with different parameters, in order to train diverse models.</p>\n<p>For official segmentation annotation, we make 2 channel masks by drawing lines representing the tube, and drawing big dots indicating the tips of tubes. See pictures below.</p>\n<p>For trachea bifurcation annotation, we make 1 channel masks by drawing big dots.</p>\n<h2>Segmentation Stage 1</h2>\n<p><img src=\"https://i.imgur.com/xdP22R1.png\" alt=\"\"></p>\n<ul>\n<li>Model 1: Mask is tube and tips - 2 channel output<ul>\n<li>Train and validate on 9k images with tube anno</li>\n<li>Pseudo label on 28k + (30k - 9k) data without tube anno</li>\n<li>10 model ensemble with a mixture of Unet and Unet++, with backbones B3-B8, and different preprocessing parameters, at image size 1024x1024 to 1536x1536.</li></ul></li>\n<li>Model 2: Trachea bifurcation (TB) - 1 channel output<ul>\n<li>Train and validate on 5k images with TB anno</li>\n<li>Pseudo label on 28k + (30k - 5k) data without TB anno</li>\n<li>Similar ensemble as model 1. But TBs are easier to segment, so the image sizes are 384x384 to 1024x1024</li></ul></li>\n</ul>\n<h2>Segmentation Stage 2</h2>\n<p><img src=\"https://i.imgur.com/r4fI8Iz.png\" alt=\"\"></p>\n<ul>\n<li>Tubes, tips and TB – 3 channel output<ul>\n<li>Train on 30k + 28k images (with combination of GT and pseudo labels)</li>\n<li>Validate tubes and tips on 9k data</li>\n<li>Validate TB separately on 5k data</li>\n<li>Predict out-of-fold on all 30k + 28k images, to be used by classification</li></ul></li>\n<li>Stage 2 needs to run in inference kernel, so there are only 5 models in the ensemble:<ul>\n<li>Unet++ B3 at 1536</li>\n<li>Unet B4 at 1536</li>\n<li>Unet++ B5 at 1024</li>\n<li>Unet++ B6 at 1024</li>\n<li>Unet B7 at 1024</li></ul></li>\n</ul>\n<p>Locally we trained 5 fold for each model, in order to get an OOF cv score. In inference, only one fold from each model is used.</p>\n<h2>Classification Stage 1</h2>\n<ul>\n<li>Input is 6 channel (3 ch original image + 3 ch predicted masks)</li>\n<li>Output is 12 classes: original 11 plus no_ETT, defined as whether all 3 ETT classes are 0</li>\n<li>Loss is weighted average of CE loss for the 4 ETT classes and BCE loss for the other 8 classes, with weight being 1:7</li>\n<li>Train on 30k data; make pseudo labels on 28k external data</li>\n<li>20 model ensemble, a mixture of EfficientNets, ResNets, ResNexts, ViTs at size 384 to 512, with various pre-processing parameters</li>\n<li>CV = <strong>0.97553</strong> with rank ensemble</li>\n</ul>\n<h2>Classification Stage 2</h2>\n<ul>\n<li>Same input, output, loss as Stage 1</li>\n<li>Training with 30k+28k data (combination of GT and pseudo labels)</li>\n<li>Since Stage 2 models need to go into inference kernel, overall model sizes are smaller than Stage 1</li>\n<li>31 model ensemble, a mix of EfficientNets, ResNets, SEResNexts, ResNexts, RegNet, Inception, RexNet, DenseNet, ViTs etc at size 384 to 512, with various pre-processing parameters</li>\n<li>CV = <strong>0.97606</strong> with rank ensemble</li>\n</ul>\n<p>Trained 5 fold locally, but only squeezed 67 folds into the inference kernel. </p>\n<h2>Post-processing</h2>\n<p>For AUC metric, sometimes it’s better to rank the probabilities of each model before ensembling. For this competition, we found that this is true only for these 5 columns:</p>\n<pre><code>            'ETT - Abnormal',\n            'NGT - Borderline',\n            'NGT - Incompletely Imaged',\n            'CVC - Normal',\n            'Swan Ganz Catheter Present'\n</code></pre>\n<p>This boosts CV score by about 0.00032</p>\n<h2>update</h2>\n<p>We have released simplified training and inference code:<br>\n<a href=\"https://www.kaggle.com/haqishen/ranzcr-1st-place-soluiton-seg-model-small-ver\" target=\"_blank\">Segmentation training</a>,  <a href=\"https://www.kaggle.com/haqishen/ranzcr-1st-place-soluiton-cls-model-small-ver\" target=\"_blank\">Classification training</a>,  <a href=\"https://www.kaggle.com/haqishen/ranzcr-1st-place-soluiton-inference-small-ver\" target=\"_blank\">Inference</a></p>",
  "messages": [
    {
      "id": "1241669",
      "postDate": "03/17/2021 06:33:17",
      "content": "<p>Congrats to all the winners. Thank you for great collaboration, my long time teammates <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> and <a href=\"https://www.kaggle.com/garybios\" target=\"_blank\">@garybios</a>. I had a blast once again working with such talented teammates.</p>\n<h2>TL;DR</h2>\n<p>2-stage segmentation and 2-stage classification pipeline. Using pseudo labels in both segmentation and classification.</p>\n<h2>Dataset definitions</h2>\n<p>Different subsets of the NIH ChestX dataset (112k images) are illustrated in this Venn diagram. </p>\n<ul>\n<li>Official dataset (30k) contains 9k images with tube segmentation ground truth. </li>\n<li>We also used Dr. Konya’s <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> trachea bifurcation annotation <a href=\"https://www.kaggle.com/sandorkonya/5k-trachea-bifurcation-on-chest-xray\" target=\"_blank\">dataset</a> which has 5k images. Thank you Doctor!</li>\n<li>For pseudo labeling, we identified 28k external images outside the Train set which contain tubes: (1) first run a segmentation model inference on all the 112k images to filter out all the images without tubes; (2) using imagehash to de-duplicate the 30k images that are already in 30k train set; (3) link patient IDs to make sure same patient in external data and original data falls into the same fold.<br>\n<img src=\"https://i.imgur.com/KJQumVO.png\" alt=\"\"></li>\n</ul>\n<h2>Pre-processing</h2>\n<p>For images, we applied <a href=\"https://www.kaggle.com/ratthachat/aptos-eye-preprocessing-in-diabetic-retinopathy#2.-Try-Ben-Graham's-preprocessing-method.\" target=\"_blank\">“Ben’s pre-processing”</a> with different parameters, in order to train diverse models.</p>\n<p>For official segmentation annotation, we make 2 channel masks by drawing lines representing the tube, and drawing big dots indicating the tips of tubes. See pictures below.</p>\n<p>For trachea bifurcation annotation, we make 1 channel masks by drawing big dots.</p>\n<h2>Segmentation Stage 1</h2>\n<p><img src=\"https://i.imgur.com/xdP22R1.png\" alt=\"\"></p>\n<ul>\n<li>Model 1: Mask is tube and tips - 2 channel output<ul>\n<li>Train and validate on 9k images with tube anno</li>\n<li>Pseudo label on 28k + (30k - 9k) data without tube anno</li>\n<li>10 model ensemble with a mixture of Unet and Unet++, with backbones B3-B8, and different preprocessing parameters, at image size 1024x1024 to 1536x1536.</li></ul></li>\n<li>Model 2: Trachea bifurcation (TB) - 1 channel output<ul>\n<li>Train and validate on 5k images with TB anno</li>\n<li>Pseudo label on 28k + (30k - 5k) data without TB anno</li>\n<li>Similar ensemble as model 1. But TBs are easier to segment, so the image sizes are 384x384 to 1024x1024</li></ul></li>\n</ul>\n<h2>Segmentation Stage 2</h2>\n<p><img src=\"https://i.imgur.com/r4fI8Iz.png\" alt=\"\"></p>\n<ul>\n<li>Tubes, tips and TB – 3 channel output<ul>\n<li>Train on 30k + 28k images (with combination of GT and pseudo labels)</li>\n<li>Validate tubes and tips on 9k data</li>\n<li>Validate TB separately on 5k data</li>\n<li>Predict out-of-fold on all 30k + 28k images, to be used by classification</li></ul></li>\n<li>Stage 2 needs to run in inference kernel, so there are only 5 models in the ensemble:<ul>\n<li>Unet++ B3 at 1536</li>\n<li>Unet B4 at 1536</li>\n<li>Unet++ B5 at 1024</li>\n<li>Unet++ B6 at 1024</li>\n<li>Unet B7 at 1024</li></ul></li>\n</ul>\n<p>Locally we trained 5 fold for each model, in order to get an OOF cv score. In inference, only one fold from each model is used.</p>\n<h2>Classification Stage 1</h2>\n<ul>\n<li>Input is 6 channel (3 ch original image + 3 ch predicted masks)</li>\n<li>Output is 12 classes: original 11 plus no_ETT, defined as whether all 3 ETT classes are 0</li>\n<li>Loss is weighted average of CE loss for the 4 ETT classes and BCE loss for the other 8 classes, with weight being 1:7</li>\n<li>Train on 30k data; make pseudo labels on 28k external data</li>\n<li>20 model ensemble, a mixture of EfficientNets, ResNets, ResNexts, ViTs at size 384 to 512, with various pre-processing parameters</li>\n<li>CV = <strong>0.97553</strong> with rank ensemble</li>\n</ul>\n<h2>Classification Stage 2</h2>\n<ul>\n<li>Same input, output, loss as Stage 1</li>\n<li>Training with 30k+28k data (combination of GT and pseudo labels)</li>\n<li>Since Stage 2 models need to go into inference kernel, overall model sizes are smaller than Stage 1</li>\n<li>31 model ensemble, a mix of EfficientNets, ResNets, SEResNexts, ResNexts, RegNet, Inception, RexNet, DenseNet, ViTs etc at size 384 to 512, with various pre-processing parameters</li>\n<li>CV = <strong>0.97606</strong> with rank ensemble</li>\n</ul>\n<p>Trained 5 fold locally, but only squeezed 67 folds into the inference kernel. </p>\n<h2>Post-processing</h2>\n<p>For AUC metric, sometimes it’s better to rank the probabilities of each model before ensembling. For this competition, we found that this is true only for these 5 columns:</p>\n<pre><code>            'ETT - Abnormal',\n            'NGT - Borderline',\n            'NGT - Incompletely Imaged',\n            'CVC - Normal',\n            'Swan Ganz Catheter Present'\n</code></pre>\n<p>This boosts CV score by about 0.00032</p>\n<h2>update</h2>\n<p>We have released simplified training and inference code:<br>\n<a href=\"https://www.kaggle.com/haqishen/ranzcr-1st-place-soluiton-seg-model-small-ver\" target=\"_blank\">Segmentation training</a>,  <a href=\"https://www.kaggle.com/haqishen/ranzcr-1st-place-soluiton-cls-model-small-ver\" target=\"_blank\">Classification training</a>,  <a href=\"https://www.kaggle.com/haqishen/ranzcr-1st-place-soluiton-inference-small-ver\" target=\"_blank\">Inference</a></p>",
      "rawMarkdown": "Congrats to all the winners. Thank you for great collaboration, my long time teammates @haqishen and @garybios. I had a blast once again working with such talented teammates.\n\n\n## TL;DR\n2-stage segmentation and 2-stage classification pipeline. Using pseudo labels in both segmentation and classification.\n\n\n## Dataset definitions\nDifferent subsets of the NIH ChestX dataset (112k images) are illustrated in this Venn diagram. \n- Official dataset (30k) contains 9k images with tube segmentation ground truth. \n- We also used Dr. Konya’s @sandorkonya trachea bifurcation annotation [dataset](https://www.kaggle.com/sandorkonya/5k-trachea-bifurcation-on-chest-xray) which has 5k images. Thank you Doctor!\n- For pseudo labeling, we identified 28k external images outside the Train set which contain tubes: (1) first run a segmentation model inference on all the 112k images to filter out all the images without tubes; (2) using imagehash to de-duplicate the 30k images that are already in 30k train set; (3) link patient IDs to make sure same patient in external data and original data falls into the same fold.\n![](https://i.imgur.com/KJQumVO.png)\n## Pre-processing\nFor images, we applied [“Ben’s pre-processing”](https://www.kaggle.com/ratthachat/aptos-eye-preprocessing-in-diabetic-retinopathy#2.-Try-Ben-Graham's-preprocessing-method.) with different parameters, in order to train diverse models.\n\nFor official segmentation annotation, we make 2 channel masks by drawing lines representing the tube, and drawing big dots indicating the tips of tubes. See pictures below.\n\nFor trachea bifurcation annotation, we make 1 channel masks by drawing big dots.\n\n\n## Segmentation Stage 1\n![](https://i.imgur.com/xdP22R1.png)\n- Model 1: Mask is tube and tips - 2 channel output\n    - Train and validate on 9k images with tube anno\n    - Pseudo label on 28k + (30k - 9k) data without tube anno\n    - 10 model ensemble with a mixture of Unet and Unet++, with backbones B3-B8, and different preprocessing parameters, at image size 1024x1024 to 1536x1536.\n- Model 2: Trachea bifurcation (TB) - 1 channel output\n    - Train and validate on 5k images with TB anno\n    - Pseudo label on 28k + (30k - 5k) data without TB anno\n    - Similar ensemble as model 1. But TBs are easier to segment, so the image sizes are 384x384 to 1024x1024\n\n## Segmentation Stage 2\n![](https://i.imgur.com/r4fI8Iz.png)\n- Tubes, tips and TB – 3 channel output\n    - Train on 30k + 28k images (with combination of GT and pseudo labels)\n    - Validate tubes and tips on 9k data\n    - Validate TB separately on 5k data\n    - Predict out-of-fold on all 30k + 28k images, to be used by classification\n- Stage 2 needs to run in inference kernel, so there are only 5 models in the ensemble:\n    - Unet++ B3 at 1536\n    - Unet B4 at 1536\n    - Unet++ B5 at 1024\n    - Unet++ B6 at 1024\n    - Unet B7 at 1024\n\nLocally we trained 5 fold for each model, in order to get an OOF cv score. In inference, only one fold from each model is used.\n\n## Classification Stage 1\n- Input is 6 channel (3 ch original image + 3 ch predicted masks)\n- Output is 12 classes: original 11 plus no_ETT, defined as whether all 3 ETT classes are 0\n- Loss is weighted average of CE loss for the 4 ETT classes and BCE loss for the other 8 classes, with weight being 1:7\n- Train on 30k data; make pseudo labels on 28k external data\n- 20 model ensemble, a mixture of EfficientNets, ResNets, ResNexts, ViTs at size 384 to 512, with various pre-processing parameters\n- CV = **0.97553** with rank ensemble\n\n## Classification Stage 2\n- Same input, output, loss as Stage 1\n- Training with 30k+28k data (combination of GT and pseudo labels)\n- Since Stage 2 models need to go into inference kernel, overall model sizes are smaller than Stage 1\n- 31 model ensemble, a mix of EfficientNets, ResNets, SEResNexts, ResNexts, RegNet, Inception, RexNet, DenseNet, ViTs etc at size 384 to 512, with various pre-processing parameters\n- CV = **0.97606** with rank ensemble\n\nTrained 5 fold locally, but only squeezed 67 folds into the inference kernel. \n\n## Post-processing\nFor AUC metric, sometimes it’s better to rank the probabilities of each model before ensembling. For this competition, we found that this is true only for these 5 columns:\n```\n            'ETT - Abnormal',\n            'NGT - Borderline',\n            'NGT - Incompletely Imaged',\n            'CVC - Normal',\n            'Swan Ganz Catheter Present'\n```\nThis boosts CV score by about 0.00032\n\n## update\nWe have released simplified training and inference code:\n[Segmentation training](https://www.kaggle.com/haqishen/ranzcr-1st-place-soluiton-seg-model-small-ver),  [Classification training](https://www.kaggle.com/haqishen/ranzcr-1st-place-soluiton-cls-model-small-ver),  [Inference](https://www.kaggle.com/haqishen/ranzcr-1st-place-soluiton-inference-small-ver)",
      "votes": null
    },
    {
      "id": "1241692",
      "postDate": "03/17/2021 06:50:55",
      "content": "<p>Congratulations on the win ! Great work !<br>\nWould you like to share your team's hardware information as well?</p>",
      "rawMarkdown": "Congratulations on the win ! Great work !\nWould you like to share your team's hardware information as well?",
      "votes": null
    },
    {
      "id": "1241703",
      "postDate": "03/17/2021 06:59:24",
      "content": "<p>Thanks.</p>\n<p>I have NVIDIA DGX Station with V100 GPUs. I think my teammates both have HP Z8G4 Workstation with NVIDIA RTX6000 GPUs and HP ZBook with NVIDIA RTX5000 GPU.</p>",
      "rawMarkdown": "Thanks.\n\nI have NVIDIA DGX Station with V100 GPUs. I think my teammates both have HP Z8G4 Workstation with NVIDIA RTX6000 GPUs and HP ZBook with NVIDIA RTX5000 GPU.",
      "votes": null
    },
    {
      "id": "1241712",
      "postDate": "03/17/2021 07:10:34",
      "content": "<p>Wow great solutions! Well deserve. Congrats <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> and team on winning. Thanks for sharing details solution </p>",
      "rawMarkdown": "Wow great solutions! Well deserve. Congrats @boliu0 and team on winning. Thanks for sharing details solution",
      "votes": null
    },
    {
      "id": "1241741",
      "postDate": "03/17/2021 07:22:04",
      "content": "<p>Wow great solutions! Well deserve. Congrats <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> and team on winning.<br>\nWe didnt find any way to use <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> dataset but learn a lot from him<br>\nAlso the idea of segmentation struck up to me 5 hours before the deadline</p>",
      "rawMarkdown": "Wow great solutions! Well deserve. Congrats @boliu0 and team on winning.\nWe didnt find any way to use @sandorkonya dataset but learn a lot from him\nAlso the idea of segmentation struck up to me 5 hours before the deadline",
      "votes": null
    },
    {
      "id": "1241778",
      "postDate": "03/17/2021 07:38:41",
      "content": "<p>Subarashii! Omedetou</p>",
      "rawMarkdown": "Subarashii! Omedetou",
      "votes": null
    },
    {
      "id": "1241817",
      "postDate": "03/17/2021 08:04:28",
      "content": "<p>Congratulations! Great solution :) One question. Did you use soft pseudo-labels or somehow make them hard pseudo-labels?</p>",
      "rawMarkdown": "Congratulations! Great solution :) One question. Did you use soft pseudo-labels or somehow make them hard pseudo-labels?",
      "votes": null
    },
    {
      "id": "1241857",
      "postDate": "03/17/2021 08:51:26",
      "content": "<p>First place! Congratulations!<br>\nQuestion.<br>\n1:Does 6 channel take more time for trainning compared to 3 channel?　I thought it would take longer in my opinion.<br>\n2:Since the image size is smaller, is the relative time to achieve high accuracy the same?</p>",
      "rawMarkdown": "First place! Congratulations!\nQuestion.\n1:Does 6 channel take more time for trainning compared to 3 channel?　I thought it would take longer in my opinion.\n2:Since the image size is smaller, is the relative time to achieve high accuracy the same?",
      "votes": null
    },
    {
      "id": "1241883",
      "postDate": "03/17/2021 09:01:09",
      "content": "<p>Thanks for sharing! Congratulations to first place!</p>\n<p>I have a question regarding segmentation stage 1: Here it says \"Pseudo label on 28k + (30k - 9k) data without tube anno\"<br>\nI was wondering where you use these pseudo labels, because for seg. stage 2 you wrote: \"Train on 30k + 28k images (with combination of GT and pseudo labels)\" <br>\n2nd question: Which augmentation did you use? </p>",
      "rawMarkdown": "Thanks for sharing! Congratulations to first place!\n\nI have a question regarding segmentation stage 1: Here it says \"Pseudo label on 28k + (30k - 9k) data without tube anno\"\nI was wondering where you use these pseudo labels, because for seg. stage 2 you wrote: \"Train on 30k + 28k images (with combination of GT and pseudo labels)\" \n2nd question: Which augmentation did you use?",
      "votes": null
    },
    {
      "id": "1241928",
      "postDate": "03/17/2021 09:29:06",
      "content": "<p>Congrats ! Thank you for sharing.</p>",
      "rawMarkdown": "Congrats ! Thank you for sharing.",
      "votes": null
    },
    {
      "id": "1241986",
      "postDate": "03/17/2021 10:18:32",
      "content": "<p>Just WOW …… Congratulations.</p>",
      "rawMarkdown": "Just WOW ...... Congratulations.",
      "votes": null
    },
    {
      "id": "1242030",
      "postDate": "03/17/2021 10:59:25",
      "content": "<p><a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> Wow ..Congratulations on First Place Finish . Keep up the great work </p>",
      "rawMarkdown": "boliu0 Wow ..Congratulations on First Place Finish . Keep up the great work",
      "votes": null
    },
    {
      "id": "1242065",
      "postDate": "03/17/2021 11:33:55",
      "content": "<p>just wondering if there was any reason to use 6 channels instead of 4 channels (as input images are grayscale?) </p>",
      "rawMarkdown": "just wondering if there was any reason to use 6 channels instead of 4 channels (as input images are grayscale?)",
      "votes": null
    },
    {
      "id": "1242078",
      "postDate": "03/17/2021 11:43:41",
      "content": "<p>We want to fully utilize the imagenet pretrained weights so we read\boriginal image as 3ch.</p>",
      "rawMarkdown": "We want to fully utilize the imagenet pretrained weights so we read\boriginal image as 3ch.",
      "votes": null
    },
    {
      "id": "1242085",
      "postDate": "03/17/2021 11:48:26",
      "content": "<p>Thanks!</p>\n<blockquote>\n  <p>Does 6 channel take more time for trainning compared to 3 channel?</p>\n</blockquote>\n<p>I haven't measured it but I think the training time may have gotten longer by about 0.1% by using 6ch instead of 3ch.</p>\n<blockquote>\n  <p>Since the image size is smaller, is the relative time to achieve high accuracy the same?</p>\n</blockquote>\n<p>In our pipeline, the cls model has no boost when using <code>img_size &gt; 512</code>.</p>",
      "rawMarkdown": "Thanks!\n\n> Does 6 channel take more time for trainning compared to 3 channel?\n\nI haven't measured it but I think the training time may have gotten longer by about 0.1% by using 6ch instead of 3ch.\n\n> Since the image size is smaller, is the relative time to achieve high accuracy the same?\n\nIn our pipeline, the cls model has no boost when using `img_size > 512`.",
      "votes": null
    },
    {
      "id": "1242107",
      "postDate": "03/17/2021 11:59:02",
      "content": "<p>0.1%! That's almost the same! That was surprising.<br>\nI'll try it myself. Thank you!</p>",
      "rawMarkdown": "0.1%! That's almost the same! That was surprising.\nI'll try it myself. Thank you!",
      "votes": null
    },
    {
      "id": "1242226",
      "postDate": "03/17/2021 13:39:06",
      "content": "<p>Thanks. We used soft pseudo labels for both segmentation and classification. They work better than hard ones as they contain more information.</p>",
      "rawMarkdown": "Thanks. We used soft pseudo labels for both segmentation and classification. They work better than hard ones as they contain more information.",
      "votes": null
    },
    {
      "id": "1242237",
      "postDate": "03/17/2021 13:44:43",
      "content": "<p>Congratulations with the first place and thanks for sharing this great summary! Very nice way to incorporate annotations and external data in the solution.</p>",
      "rawMarkdown": "Congratulations with the first place and thanks for sharing this great summary! Very nice way to incorporate annotations and external data in the solution.",
      "votes": null
    },
    {
      "id": "1242240",
      "postDate": "03/17/2021 13:46:37",
      "content": "<p>Thanks. </p>\n<p>Q1: We have tube annotation for the 9k images, so we only need to make pseudo labels for the other 28k+30k-9k images. The goal for stage 1 is to make pseudo labels, to enable stage 2 to train on all 28k+30k data.</p>\n<p>Then in stage 2, we have labels for all 28k+30k data (among which 9k are GT labels, the rest being pseudo labels). So we train on all of them.</p>\n<p>Q2. For segmentation: HorizontalFlip, RandomBrightness, ShiftScaleRotate, Cutout<br>\nFor classification: all the above, plus RandomContrast, OpticalDistortion, GridDistortion, HueSaturationValue</p>",
      "rawMarkdown": "Thanks. \n\nQ1: We have tube annotation for the 9k images, so we only need to make pseudo labels for the other 28k+30k-9k images. The goal for stage 1 is to make pseudo labels, to enable stage 2 to train on all 28k+30k data.\n\nThen in stage 2, we have labels for all 28k+30k data (among which 9k are GT labels, the rest being pseudo labels). So we train on all of them.\n\nQ2. For segmentation: HorizontalFlip, RandomBrightness, ShiftScaleRotate, Cutout\nFor classification: all the above, plus RandomContrast, OpticalDistortion, GridDistortion, HueSaturationValue",
      "votes": null
    },
    {
      "id": "1242266",
      "postDate": "03/17/2021 14:02:19",
      "content": "<p>Congratulations, Lots of hard work. Thanks for sharing with us your team's great idea and solution. Can you share your codebase link? Thanks in advance. </p>",
      "rawMarkdown": "Congratulations, Lots of hard work. Thanks for sharing with us your team's great idea and solution. Can you share your codebase link? Thanks in advance.",
      "votes": null
    },
    {
      "id": "1242318",
      "postDate": "03/17/2021 14:42:35",
      "content": "<p>Thank you :)</p>",
      "rawMarkdown": "Thank you :)",
      "votes": null
    },
    {
      "id": "1242339",
      "postDate": "03/17/2021 14:55:00",
      "content": "<p>Wow! Just wowed! Congratulations!</p>",
      "rawMarkdown": "Wow! Just wowed! Congratulations!",
      "votes": null
    },
    {
      "id": "1242418",
      "postDate": "03/17/2021 15:48:52",
      "content": "<p>I'm still a little confused here. As I understand (please correct me if wrong), the stored images only have single channel information, which you can then load with cv2 as 3 channel or 1 channel, where the 3 channel loading just duplicates the single channel information. It intuitively makes sense to use 3 channel as you said to fully utilize imagenet pretrained weights whenever you're training on standard images. But once you move to adding in the 3ch segmentation mask, why should there be a difference whether 4 channel (1+3) vs 6 channel (3+3).</p>",
      "rawMarkdown": "I'm still a little confused here. As I understand (please correct me if wrong), the stored images only have single channel information, which you can then load with cv2 as 3 channel or 1 channel, where the 3 channel loading just duplicates the single channel information. It intuitively makes sense to use 3 channel as you said to fully utilize imagenet pretrained weights whenever you're training on standard images. But once you move to adding in the 3ch segmentation mask, why should there be a difference whether 4 channel (1+3) vs 6 channel (3+3).",
      "votes": null
    },
    {
      "id": "1242421",
      "postDate": "03/17/2021 15:49:54",
      "content": "<p>How were the Segmentation 1 models ensembled for pseudo labels/masks? </p>\n<p>Did you just average the masks? Was any smoothing/filter used for the mask predictions? </p>\n<p>The solution is really impressive…5 segmentation models and 67 classification models! Congrats!</p>",
      "rawMarkdown": "How were the Segmentation 1 models ensembled for pseudo labels/masks? \n\nDid you just average the masks? Was any smoothing/filter used for the mask predictions? \n\nThe solution is really impressive...5 segmentation models and 67 classification models! Congrats!",
      "votes": null
    },
    {
      "id": "1242425",
      "postDate": "03/17/2021 15:51:39",
      "content": "<p>Thank you, and I learn a lot from the kaggle community aswell =)<br>\nWin-win =)</p>",
      "rawMarkdown": "Thank you, and I learn a lot from the kaggle community aswell =)\nWin-win =)",
      "votes": null
    },
    {
      "id": "1242450",
      "postDate": "03/17/2021 16:13:45",
      "content": "<p>Thank you.</p>\n<p>Just simple average of soft labels, for both segmentation ensemble and classification ensemble. No weighting. No thresholding.</p>",
      "rawMarkdown": "Thank you.\n\nJust simple average of soft labels, for both segmentation ensemble and classification ensemble. No weighting. No thresholding.",
      "votes": null
    },
    {
      "id": "1242457",
      "postDate": "03/17/2021 16:16:52",
      "content": "<p>You're right. When we load the original image as 3 channel numpy array, the 3 channels are identical.</p>\n<p>It's just that when we started experiments, we were using 3 channel input (without masks), then later after adding 3 channel mask, we just kept it as 3+3 channels. But I believe 1+3 channels would work equally well as you suggested.</p>",
      "rawMarkdown": "You're right. When we load the original image as 3 channel numpy array, the 3 channels are identical.\n\nIt's just that when we started experiments, we were using 3 channel input (without masks), then later after adding 3 channel mask, we just kept it as 3+3 channels. But I believe 1+3 channels would work equally well as you suggested.",
      "votes": null
    },
    {
      "id": "1242556",
      "postDate": "03/17/2021 17:16:10",
      "content": "<p>Now I got it, thanks :)</p>",
      "rawMarkdown": "Now I got it, thanks :)",
      "votes": null
    },
    {
      "id": "1242620",
      "postDate": "03/17/2021 18:03:42",
      "content": "<p>I was wondering, since you trained more than 40(?) different model architectures,  how do you choose hyperparameters for your model? (img_size, optimizer, learning rate.. so many more). So far I've always done it \"by hand\", trying things and see if they work. I guess partly it is true for you as well, but do you also use some automated procedure? (e.g. train for a certain number of steps and pick the best combination)</p>\n<p>Congratulations btw, you reached a level of skill where Machine Learning becomes art. </p>",
      "rawMarkdown": "I was wondering, since you trained more than 40(?) different model architectures,  how do you choose hyperparameters for your model? (img_size, optimizer, learning rate.. so many more). So far I've always done it \"by hand\", trying things and see if they work. I guess partly it is true for you as well, but do you also use some automated procedure? (e.g. train for a certain number of steps and pick the best combination)\n\nCongratulations btw, you reached a level of skill where Machine Learning becomes art.",
      "votes": null
    },
    {
      "id": "1242674",
      "postDate": "03/17/2021 18:29:22",
      "content": "<p>Thanks.</p>\n<p>Yeah we also do this by hand. The AutoCV is still under development. 😉</p>\n<p>For optimizer, we like Adam. For image size, you can start with a small one like 256 or 384 depending on your hardware, then gradually increase to 512 or even larger after you tune your model on small images.</p>\n<p>Best learning rate varies somewhat for different model architecture, so they need to be tuned on each model. But you only need to run one fold per model for this. And it doesn't have to be exact. A slightly \"off\" learning rate for a single model won't affect ensemble's score by much.</p>",
      "rawMarkdown": "Thanks.\n\nYeah we also do this by hand. The AutoCV is still under development. 😉\n\nFor optimizer, we like Adam. For image size, you can start with a small one like 256 or 384 depending on your hardware, then gradually increase to 512 or even larger after you tune your model on small images.\n\nBest learning rate varies somewhat for different model architecture, so they need to be tuned on each model. But you only need to run one fold per model for this. And it doesn't have to be exact. A slightly \"off\" learning rate for a single model won't affect ensemble's score by much.",
      "votes": null
    },
    {
      "id": "1242754",
      "postDate": "03/17/2021 19:07:38",
      "content": "<p>Great thank you! And awesome solution!</p>",
      "rawMarkdown": "Great thank you! And awesome solution!",
      "votes": null
    },
    {
      "id": "1243532",
      "postDate": "03/18/2021 09:45:03",
      "content": "<p>Thank you for your fantastic solution. Can definitely learn a lot from it.</p>\n<blockquote>\n  <p>In our pipeline, the cls model has no boost when using img_size &gt; 512</p>\n</blockquote>\n<p>Do you have an intuition for it?</p>\n<p>I almost felt it should be the opposite, i.e. low resolution for segmentation but high resolution for classification. </p>\n<p>Also what downscaling method do you use to resize the segmentation output for classification, INTER_NEAREST, INTER_LINEAR, etc, and do they make any difference?</p>\n<p>Many thanks!</p>",
      "rawMarkdown": "Thank you for your fantastic solution. Can definitely learn a lot from it.\n\n> In our pipeline, the cls model has no boost when using img_size > 512\n\nDo you have an intuition for it?\n\nI almost felt it should be the opposite, i.e. low resolution for segmentation but high resolution for classification. \n\nAlso what downscaling method do you use to resize the segmentation output for classification, INTER_NEAREST, INTER_LINEAR, etc, and do they make any difference?\n\nMany thanks!",
      "votes": null
    },
    {
      "id": "1243778",
      "postDate": "03/18/2021 13:34:47",
      "content": "<p>Before experiments we can hardly tell anything.</p>\n<p>So we've compared (1024seg + 512cls) with (512seg + 1024cls) and the result is —— the former is far better than the latter.</p>\n<p>Based on the results of this experiment we can in turn make some explanations for the dataset. For example, segmentation requires more fine-grained detail, etc.</p>\n<p>So sometimes don't trust your intuition. Just trust the results of the experiment.</p>\n<p>While, we didn't pay any attention on downscaling method, just used the default one.</p>",
      "rawMarkdown": "Before experiments we can hardly tell anything.\n\nSo we've compared (1024seg + 512cls) with (512seg + 1024cls) and the result is —— the former is far better than the latter.\n\nBased on the results of this experiment we can in turn make some explanations for the dataset. For example, segmentation requires more fine-grained detail, etc.\n\nSo sometimes don't trust your intuition. Just trust the results of the experiment.\n\nWhile, we didn't pay any attention on downscaling method, just used the default one.",
      "votes": null
    },
    {
      "id": "1243938",
      "postDate": "03/18/2021 15:43:07",
      "content": "<p>Congrats! After seeing your 0.928 LB score we already knew you can win this :)</p>",
      "rawMarkdown": "Congrats! After seeing your 0.928 LB score we already knew you can win this :)",
      "votes": null
    },
    {
      "id": "1243954",
      "postDate": "03/18/2021 15:55:01",
      "content": "<p>Thank you. Oh, you noticed 😏</p>",
      "rawMarkdown": "Thank you. Oh, you noticed 😏",
      "votes": null
    },
    {
      "id": "1244151",
      "postDate": "03/18/2021 18:10:33",
      "content": "<p>We dont need a leaderboard because either you guys or <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> <a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> would take up the money. the host can freeze you guy at the top. thats more likely</p>",
      "rawMarkdown": "We dont need a leaderboard because either you guys or @bestfitting @wowfattie @philippsinger and @christofhenkel would take up the money. the host can freeze you guy at the top. thats more likely",
      "votes": null
    },
    {
      "id": "1244640",
      "postDate": "03/19/2021 05:40:08",
      "content": "<p>Great solution, congratulations! <br>\nI was wondering, how do you overcome submission time limit using more than 60 models in an inference? Do you use ONNX or maybe even pruning?</p>",
      "rawMarkdown": "Great solution, congratulations! \nI was wondering, how do you overcome submission time limit using more than 60 models in an inference? Do you use ONNX or maybe even pruning?",
      "votes": null
    },
    {
      "id": "1245138",
      "postDate": "03/19/2021 13:41:53",
      "content": "<p>Hi, thank you.</p>\n<p>No we didn't use ONNX or pruning. We put all the models on the GPU, then use one dataloader to iterate all test data only once. We maximized the 9 hour limit by squeezing in as many models as possible without timeout.</p>",
      "rawMarkdown": "Hi, thank you.\n\nNo we didn't use ONNX or pruning. We put all the models on the GPU, then use one dataloader to iterate all test data only once. We maximized the 9 hour limit by squeezing in as many models as possible without timeout.",
      "votes": null
    },
    {
      "id": "1261979",
      "postDate": "04/03/2021 16:36:46",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> </p>",
      "rawMarkdown": "Congratulations @boliu0",
      "votes": null
    },
    {
      "id": "1264319",
      "postDate": "04/06/2021 04:36:04",
      "content": "<p>Great Job! Congrats! <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a>  </p>\n<p>I wonder: For the 28k external data, (1) first run a segmentation model inference on all the 112k images to filter out all the images without tubes… How did you build this model, which dataset did you train on, or is there a one on Kaggle to do this? </p>\n<p>Thanks for sharing the solution.</p>",
      "rawMarkdown": "Great Job! Congrats! @boliu0 @haqishen  \n\nI wonder: For the 28k external data, (1) first run a segmentation model inference on all the 112k images to filter out all the images without tubes... How did you build this model, which dataset did you train on, or is there a one on Kaggle to do this? \n\nThanks for sharing the solution.",
      "votes": null
    },
    {
      "id": "1279718",
      "postDate": "04/21/2021 07:25:39",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> Congratulations! Quite an impressive solution! <br>\nI am rather new to AI and computer vision in particular. Is it possible to share with us the NIH data that you extracted (28k) or to share the steps needed to acquire this data from the full NIH (112k) dataset? </p>",
      "rawMarkdown": "Hello @boliu0 Congratulations! Quite an impressive solution! \nI am rather new to AI and computer vision in particular. Is it possible to share with us the NIH data that you extracted (28k) or to share the steps needed to acquire this data from the full NIH (112k) dataset?",
      "votes": null
    },
    {
      "id": "1280057",
      "postDate": "04/21/2021 14:03:21",
      "content": "<p>Hi, for some reason, Kaggle disabled attachment in forum.</p>\n<p>I uploaded the 28k images' IDs to a dataset: <a href=\"https://www.kaggle.com/boliu0/ranzcr-external-data-id\" target=\"_blank\">https://www.kaggle.com/boliu0/ranzcr-external-data-id</a></p>",
      "rawMarkdown": "Hi, for some reason, Kaggle disabled attachment in forum.\n\nI uploaded the 28k images' IDs to a dataset: https://www.kaggle.com/boliu0/ranzcr-external-data-id",
      "votes": null
    },
    {
      "id": "1280060",
      "postDate": "04/21/2021 14:05:33",
      "content": "<p>Hi, for this purpose, we just used a single model (Unet++, B5, 1024) trained on official annotation. </p>",
      "rawMarkdown": "Hi, for this purpose, we just used a single model (Unet++, B5, 1024) trained on official annotation.",
      "votes": null
    },
    {
      "id": "1566563",
      "postDate": "11/01/2021 02:49:47",
      "content": "<p>Excuse me, I want to understand whether pseudo masking was done with manual in the Segmentation Stage 1.<br>\nCongratulation win the champion!!</p>",
      "rawMarkdown": "Excuse me, I want to understand whether pseudo masking was done with manual in the Segmentation Stage 1.\nCongratulation win the champion!!",
      "votes": null
    },
    {
      "id": "1571730",
      "postDate": "11/05/2021 05:50:17",
      "content": "<p>I don't know if it has been answered  before, but I am confused about how can input with more than 3 channels be used in a pretrained model,<br>\nso I checked the code,there you have done<br>\n<code>self.enet.conv_stem.weight = nn.Parameter(self.enet.conv_stem.weight.repeat(1,n_ch//3+1,1,1)[:, :n_ch])</code><br>\ncan someone explain what its doing?</p>",
      "rawMarkdown": "I don't know if it has been answered  before, but I am confused about how can input with more than 3 channels be used in a pretrained model,\nso I checked the code,there you have done\n`self.enet.conv_stem.weight = nn.Parameter(self.enet.conv_stem.weight.repeat(1,n_ch//3+1,1,1)[:, :n_ch])`\ncan someone explain what its doing?",
      "votes": null
    },
    {
      "id": "2063518",
      "postDate": "12/13/2022 02:31:35",
      "content": "<p>This is really great! Thank you for sharing the idea.</p>",
      "rawMarkdown": "This is really great! Thank you for sharing the idea.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1241692,
      "author_name": "lhagiimn",
      "author_url": "",
      "post_date": "03/17/2021 06:50:55",
      "content": "<p>Congratulations on the win ! Great work !<br>\nWould you like to share your team's hardware information as well?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1241703,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "03/17/2021 06:59:24",
          "content": "<p>Thanks.</p>\n<p>I have NVIDIA DGX Station with V100 GPUs. I think my teammates both have HP Z8G4 Workstation with NVIDIA RTX6000 GPUs and HP ZBook with NVIDIA RTX5000 GPU.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1241712,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "03/17/2021 07:10:34",
      "content": "<p>Wow great solutions! Well deserve. Congrats <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> and team on winning. Thanks for sharing details solution </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1241741,
      "author_name": "morizin",
      "author_url": "",
      "post_date": "03/17/2021 07:22:04",
      "content": "<p>Wow great solutions! Well deserve. Congrats <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> and team on winning.<br>\nWe didnt find any way to use <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> dataset but learn a lot from him<br>\nAlso the idea of segmentation struck up to me 5 hours before the deadline</p>",
      "votes": null,
      "replies": [
        {
          "id": 1242425,
          "author_name": "sandorkonya",
          "author_url": "",
          "post_date": "03/17/2021 15:51:39",
          "content": "<p>Thank you, and I learn a lot from the kaggle community aswell =)<br>\nWin-win =)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1241778,
      "author_name": "nvminno",
      "author_url": "",
      "post_date": "03/17/2021 07:38:41",
      "content": "<p>Subarashii! Omedetou</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1241817,
      "author_name": "ademyanchuk",
      "author_url": "",
      "post_date": "03/17/2021 08:04:28",
      "content": "<p>Congratulations! Great solution :) One question. Did you use soft pseudo-labels or somehow make them hard pseudo-labels?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1242226,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "03/17/2021 13:39:06",
          "content": "<p>Thanks. We used soft pseudo labels for both segmentation and classification. They work better than hard ones as they contain more information.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1242318,
          "author_name": "ademyanchuk",
          "author_url": "",
          "post_date": "03/17/2021 14:42:35",
          "content": "<p>Thank you :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1241857,
      "author_name": "shinoda18",
      "author_url": "",
      "post_date": "03/17/2021 08:51:26",
      "content": "<p>First place! Congratulations!<br>\nQuestion.<br>\n1:Does 6 channel take more time for trainning compared to 3 channel?　I thought it would take longer in my opinion.<br>\n2:Since the image size is smaller, is the relative time to achieve high accuracy the same?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1242085,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "03/17/2021 11:48:26",
          "content": "<p>Thanks!</p>\n<blockquote>\n  <p>Does 6 channel take more time for trainning compared to 3 channel?</p>\n</blockquote>\n<p>I haven't measured it but I think the training time may have gotten longer by about 0.1% by using 6ch instead of 3ch.</p>\n<blockquote>\n  <p>Since the image size is smaller, is the relative time to achieve high accuracy the same?</p>\n</blockquote>\n<p>In our pipeline, the cls model has no boost when using <code>img_size &gt; 512</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1242107,
          "author_name": "shinoda18",
          "author_url": "",
          "post_date": "03/17/2021 11:59:02",
          "content": "<p>0.1%! That's almost the same! That was surprising.<br>\nI'll try it myself. Thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1243532,
          "author_name": "yl1202",
          "author_url": "",
          "post_date": "03/18/2021 09:45:03",
          "content": "<p>Thank you for your fantastic solution. Can definitely learn a lot from it.</p>\n<blockquote>\n  <p>In our pipeline, the cls model has no boost when using img_size &gt; 512</p>\n</blockquote>\n<p>Do you have an intuition for it?</p>\n<p>I almost felt it should be the opposite, i.e. low resolution for segmentation but high resolution for classification. </p>\n<p>Also what downscaling method do you use to resize the segmentation output for classification, INTER_NEAREST, INTER_LINEAR, etc, and do they make any difference?</p>\n<p>Many thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1243778,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "03/18/2021 13:34:47",
          "content": "<p>Before experiments we can hardly tell anything.</p>\n<p>So we've compared (1024seg + 512cls) with (512seg + 1024cls) and the result is —— the former is far better than the latter.</p>\n<p>Based on the results of this experiment we can in turn make some explanations for the dataset. For example, segmentation requires more fine-grained detail, etc.</p>\n<p>So sometimes don't trust your intuition. Just trust the results of the experiment.</p>\n<p>While, we didn't pay any attention on downscaling method, just used the default one.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1241883,
      "author_name": "maschwa",
      "author_url": "",
      "post_date": "03/17/2021 09:01:09",
      "content": "<p>Thanks for sharing! Congratulations to first place!</p>\n<p>I have a question regarding segmentation stage 1: Here it says \"Pseudo label on 28k + (30k - 9k) data without tube anno\"<br>\nI was wondering where you use these pseudo labels, because for seg. stage 2 you wrote: \"Train on 30k + 28k images (with combination of GT and pseudo labels)\" <br>\n2nd question: Which augmentation did you use? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1242240,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "03/17/2021 13:46:37",
          "content": "<p>Thanks. </p>\n<p>Q1: We have tube annotation for the 9k images, so we only need to make pseudo labels for the other 28k+30k-9k images. The goal for stage 1 is to make pseudo labels, to enable stage 2 to train on all 28k+30k data.</p>\n<p>Then in stage 2, we have labels for all 28k+30k data (among which 9k are GT labels, the rest being pseudo labels). So we train on all of them.</p>\n<p>Q2. For segmentation: HorizontalFlip, RandomBrightness, ShiftScaleRotate, Cutout<br>\nFor classification: all the above, plus RandomContrast, OpticalDistortion, GridDistortion, HueSaturationValue</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1242556,
          "author_name": "maschwa",
          "author_url": "",
          "post_date": "03/17/2021 17:16:10",
          "content": "<p>Now I got it, thanks :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1241928,
      "author_name": "ttahara",
      "author_url": "",
      "post_date": "03/17/2021 09:29:06",
      "content": "<p>Congrats ! Thank you for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1241986,
      "author_name": "ammarali32",
      "author_url": "",
      "post_date": "03/17/2021 10:18:32",
      "content": "<p>Just WOW …… Congratulations.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1242030,
      "author_name": "usharengaraju",
      "author_url": "",
      "post_date": "03/17/2021 10:59:25",
      "content": "<p><a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> Wow ..Congratulations on First Place Finish . Keep up the great work </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1242065,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "03/17/2021 11:33:55",
      "content": "<p>just wondering if there was any reason to use 6 channels instead of 4 channels (as input images are grayscale?) </p>",
      "votes": null,
      "replies": [
        {
          "id": 1242078,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "03/17/2021 11:43:41",
          "content": "<p>We want to fully utilize the imagenet pretrained weights so we read\boriginal image as 3ch.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1242418,
          "author_name": "jcsagar",
          "author_url": "",
          "post_date": "03/17/2021 15:48:52",
          "content": "<p>I'm still a little confused here. As I understand (please correct me if wrong), the stored images only have single channel information, which you can then load with cv2 as 3 channel or 1 channel, where the 3 channel loading just duplicates the single channel information. It intuitively makes sense to use 3 channel as you said to fully utilize imagenet pretrained weights whenever you're training on standard images. But once you move to adding in the 3ch segmentation mask, why should there be a difference whether 4 channel (1+3) vs 6 channel (3+3).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1242457,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "03/17/2021 16:16:52",
          "content": "<p>You're right. When we load the original image as 3 channel numpy array, the 3 channels are identical.</p>\n<p>It's just that when we started experiments, we were using 3 channel input (without masks), then later after adding 3 channel mask, we just kept it as 3+3 channels. But I believe 1+3 channels would work equally well as you suggested.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1242754,
          "author_name": "jcsagar",
          "author_url": "",
          "post_date": "03/17/2021 19:07:38",
          "content": "<p>Great thank you! And awesome solution!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1242237,
      "author_name": "kozodoi",
      "author_url": "",
      "post_date": "03/17/2021 13:44:43",
      "content": "<p>Congratulations with the first place and thanks for sharing this great summary! Very nice way to incorporate annotations and external data in the solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1242266,
      "author_name": "durbin164",
      "author_url": "",
      "post_date": "03/17/2021 14:02:19",
      "content": "<p>Congratulations, Lots of hard work. Thanks for sharing with us your team's great idea and solution. Can you share your codebase link? Thanks in advance. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1242339,
      "author_name": "drcodikpollonny",
      "author_url": "",
      "post_date": "03/17/2021 14:55:00",
      "content": "<p>Wow! Just wowed! Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1242421,
      "author_name": "jy2tong",
      "author_url": "",
      "post_date": "03/17/2021 15:49:54",
      "content": "<p>How were the Segmentation 1 models ensembled for pseudo labels/masks? </p>\n<p>Did you just average the masks? Was any smoothing/filter used for the mask predictions? </p>\n<p>The solution is really impressive…5 segmentation models and 67 classification models! Congrats!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1242450,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "03/17/2021 16:13:45",
          "content": "<p>Thank you.</p>\n<p>Just simple average of soft labels, for both segmentation ensemble and classification ensemble. No weighting. No thresholding.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1242620,
      "author_name": "nickthenick23",
      "author_url": "",
      "post_date": "03/17/2021 18:03:42",
      "content": "<p>I was wondering, since you trained more than 40(?) different model architectures,  how do you choose hyperparameters for your model? (img_size, optimizer, learning rate.. so many more). So far I've always done it \"by hand\", trying things and see if they work. I guess partly it is true for you as well, but do you also use some automated procedure? (e.g. train for a certain number of steps and pick the best combination)</p>\n<p>Congratulations btw, you reached a level of skill where Machine Learning becomes art. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1242674,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "03/17/2021 18:29:22",
          "content": "<p>Thanks.</p>\n<p>Yeah we also do this by hand. The AutoCV is still under development. 😉</p>\n<p>For optimizer, we like Adam. For image size, you can start with a small one like 256 or 384 depending on your hardware, then gradually increase to 512 or even larger after you tune your model on small images.</p>\n<p>Best learning rate varies somewhat for different model architecture, so they need to be tuned on each model. But you only need to run one fold per model for this. And it doesn't have to be exact. A slightly \"off\" learning rate for a single model won't affect ensemble's score by much.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1243938,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "03/18/2021 15:43:07",
      "content": "<p>Congrats! After seeing your 0.928 LB score we already knew you can win this :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1243954,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "03/18/2021 15:55:01",
          "content": "<p>Thank you. Oh, you noticed 😏</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1244151,
          "author_name": "morizin",
          "author_url": "",
          "post_date": "03/18/2021 18:10:33",
          "content": "<p>We dont need a leaderboard because either you guys or <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> <a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> would take up the money. the host can freeze you guy at the top. thats more likely</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1244640,
      "author_name": "kseniar",
      "author_url": "",
      "post_date": "03/19/2021 05:40:08",
      "content": "<p>Great solution, congratulations! <br>\nI was wondering, how do you overcome submission time limit using more than 60 models in an inference? Do you use ONNX or maybe even pruning?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1245138,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "03/19/2021 13:41:53",
          "content": "<p>Hi, thank you.</p>\n<p>No we didn't use ONNX or pruning. We put all the models on the GPU, then use one dataloader to iterate all test data only once. We maximized the 9 hour limit by squeezing in as many models as possible without timeout.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1261979,
      "author_name": "sohailds",
      "author_url": "",
      "post_date": "04/03/2021 16:36:46",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1264319,
      "author_name": "tcallioglu",
      "author_url": "",
      "post_date": "04/06/2021 04:36:04",
      "content": "<p>Great Job! Congrats! <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a>  </p>\n<p>I wonder: For the 28k external data, (1) first run a segmentation model inference on all the 112k images to filter out all the images without tubes… How did you build this model, which dataset did you train on, or is there a one on Kaggle to do this? </p>\n<p>Thanks for sharing the solution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1280060,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "04/21/2021 14:05:33",
          "content": "<p>Hi, for this purpose, we just used a single model (Unet++, B5, 1024) trained on official annotation. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1279718,
      "author_name": "melissabeaini",
      "author_url": "",
      "post_date": "04/21/2021 07:25:39",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> Congratulations! Quite an impressive solution! <br>\nI am rather new to AI and computer vision in particular. Is it possible to share with us the NIH data that you extracted (28k) or to share the steps needed to acquire this data from the full NIH (112k) dataset? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1280057,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "04/21/2021 14:03:21",
          "content": "<p>Hi, for some reason, Kaggle disabled attachment in forum.</p>\n<p>I uploaded the 28k images' IDs to a dataset: <a href=\"https://www.kaggle.com/boliu0/ranzcr-external-data-id\" target=\"_blank\">https://www.kaggle.com/boliu0/ranzcr-external-data-id</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1566563,
      "author_name": "kaiamster",
      "author_url": "",
      "post_date": "11/01/2021 02:49:47",
      "content": "<p>Excuse me, I want to understand whether pseudo masking was done with manual in the Segmentation Stage 1.<br>\nCongratulation win the champion!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1571730,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "11/05/2021 05:50:17",
      "content": "<p>I don't know if it has been answered  before, but I am confused about how can input with more than 3 channels be used in a pretrained model,<br>\nso I checked the code,there you have done<br>\n<code>self.enet.conv_stem.weight = nn.Parameter(self.enet.conv_stem.weight.repeat(1,n_ch//3+1,1,1)[:, :n_ch])</code><br>\ncan someone explain what its doing?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2063518,
      "author_name": "kimjin2510",
      "author_url": "",
      "post_date": "12/13/2022 02:31:35",
      "content": "<p>This is really great! Thank you for sharing the idea.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1241669": "Congrats to all the winners. Thank you for great collaboration, my long time teammates @haqishen and @garybios. I had a blast once again working with such talented teammates.\n\n\n## TL;DR\n2-stage segmentation and 2-stage classification pipeline. Using pseudo labels in both segmentation and classification.\n\n\n## Dataset definitions\nDifferent subsets of the NIH ChestX dataset (112k images) are illustrated in this Venn diagram. \n- Official dataset (30k) contains 9k images with tube segmentation ground truth. \n- We also used Dr. Konya’s @sandorkonya trachea bifurcation annotation [dataset](https://www.kaggle.com/sandorkonya/5k-trachea-bifurcation-on-chest-xray) which has 5k images. Thank you Doctor!\n- For pseudo labeling, we identified 28k external images outside the Train set which contain tubes: (1) first run a segmentation model inference on all the 112k images to filter out all the images without tubes; (2) using imagehash to de-duplicate the 30k images that are already in 30k train set; (3) link patient IDs to make sure same patient in external data and original data falls into the same fold.\n![](https://i.imgur.com/KJQumVO.png)\n## Pre-processing\nFor images, we applied [“Ben’s pre-processing”](https://www.kaggle.com/ratthachat/aptos-eye-preprocessing-in-diabetic-retinopathy#2.-Try-Ben-Graham's-preprocessing-method.) with different parameters, in order to train diverse models.\n\nFor official segmentation annotation, we make 2 channel masks by drawing lines representing the tube, and drawing big dots indicating the tips of tubes. See pictures below.\n\nFor trachea bifurcation annotation, we make 1 channel masks by drawing big dots.\n\n\n## Segmentation Stage 1\n![](https://i.imgur.com/xdP22R1.png)\n- Model 1: Mask is tube and tips - 2 channel output\n    - Train and validate on 9k images with tube anno\n    - Pseudo label on 28k + (30k - 9k) data without tube anno\n    - 10 model ensemble with a mixture of Unet and Unet++, with backbones B3-B8, and different preprocessing parameters, at image size 1024x1024 to 1536x1536.\n- Model 2: Trachea bifurcation (TB) - 1 channel output\n    - Train and validate on 5k images with TB anno\n    - Pseudo label on 28k + (30k - 5k) data without TB anno\n    - Similar ensemble as model 1. But TBs are easier to segment, so the image sizes are 384x384 to 1024x1024\n\n## Segmentation Stage 2\n![](https://i.imgur.com/r4fI8Iz.png)\n- Tubes, tips and TB – 3 channel output\n    - Train on 30k + 28k images (with combination of GT and pseudo labels)\n    - Validate tubes and tips on 9k data\n    - Validate TB separately on 5k data\n    - Predict out-of-fold on all 30k + 28k images, to be used by classification\n- Stage 2 needs to run in inference kernel, so there are only 5 models in the ensemble:\n    - Unet++ B3 at 1536\n    - Unet B4 at 1536\n    - Unet++ B5 at 1024\n    - Unet++ B6 at 1024\n    - Unet B7 at 1024\n\nLocally we trained 5 fold for each model, in order to get an OOF cv score. In inference, only one fold from each model is used.\n\n## Classification Stage 1\n- Input is 6 channel (3 ch original image + 3 ch predicted masks)\n- Output is 12 classes: original 11 plus no_ETT, defined as whether all 3 ETT classes are 0\n- Loss is weighted average of CE loss for the 4 ETT classes and BCE loss for the other 8 classes, with weight being 1:7\n- Train on 30k data; make pseudo labels on 28k external data\n- 20 model ensemble, a mixture of EfficientNets, ResNets, ResNexts, ViTs at size 384 to 512, with various pre-processing parameters\n- CV = **0.97553** with rank ensemble\n\n## Classification Stage 2\n- Same input, output, loss as Stage 1\n- Training with 30k+28k data (combination of GT and pseudo labels)\n- Since Stage 2 models need to go into inference kernel, overall model sizes are smaller than Stage 1\n- 31 model ensemble, a mix of EfficientNets, ResNets, SEResNexts, ResNexts, RegNet, Inception, RexNet, DenseNet, ViTs etc at size 384 to 512, with various pre-processing parameters\n- CV = **0.97606** with rank ensemble\n\nTrained 5 fold locally, but only squeezed 67 folds into the inference kernel. \n\n## Post-processing\nFor AUC metric, sometimes it’s better to rank the probabilities of each model before ensembling. For this competition, we found that this is true only for these 5 columns:\n```\n            'ETT - Abnormal',\n            'NGT - Borderline',\n            'NGT - Incompletely Imaged',\n            'CVC - Normal',\n            'Swan Ganz Catheter Present'\n```\nThis boosts CV score by about 0.00032\n\n## update\nWe have released simplified training and inference code:\n[Segmentation training](https://www.kaggle.com/haqishen/ranzcr-1st-place-soluiton-seg-model-small-ver),  [Classification training](https://www.kaggle.com/haqishen/ranzcr-1st-place-soluiton-cls-model-small-ver),  [Inference](https://www.kaggle.com/haqishen/ranzcr-1st-place-soluiton-inference-small-ver)",
    "1241692": "Congratulations on the win ! Great work !\nWould you like to share your team's hardware information as well?",
    "1241703": "Thanks.\n\nI have NVIDIA DGX Station with V100 GPUs. I think my teammates both have HP Z8G4 Workstation with NVIDIA RTX6000 GPUs and HP ZBook with NVIDIA RTX5000 GPU.",
    "1241712": "Wow great solutions! Well deserve. Congrats @boliu0 and team on winning. Thanks for sharing details solution",
    "1241741": "Wow great solutions! Well deserve. Congrats @boliu0 and team on winning.\nWe didnt find any way to use @sandorkonya dataset but learn a lot from him\nAlso the idea of segmentation struck up to me 5 hours before the deadline",
    "1241778": "Subarashii! Omedetou",
    "1241817": "Congratulations! Great solution :) One question. Did you use soft pseudo-labels or somehow make them hard pseudo-labels?",
    "1241857": "First place! Congratulations!\nQuestion.\n1:Does 6 channel take more time for trainning compared to 3 channel?　I thought it would take longer in my opinion.\n2:Since the image size is smaller, is the relative time to achieve high accuracy the same?",
    "1241883": "Thanks for sharing! Congratulations to first place!\n\nI have a question regarding segmentation stage 1: Here it says \"Pseudo label on 28k + (30k - 9k) data without tube anno\"\nI was wondering where you use these pseudo labels, because for seg. stage 2 you wrote: \"Train on 30k + 28k images (with combination of GT and pseudo labels)\" \n2nd question: Which augmentation did you use?",
    "1241928": "Congrats ! Thank you for sharing.",
    "1241986": "Just WOW ...... Congratulations.",
    "1242030": "boliu0 Wow ..Congratulations on First Place Finish . Keep up the great work",
    "1242065": "just wondering if there was any reason to use 6 channels instead of 4 channels (as input images are grayscale?)",
    "1242078": "We want to fully utilize the imagenet pretrained weights so we read\boriginal image as 3ch.",
    "1242085": "Thanks!\n\n> Does 6 channel take more time for trainning compared to 3 channel?\n\nI haven't measured it but I think the training time may have gotten longer by about 0.1% by using 6ch instead of 3ch.\n\n> Since the image size is smaller, is the relative time to achieve high accuracy the same?\n\nIn our pipeline, the cls model has no boost when using `img_size > 512`.",
    "1242107": "0.1%! That's almost the same! That was surprising.\nI'll try it myself. Thank you!",
    "1242226": "Thanks. We used soft pseudo labels for both segmentation and classification. They work better than hard ones as they contain more information.",
    "1242237": "Congratulations with the first place and thanks for sharing this great summary! Very nice way to incorporate annotations and external data in the solution.",
    "1242240": "Thanks. \n\nQ1: We have tube annotation for the 9k images, so we only need to make pseudo labels for the other 28k+30k-9k images. The goal for stage 1 is to make pseudo labels, to enable stage 2 to train on all 28k+30k data.\n\nThen in stage 2, we have labels for all 28k+30k data (among which 9k are GT labels, the rest being pseudo labels). So we train on all of them.\n\nQ2. For segmentation: HorizontalFlip, RandomBrightness, ShiftScaleRotate, Cutout\nFor classification: all the above, plus RandomContrast, OpticalDistortion, GridDistortion, HueSaturationValue",
    "1242266": "Congratulations, Lots of hard work. Thanks for sharing with us your team's great idea and solution. Can you share your codebase link? Thanks in advance.",
    "1242318": "Thank you :)",
    "1242339": "Wow! Just wowed! Congratulations!",
    "1242418": "I'm still a little confused here. As I understand (please correct me if wrong), the stored images only have single channel information, which you can then load with cv2 as 3 channel or 1 channel, where the 3 channel loading just duplicates the single channel information. It intuitively makes sense to use 3 channel as you said to fully utilize imagenet pretrained weights whenever you're training on standard images. But once you move to adding in the 3ch segmentation mask, why should there be a difference whether 4 channel (1+3) vs 6 channel (3+3).",
    "1242421": "How were the Segmentation 1 models ensembled for pseudo labels/masks? \n\nDid you just average the masks? Was any smoothing/filter used for the mask predictions? \n\nThe solution is really impressive...5 segmentation models and 67 classification models! Congrats!",
    "1242425": "Thank you, and I learn a lot from the kaggle community aswell =)\nWin-win =)",
    "1242450": "Thank you.\n\nJust simple average of soft labels, for both segmentation ensemble and classification ensemble. No weighting. No thresholding.",
    "1242457": "You're right. When we load the original image as 3 channel numpy array, the 3 channels are identical.\n\nIt's just that when we started experiments, we were using 3 channel input (without masks), then later after adding 3 channel mask, we just kept it as 3+3 channels. But I believe 1+3 channels would work equally well as you suggested.",
    "1242556": "Now I got it, thanks :)",
    "1242620": "I was wondering, since you trained more than 40(?) different model architectures,  how do you choose hyperparameters for your model? (img_size, optimizer, learning rate.. so many more). So far I've always done it \"by hand\", trying things and see if they work. I guess partly it is true for you as well, but do you also use some automated procedure? (e.g. train for a certain number of steps and pick the best combination)\n\nCongratulations btw, you reached a level of skill where Machine Learning becomes art.",
    "1242674": "Thanks.\n\nYeah we also do this by hand. The AutoCV is still under development. 😉\n\nFor optimizer, we like Adam. For image size, you can start with a small one like 256 or 384 depending on your hardware, then gradually increase to 512 or even larger after you tune your model on small images.\n\nBest learning rate varies somewhat for different model architecture, so they need to be tuned on each model. But you only need to run one fold per model for this. And it doesn't have to be exact. A slightly \"off\" learning rate for a single model won't affect ensemble's score by much.",
    "1242754": "Great thank you! And awesome solution!",
    "1243532": "Thank you for your fantastic solution. Can definitely learn a lot from it.\n\n> In our pipeline, the cls model has no boost when using img_size > 512\n\nDo you have an intuition for it?\n\nI almost felt it should be the opposite, i.e. low resolution for segmentation but high resolution for classification. \n\nAlso what downscaling method do you use to resize the segmentation output for classification, INTER_NEAREST, INTER_LINEAR, etc, and do they make any difference?\n\nMany thanks!",
    "1243778": "Before experiments we can hardly tell anything.\n\nSo we've compared (1024seg + 512cls) with (512seg + 1024cls) and the result is —— the former is far better than the latter.\n\nBased on the results of this experiment we can in turn make some explanations for the dataset. For example, segmentation requires more fine-grained detail, etc.\n\nSo sometimes don't trust your intuition. Just trust the results of the experiment.\n\nWhile, we didn't pay any attention on downscaling method, just used the default one.",
    "1243938": "Congrats! After seeing your 0.928 LB score we already knew you can win this :)",
    "1243954": "Thank you. Oh, you noticed 😏",
    "1244151": "We dont need a leaderboard because either you guys or @bestfitting @wowfattie @philippsinger and @christofhenkel would take up the money. the host can freeze you guy at the top. thats more likely",
    "1244640": "Great solution, congratulations! \nI was wondering, how do you overcome submission time limit using more than 60 models in an inference? Do you use ONNX or maybe even pruning?",
    "1245138": "Hi, thank you.\n\nNo we didn't use ONNX or pruning. We put all the models on the GPU, then use one dataloader to iterate all test data only once. We maximized the 9 hour limit by squeezing in as many models as possible without timeout.",
    "1261979": "Congratulations @boliu0",
    "1264319": "Great Job! Congrats! @boliu0 @haqishen  \n\nI wonder: For the 28k external data, (1) first run a segmentation model inference on all the 112k images to filter out all the images without tubes... How did you build this model, which dataset did you train on, or is there a one on Kaggle to do this? \n\nThanks for sharing the solution.",
    "1279718": "Hello @boliu0 Congratulations! Quite an impressive solution! \nI am rather new to AI and computer vision in particular. Is it possible to share with us the NIH data that you extracted (28k) or to share the steps needed to acquire this data from the full NIH (112k) dataset?",
    "1280057": "Hi, for some reason, Kaggle disabled attachment in forum.\n\nI uploaded the 28k images' IDs to a dataset: https://www.kaggle.com/boliu0/ranzcr-external-data-id",
    "1280060": "Hi, for this purpose, we just used a single model (Unet++, B5, 1024) trained on official annotation.",
    "1566563": "Excuse me, I want to understand whether pseudo masking was done with manual in the Segmentation Stage 1.\nCongratulation win the champion!!",
    "1571730": "I don't know if it has been answered  before, but I am confused about how can input with more than 3 channels be used in a pretrained model,\nso I checked the code,there you have done\n`self.enet.conv_stem.weight = nn.Parameter(self.enet.conv_stem.weight.repeat(1,n_ch//3+1,1,1)[:, :n_ch])`\ncan someone explain what its doing?",
    "2063518": "This is really great! Thank you for sharing the idea."
  },
  "source": "meta"
}