{
  "id": 230033,
  "title": "5th place solution",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/writeups/guanshuo-xu-5th-place-solution",
  "author_name": "",
  "post_date": "2021-05-18T15:45:25.423Z",
  "votes": 31,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Apologize for the late sharing, as I was fighting in another competition.<br>\nI have made figures but the attachment function is disabled, I will edit this post when the function is back.</p>\n<p><strong>Overview</strong><br>\nMy training pipeline can be broken into three steps. In step1, the model was initially trained for some dozens of epochs with all the 11 targets, we can think of this as some type of pretraining or warmup. In step2, the trained model was split into three expert models, based on the three types of lines, namely, ETT, NGT and CVC. The “Swan Ganz Catheter Present” was merged into the CVC model for simplicity. In step3, each expert model performed independent pseudo-labeling on the NIH data and finetuned several epochs. For inference, each expert model was responsible for predicting the subset of labels it was trained on. </p>\n<p><a><img src=\"https://i.ibb.co/vh5N9Hd/Picture1.png\"></a></p>\n<p><strong>Step 1 and Regularization</strong><br>\nFor better results, it’s important to regularize the training of the neural network and force it to focus on the lines and endpoints, especially if the input image size is not large enough. In my work, I made use of the train_annotations.csv as segmentation masks. To make it more economic, rather than training a complete segmentation model, I chose to downsize the segmentation mask and let the second last block of the CNN to fit on it as an auxiliary task. Since we know there is only one ETT for each image, it’s not necessary to draw the ETT in the segmentation mask, instead, for each image, I extracted the endpoint of the ETT from the annotation file as a single point with two values and add a regression task for guiding the ETT training. So, the segmentation mask only included CVC, NGT, and SWAN, and all of these lines were drawn in a single mask, and we could use a 1x1 conv layer to bridge the features maps and the mask. Later in the competition, Dr. Konya shared his “5k trachea bifurcation annotation dataset” (<a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221007)\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221007)</a>. This could be served as an additional landmark or reference points, so I added an additional regression head forcing the network to predict the trachea bifurcation. In all these regularization tasks, the losses without annotations were ignored during training. The figure below shows my CNN training in general. The names of layer1, layer2 … is more of the naming convention for resnet implementations. The sub-figure in upperleft represents the step1 training.</p>\n<p><strong>Step 2: Expert Models</strong><br>\nMy initial idea for step 2 was just to finetune the step1 model from different epochs because different line types might converge at different epochs. Later, I discovered that the performance could be further improved by only using targets of specific line types. Probably there was some interference in optimization if all the targets are trained together. So I adjusted targets as well as regularization tasks and trained three expert models. The details can be found in the figure below. I crossed out the irrelevant parts of the expert models in the figure.</p>\n<p><a><img src=\"https://i.ibb.co/vHBZMFJ/Picture2.png\"></a><br>\n<a><img src=\"https://i.ibb.co/rZKzMfr/Picture4.png\"></a><br>\n<a><img src=\"https://i.ibb.co/CBP40W9/Picture3.png\"></a></p>\n<p><strong>Step 3: Pseudo-labeling</strong><br>\nSince we know all the private test data are in the NIH data, it’s natural to make use of them by pseudo-labeling. In step3, I first excluded this competitions training data from the NIH data with imagehash, then I generated pseudo labels from each of the expert model on the leftover NIH data which should contain all the test data. Since it wouldn’t be of much use to add data with these lines of interest absent, I used summation of the relevant predictions for filtering. For example, I summed the three CVC predictions and compared it with a threshold to only accept data with CVC existed. For CVC there were around 25000 extra data filtered for pseudo-labeling, and interestingly, there were only around 2500 data left for ETT or NGT which implies a much smaller ratio compared with the test:train ratio in the competition data.</p>\n<p><strong>Models and Results</strong><br>\nI created a 6 fold split and managed to complete training models on 4 of them. The four models were three resnet200d and one efficientnet-b7 with image input size 672x672. The results are CV: 0.9712, public LB: 0.9682, private LB 0.9756. My public results were much lower than expected. It was a little frustrating during the competition as I saw many participants got a better score early and maybe easily. I checked many times my submission kernel and could not find any problem. I’m glad that I didn’t give up during this journey.</p>",
  "messages": [
    {
      "id": "1259880",
      "postDate": "04/01/2021 18:05:10",
      "content": "<p>Apologize for the late sharing, as I was fighting in another competition.<br>\nI have made figures but the attachment function is disabled, I will edit this post when the function is back.</p>\n<p><strong>Overview</strong><br>\nMy training pipeline can be broken into three steps. In step1, the model was initially trained for some dozens of epochs with all the 11 targets, we can think of this as some type of pretraining or warmup. In step2, the trained model was split into three expert models, based on the three types of lines, namely, ETT, NGT and CVC. The “Swan Ganz Catheter Present” was merged into the CVC model for simplicity. In step3, each expert model performed independent pseudo-labeling on the NIH data and finetuned several epochs. For inference, each expert model was responsible for predicting the subset of labels it was trained on. </p>\n<p><a><img src=\"https://i.ibb.co/vh5N9Hd/Picture1.png\"></a></p>\n<p><strong>Step 1 and Regularization</strong><br>\nFor better results, it’s important to regularize the training of the neural network and force it to focus on the lines and endpoints, especially if the input image size is not large enough. In my work, I made use of the train_annotations.csv as segmentation masks. To make it more economic, rather than training a complete segmentation model, I chose to downsize the segmentation mask and let the second last block of the CNN to fit on it as an auxiliary task. Since we know there is only one ETT for each image, it’s not necessary to draw the ETT in the segmentation mask, instead, for each image, I extracted the endpoint of the ETT from the annotation file as a single point with two values and add a regression task for guiding the ETT training. So, the segmentation mask only included CVC, NGT, and SWAN, and all of these lines were drawn in a single mask, and we could use a 1x1 conv layer to bridge the features maps and the mask. Later in the competition, Dr. Konya shared his “5k trachea bifurcation annotation dataset” (<a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221007)\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221007)</a>. This could be served as an additional landmark or reference points, so I added an additional regression head forcing the network to predict the trachea bifurcation. In all these regularization tasks, the losses without annotations were ignored during training. The figure below shows my CNN training in general. The names of layer1, layer2 … is more of the naming convention for resnet implementations. The sub-figure in upperleft represents the step1 training.</p>\n<p><strong>Step 2: Expert Models</strong><br>\nMy initial idea for step 2 was just to finetune the step1 model from different epochs because different line types might converge at different epochs. Later, I discovered that the performance could be further improved by only using targets of specific line types. Probably there was some interference in optimization if all the targets are trained together. So I adjusted targets as well as regularization tasks and trained three expert models. The details can be found in the figure below. I crossed out the irrelevant parts of the expert models in the figure.</p>\n<p><a><img src=\"https://i.ibb.co/vHBZMFJ/Picture2.png\"></a><br>\n<a><img src=\"https://i.ibb.co/rZKzMfr/Picture4.png\"></a><br>\n<a><img src=\"https://i.ibb.co/CBP40W9/Picture3.png\"></a></p>\n<p><strong>Step 3: Pseudo-labeling</strong><br>\nSince we know all the private test data are in the NIH data, it’s natural to make use of them by pseudo-labeling. In step3, I first excluded this competitions training data from the NIH data with imagehash, then I generated pseudo labels from each of the expert model on the leftover NIH data which should contain all the test data. Since it wouldn’t be of much use to add data with these lines of interest absent, I used summation of the relevant predictions for filtering. For example, I summed the three CVC predictions and compared it with a threshold to only accept data with CVC existed. For CVC there were around 25000 extra data filtered for pseudo-labeling, and interestingly, there were only around 2500 data left for ETT or NGT which implies a much smaller ratio compared with the test:train ratio in the competition data.</p>\n<p><strong>Models and Results</strong><br>\nI created a 6 fold split and managed to complete training models on 4 of them. The four models were three resnet200d and one efficientnet-b7 with image input size 672x672. The results are CV: 0.9712, public LB: 0.9682, private LB 0.9756. My public results were much lower than expected. It was a little frustrating during the competition as I saw many participants got a better score early and maybe easily. I checked many times my submission kernel and could not find any problem. I’m glad that I didn’t give up during this journey.</p>",
      "rawMarkdown": "Apologize for the late sharing, as I was fighting in another competition.\nI have made figures but the attachment function is disabled, I will edit this post when the function is back.\n\n**Overview**\nMy training pipeline can be broken into three steps. In step1, the model was initially trained for some dozens of epochs with all the 11 targets, we can think of this as some type of pretraining or warmup. In step2, the trained model was split into three expert models, based on the three types of lines, namely, ETT, NGT and CVC. The “Swan Ganz Catheter Present” was merged into the CVC model for simplicity. In step3, each expert model performed independent pseudo-labeling on the NIH data and finetuned several epochs. For inference, each expert model was responsible for predicting the subset of labels it was trained on. \n\n<a><img src=\"https://i.ibb.co/vh5N9Hd/Picture1.png\"></a>\n\n**Step 1 and Regularization**\nFor better results, it’s important to regularize the training of the neural network and force it to focus on the lines and endpoints, especially if the input image size is not large enough. In my work, I made use of the train_annotations.csv as segmentation masks. To make it more economic, rather than training a complete segmentation model, I chose to downsize the segmentation mask and let the second last block of the CNN to fit on it as an auxiliary task. Since we know there is only one ETT for each image, it’s not necessary to draw the ETT in the segmentation mask, instead, for each image, I extracted the endpoint of the ETT from the annotation file as a single point with two values and add a regression task for guiding the ETT training. So, the segmentation mask only included CVC, NGT, and SWAN, and all of these lines were drawn in a single mask, and we could use a 1x1 conv layer to bridge the features maps and the mask. Later in the competition, Dr. Konya shared his “5k trachea bifurcation annotation dataset” (https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221007). This could be served as an additional landmark or reference points, so I added an additional regression head forcing the network to predict the trachea bifurcation. In all these regularization tasks, the losses without annotations were ignored during training. The figure below shows my CNN training in general. The names of layer1, layer2 … is more of the naming convention for resnet implementations. The sub-figure in upperleft represents the step1 training.\n\n**Step 2: Expert Models**\nMy initial idea for step 2 was just to finetune the step1 model from different epochs because different line types might converge at different epochs. Later, I discovered that the performance could be further improved by only using targets of specific line types. Probably there was some interference in optimization if all the targets are trained together. So I adjusted targets as well as regularization tasks and trained three expert models. The details can be found in the figure below. I crossed out the irrelevant parts of the expert models in the figure.\n\n<a><img src=\"https://i.ibb.co/vHBZMFJ/Picture2.png\"></a>\n<a><img src=\"https://i.ibb.co/rZKzMfr/Picture4.png\" height=\"128\" width=\"128\"></a>\n<a><img src=\"https://i.ibb.co/CBP40W9/Picture3.png\"></a>\n\n**Step 3: Pseudo-labeling**\nSince we know all the private test data are in the NIH data, it’s natural to make use of them by pseudo-labeling. In step3, I first excluded this competitions training data from the NIH data with imagehash, then I generated pseudo labels from each of the expert model on the leftover NIH data which should contain all the test data. Since it wouldn’t be of much use to add data with these lines of interest absent, I used summation of the relevant predictions for filtering. For example, I summed the three CVC predictions and compared it with a threshold to only accept data with CVC existed. For CVC there were around 25000 extra data filtered for pseudo-labeling, and interestingly, there were only around 2500 data left for ETT or NGT which implies a much smaller ratio compared with the test:train ratio in the competition data.\n\n**Models and Results**\nI created a 6 fold split and managed to complete training models on 4 of them. The four models were three resnet200d and one efficientnet-b7 with image input size 672x672. The results are CV: 0.9712, public LB: 0.9682, private LB 0.9756. My public results were much lower than expected. It was a little frustrating during the competition as I saw many participants got a better score early and maybe easily. I checked many times my submission kernel and could not find any problem. I’m glad that I didn’t give up during this journey.",
      "votes": null
    },
    {
      "id": "1262599",
      "postDate": "04/04/2021 13:50:25",
      "content": "<p>Thanks for sharing! Waiting for the figures once the attachments are enabled again. </p>\n<p>In the final section, you mentioned how your public score didn't reflect your cross validation and indeed, your final private score ended being very good!</p>\n<p>Have you found some explanation to this discrepancy? Is your CV setup a little different to everyone else?</p>",
      "rawMarkdown": "Thanks for sharing! Waiting for the figures once the attachments are enabled again. \n\nIn the final section, you mentioned how your public score didn't reflect your cross validation and indeed, your final private score ended being very good!\n\nHave you found some explanation to this discrepancy? Is your CV setup a little different to everyone else?",
      "votes": null
    },
    {
      "id": "1262618",
      "postDate": "04/04/2021 14:18:29",
      "content": "<p>I have no idea to be honest. For CV split I simply used sklearn's kfold on patient ids.</p>",
      "rawMarkdown": "I have no idea to be honest. For CV split I simply used sklearn's kfold on patient ids.",
      "votes": null
    },
    {
      "id": "1262890",
      "postDate": "04/04/2021 20:38:27",
      "content": "<p><a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html\" target=\"_blank\">GroupKFold</a> I presume? </p>\n<p>I have used this as well but only trained one stage, I guess I should have spent more time on this :D</p>",
      "rawMarkdown": "[GroupKFold](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html) I presume? \n\nI have used this as well but only trained one stage, I guess I should have spent more time on this :D",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1262599,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "04/04/2021 13:50:25",
      "content": "<p>Thanks for sharing! Waiting for the figures once the attachments are enabled again. </p>\n<p>In the final section, you mentioned how your public score didn't reflect your cross validation and indeed, your final private score ended being very good!</p>\n<p>Have you found some explanation to this discrepancy? Is your CV setup a little different to everyone else?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1262618,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "04/04/2021 14:18:29",
          "content": "<p>I have no idea to be honest. For CV split I simply used sklearn's kfold on patient ids.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1262890,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "04/04/2021 20:38:27",
          "content": "<p><a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html\" target=\"_blank\">GroupKFold</a> I presume? </p>\n<p>I have used this as well but only trained one stage, I guess I should have spent more time on this :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1259880": "Apologize for the late sharing, as I was fighting in another competition.\nI have made figures but the attachment function is disabled, I will edit this post when the function is back.\n\n**Overview**\nMy training pipeline can be broken into three steps. In step1, the model was initially trained for some dozens of epochs with all the 11 targets, we can think of this as some type of pretraining or warmup. In step2, the trained model was split into three expert models, based on the three types of lines, namely, ETT, NGT and CVC. The “Swan Ganz Catheter Present” was merged into the CVC model for simplicity. In step3, each expert model performed independent pseudo-labeling on the NIH data and finetuned several epochs. For inference, each expert model was responsible for predicting the subset of labels it was trained on. \n\n<a><img src=\"https://i.ibb.co/vh5N9Hd/Picture1.png\"></a>\n\n**Step 1 and Regularization**\nFor better results, it’s important to regularize the training of the neural network and force it to focus on the lines and endpoints, especially if the input image size is not large enough. In my work, I made use of the train_annotations.csv as segmentation masks. To make it more economic, rather than training a complete segmentation model, I chose to downsize the segmentation mask and let the second last block of the CNN to fit on it as an auxiliary task. Since we know there is only one ETT for each image, it’s not necessary to draw the ETT in the segmentation mask, instead, for each image, I extracted the endpoint of the ETT from the annotation file as a single point with two values and add a regression task for guiding the ETT training. So, the segmentation mask only included CVC, NGT, and SWAN, and all of these lines were drawn in a single mask, and we could use a 1x1 conv layer to bridge the features maps and the mask. Later in the competition, Dr. Konya shared his “5k trachea bifurcation annotation dataset” (https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221007). This could be served as an additional landmark or reference points, so I added an additional regression head forcing the network to predict the trachea bifurcation. In all these regularization tasks, the losses without annotations were ignored during training. The figure below shows my CNN training in general. The names of layer1, layer2 … is more of the naming convention for resnet implementations. The sub-figure in upperleft represents the step1 training.\n\n**Step 2: Expert Models**\nMy initial idea for step 2 was just to finetune the step1 model from different epochs because different line types might converge at different epochs. Later, I discovered that the performance could be further improved by only using targets of specific line types. Probably there was some interference in optimization if all the targets are trained together. So I adjusted targets as well as regularization tasks and trained three expert models. The details can be found in the figure below. I crossed out the irrelevant parts of the expert models in the figure.\n\n<a><img src=\"https://i.ibb.co/vHBZMFJ/Picture2.png\"></a>\n<a><img src=\"https://i.ibb.co/rZKzMfr/Picture4.png\" height=\"128\" width=\"128\"></a>\n<a><img src=\"https://i.ibb.co/CBP40W9/Picture3.png\"></a>\n\n**Step 3: Pseudo-labeling**\nSince we know all the private test data are in the NIH data, it’s natural to make use of them by pseudo-labeling. In step3, I first excluded this competitions training data from the NIH data with imagehash, then I generated pseudo labels from each of the expert model on the leftover NIH data which should contain all the test data. Since it wouldn’t be of much use to add data with these lines of interest absent, I used summation of the relevant predictions for filtering. For example, I summed the three CVC predictions and compared it with a threshold to only accept data with CVC existed. For CVC there were around 25000 extra data filtered for pseudo-labeling, and interestingly, there were only around 2500 data left for ETT or NGT which implies a much smaller ratio compared with the test:train ratio in the competition data.\n\n**Models and Results**\nI created a 6 fold split and managed to complete training models on 4 of them. The four models were three resnet200d and one efficientnet-b7 with image input size 672x672. The results are CV: 0.9712, public LB: 0.9682, private LB 0.9756. My public results were much lower than expected. It was a little frustrating during the competition as I saw many participants got a better score early and maybe easily. I checked many times my submission kernel and could not find any problem. I’m glad that I didn’t give up during this journey.",
    "1262599": "Thanks for sharing! Waiting for the figures once the attachments are enabled again. \n\nIn the final section, you mentioned how your public score didn't reflect your cross validation and indeed, your final private score ended being very good!\n\nHave you found some explanation to this discrepancy? Is your CV setup a little different to everyone else?",
    "1262618": "I have no idea to be honest. For CV split I simply used sklearn's kfold on patient ids.",
    "1262890": "[GroupKFold](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html) I presume? \n\nI have used this as well but only trained one stage, I guess I should have spent more time on this :D"
  },
  "source": "meta"
}