{
  "id": 218564,
  "title": "Major learnings or key takeaways of this competition",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/218564",
  "author_name": "Izzy Adesanya",
  "post_date": "2021-02-11T06:57:48.186000",
  "votes": 40,
  "comment_count": 10,
  "views": 0,
  "content": "<p>There is still a long way to go for this competition but still, I have consolidated a list of things that I have learned in this competition. Hope it will be useful for everyone too.</p>\n<p>The following are some points that I have tried, have read from the discussion forum and from the public notebooks published:</p>\n<ol>\n<li><p>I tried the big models like efficientnetb7, resnest, TResNet etc. with most noteable performance getting from resnet200d model.</p></li>\n<li><p>Try with <strong>larger image sizes</strong> 500-700. It works much better than smaller image size in this competition.</p></li>\n<li><p>Some patient has 172 images and some only has 1. We want to make sure that each patient's images do not appear in multiple folds to <strong>avoid data leakage</strong>. Here's what we can do:<br>\na. Make sure that the labels are stratified<br>\nb. No patient has appeared in two folds.</p></li>\n<li><p>Try using <strong>kaggle's cache</strong> to speed up your training process.</p></li>\n<li><p><strong>Training augmentation</strong>: Light augmentation is only needed here. Simple random left-right and top-bottom flipping will do enough good.</p></li>\n<li><p><strong>Scheduling</strong>: Reduce learning rate by 10x every 3 epochs where the valid AUC did not improve</p></li>\n<li><p>Optimizer: Adam with an initial learning rate of 0.001.</p></li>\n<li><p>Try a 3-stage training process as suggested by <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> or 4-stage training as suggested by <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a>.</p></li>\n<li><p>Heat map approach as hinted by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/210312\" target=\"_blank\">here</a>.Sounds like a good approach to narrow the position to a very fine area (Maybe won't do much help in this problem but nontheless a good tool to beware of).</p></li>\n</ol>\n<p>Please feel free to comment below if I have missed out on anything in the list or wrongly stated something.<br>\nHappy Learning !! ;) </p>",
  "messages": [
    {
      "id": 1195958,
      "postDate": "2021-02-11T06:57:48.187Z",
      "content": "<p>There is still a long way to go for this competition but still, I have consolidated a list of things that I have learned in this competition. Hope it will be useful for everyone too.</p>\n<p>The following are some points that I have tried, have read from the discussion forum and from the public notebooks published:</p>\n<ol>\n<li><p>I tried the big models like efficientnetb7, resnest, TResNet etc. with most noteable performance getting from resnet200d model.</p></li>\n<li><p>Try with <strong>larger image sizes</strong> 500-700. It works much better than smaller image size in this competition.</p></li>\n<li><p>Some patient has 172 images and some only has 1. We want to make sure that each patient's images do not appear in multiple folds to <strong>avoid data leakage</strong>. Here's what we can do:<br>\na. Make sure that the labels are stratified<br>\nb. No patient has appeared in two folds.</p></li>\n<li><p>Try using <strong>kaggle's cache</strong> to speed up your training process.</p></li>\n<li><p><strong>Training augmentation</strong>: Light augmentation is only needed here. Simple random left-right and top-bottom flipping will do enough good.</p></li>\n<li><p><strong>Scheduling</strong>: Reduce learning rate by 10x every 3 epochs where the valid AUC did not improve</p></li>\n<li><p>Optimizer: Adam with an initial learning rate of 0.001.</p></li>\n<li><p>Try a 3-stage training process as suggested by <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> or 4-stage training as suggested by <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a>.</p></li>\n<li><p>Heat map approach as hinted by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/210312\" target=\"_blank\">here</a>.Sounds like a good approach to narrow the position to a very fine area (Maybe won't do much help in this problem but nontheless a good tool to beware of).</p></li>\n</ol>\n<p>Please feel free to comment below if I have missed out on anything in the list or wrongly stated something.<br>\nHappy Learning !! ;) </p>",
      "rawMarkdown": "There is still a long way to go for this competition but still, I have consolidated a list of things that I have learned in this competition. Hope it will be useful for everyone too.\n\nThe following are some points that I have tried, have read from the discussion forum and from the public notebooks published:\n\n1. I tried the big models like efficientnetb7, resnest, TResNet etc. with most noteable performance getting from resnet200d model.\n\n2. Try with **larger image sizes** 500-700. It works much better than smaller image size in this competition.\n\n3. Some patient has 172 images and some only has 1. We want to make sure that each patient's images do not appear in multiple folds to **avoid data leakage**. Here's what we can do:\na. Make sure that the labels are stratified\nb. No patient has appeared in two folds.\n\n4. Try using **kaggle's cache** to speed up your training process.\n\n5. **Training augmentation**: Light augmentation is only needed here. Simple random left-right and top-bottom flipping will do enough good.\n\n6. **Scheduling**: Reduce learning rate by 10x every 3 epochs where the valid AUC did not improve\n\n7. Optimizer: Adam with an initial learning rate of 0.001.\n\n8. Try a 3-stage training process as suggested by @yasufuminakama or 4-stage training as suggested by @ammarali32.\n\n9. Heat map approach as hinted by @hengck23 [here](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/210312).Sounds like a good approach to narrow the position to a very fine area (Maybe won't do much help in this problem but nontheless a good tool to beware of).\n\nPlease feel free to comment below if I have missed out on anything in the list or wrongly stated something.\nHappy Learning !! ;) ",
      "votes": 40
    },
    {
      "id": 1196926,
      "postDate": "2021-02-11T19:22:34.423Z",
      "content": "<p>I think these are most of the key points so far. I would add that some people have taken intereste in some external datasets like chest14. Unclear exactly what to do with extra unlabeled data, maybe pretrain on it in some way or use them as additional negative samples for all of the classes, assuming that these xrays dont have catheters in them. </p>",
      "rawMarkdown": "I think these are most of the key points so far. I would add that some people have taken intereste in some external datasets like chest14. Unclear exactly what to do with extra unlabeled data, maybe pretrain on it in some way or use them as additional negative samples for all of the classes, assuming that these xrays dont have catheters in them. ",
      "votes": 4,
      "replies": [
        {
          "id": 1199017,
          "postDate": "2021-02-13T13:42:49.477Z",
          "content": "<p>Thanks for the additional point. I myself haven't tried the external dataset yet. <br>\nI will search about it a bit more and update here. Or maybe someone who has tried it can comment.</p>",
          "rawMarkdown": "Thanks for the additional point. I myself haven't tried the external dataset yet. \nI will search about it a bit more and update here. Or maybe someone who has tried it can comment."
        }
      ]
    },
    {
      "id": 1211735,
      "postDate": "2021-02-20T13:45:21.513Z",
      "content": "<p>any idea why resnet200d is best or why it works better than other?</p>",
      "rawMarkdown": "any idea why resnet200d is best or why it works better than other?",
      "votes": 1,
      "replies": [
        {
          "id": 1213126,
          "postDate": "2021-02-21T21:12:43.697Z",
          "content": "<p>One hint can be drawn from recent paper \"ChexTransfer\" by Stanford group. Older model architectural families tend to perform better on medical imaging tasks than newer versions.</p>\n<p>Authors found newer architectural families like Efficientnet etc were fine tunned to perform better on imageNet, but that necessirly doesnt translate to model performance in medical imaging.</p>",
          "rawMarkdown": "One hint can be drawn from recent paper \"ChexTransfer\" by Stanford group. Older model architectural families tend to perform better on medical imaging tasks than newer versions.\n\nAuthors found newer architectural families like Efficientnet etc were fine tunned to perform better on imageNet, but that necessirly doesnt translate to model performance in medical imaging.",
          "votes": 2
        },
        {
          "id": 1213129,
          "postDate": "2021-02-21T21:13:25.427Z",
          "content": "<p>Link to the paper - <a href=\"https://arxiv.org/abs/2101.06871\" target=\"_blank\">https://arxiv.org/abs/2101.06871</a><br>\nCheXtransfer: Performance and Parameter Efficiency of ImageNet Models for Chest X-Ray Interpretation</p>",
          "rawMarkdown": "Link to the paper - https://arxiv.org/abs/2101.06871\nCheXtransfer: Performance and Parameter Efficiency of ImageNet Models for Chest X-Ray Interpretation",
          "votes": 2
        }
      ]
    },
    {
      "id": 1222923,
      "postDate": "2021-03-02T09:52:37.017Z",
      "content": "<p>Since it is a faint image with black background, do you recommend image Normalization ?<br>\nAfter applying image normalization it looks very dark with only small portions of the chest visible.<br>\nHave you done normalization ?</p>",
      "rawMarkdown": "Since it is a faint image with black background, do you recommend image Normalization ?\nAfter applying image normalization it looks very dark with only small portions of the chest visible.\nHave you done normalization ?"
    },
    {
      "id": 1197181,
      "postDate": "2021-02-12T02:00:03.133Z",
      "content": "<p>Nice notebook! Can you tell me what to do to “Make sure that the labels are stratified”?</p>",
      "rawMarkdown": "Nice notebook! Can you tell me what to do to “Make sure that the labels are stratified”?",
      "replies": [
        {
          "id": 1199578,
          "postDate": "2021-02-13T23:55:01.047Z",
          "content": "<p><a href=\"https://www.kaggle.com/whutddmm\" target=\"_blank\">@whutddmm</a> when you split the data into N fold, you usually like to keep the distribution of the Y across the N folds</p>\n<p>See <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedKFold.html\" target=\"_blank\">sckit-learn StratifiedKFold </a></p>\n<p>For this problem, you have also to group by Patient ID, to ensure that all the images of the same patient are in the same fold.  scikit-learn comes with a <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html\" target=\"_blank\">GroupKFold function</a> </p>\n<p>If you're interested in a little bit more, apart from both stratifying the Y and doing the group-by patient, I wrote this <a href=\"https://www.kaggle.com/virilo/ranzcr-clip-stratified-kfold-to-team-up-v3/\" target=\"_blank\">kernel</a> regarding this topic</p>",
          "rawMarkdown": "@whutddmm when you split the data into N fold, you usually like to keep the distribution of the Y across the N folds\n\nSee [sckit-learn StratifiedKFold ](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedKFold.html)\n\nFor this problem, you have also to group by Patient ID, to ensure that all the images of the same patient are in the same fold.  scikit-learn comes with a [GroupKFold function](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html) \n\nIf you're interested in a little bit more, apart from both stratifying the Y and doing the group-by patient, I wrote this [kernel](https://www.kaggle.com/virilo/ranzcr-clip-stratified-kfold-to-team-up-v3/) regarding this topic",
          "votes": 2
        },
        {
          "id": 1204913,
          "postDate": "2021-02-16T12:18:44.267Z",
          "content": "<p>Got it! Thanks for your reply.</p>",
          "rawMarkdown": "Got it! Thanks for your reply."
        }
      ]
    },
    {
      "id": 1217445,
      "postDate": "2021-02-25T04:34:35.277Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1196926,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2021-02-11T19:22:34.423000",
      "content": "<p>I think these are most of the key points so far. I would add that some people have taken intereste in some external datasets like chest14. Unclear exactly what to do with extra unlabeled data, maybe pretrain on it in some way or use them as additional negative samples for all of the classes, assuming that these xrays dont have catheters in them. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 1199017,
          "author_name": "Izzy Adesanya",
          "author_url": "",
          "post_date": "2021-02-13T13:42:49.477000",
          "content": "<p>Thanks for the additional point. I myself haven't tried the external dataset yet. <br>\nI will search about it a bit more and update here. Or maybe someone who has tried it can comment.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1211735,
      "author_name": "Yi Wu",
      "author_url": "",
      "post_date": "2021-02-20T13:45:21.513000",
      "content": "<p>any idea why resnet200d is best or why it works better than other?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1213126,
          "author_name": "Dr. Amritpal Singh",
          "author_url": "",
          "post_date": "2021-02-21T21:12:43.697000",
          "content": "<p>One hint can be drawn from recent paper \"ChexTransfer\" by Stanford group. Older model architectural families tend to perform better on medical imaging tasks than newer versions.</p>\n<p>Authors found newer architectural families like Efficientnet etc were fine tunned to perform better on imageNet, but that necessirly doesnt translate to model performance in medical imaging.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1213129,
          "author_name": "Dr. Amritpal Singh",
          "author_url": "",
          "post_date": "2021-02-21T21:13:25.427000",
          "content": "<p>Link to the paper - <a href=\"https://arxiv.org/abs/2101.06871\" target=\"_blank\">https://arxiv.org/abs/2101.06871</a><br>\nCheXtransfer: Performance and Parameter Efficiency of ImageNet Models for Chest X-Ray Interpretation</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1222923,
      "author_name": "Rishi Chandra",
      "author_url": "",
      "post_date": "2021-03-02T09:52:37.017000",
      "content": "<p>Since it is a faint image with black background, do you recommend image Normalization ?<br>\nAfter applying image normalization it looks very dark with only small portions of the chest visible.<br>\nHave you done normalization ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1197181,
      "author_name": "ddmm",
      "author_url": "",
      "post_date": "2021-02-12T02:00:03.133000",
      "content": "<p>Nice notebook! Can you tell me what to do to “Make sure that the labels are stratified”?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1199578,
          "author_name": "Virilo Tejedor Aguilera",
          "author_url": "",
          "post_date": "2021-02-13T23:55:01.047000",
          "content": "<p><a href=\"https://www.kaggle.com/whutddmm\" target=\"_blank\">@whutddmm</a> when you split the data into N fold, you usually like to keep the distribution of the Y across the N folds</p>\n<p>See <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedKFold.html\" target=\"_blank\">sckit-learn StratifiedKFold </a></p>\n<p>For this problem, you have also to group by Patient ID, to ensure that all the images of the same patient are in the same fold.  scikit-learn comes with a <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html\" target=\"_blank\">GroupKFold function</a> </p>\n<p>If you're interested in a little bit more, apart from both stratifying the Y and doing the group-by patient, I wrote this <a href=\"https://www.kaggle.com/virilo/ranzcr-clip-stratified-kfold-to-team-up-v3/\" target=\"_blank\">kernel</a> regarding this topic</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1204913,
          "author_name": "ddmm",
          "author_url": "",
          "post_date": "2021-02-16T12:18:44.267000",
          "content": "<p>Got it! Thanks for your reply.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1217445,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-25T04:34:35.277000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1195958": "There is still a long way to go for this competition but still, I have consolidated a list of things that I have learned in this competition. Hope it will be useful for everyone too.\n\nThe following are some points that I have tried, have read from the discussion forum and from the public notebooks published:\n\n1. I tried the big models like efficientnetb7, resnest, TResNet etc. with most noteable performance getting from resnet200d model.\n\n2. Try with **larger image sizes** 500-700. It works much better than smaller image size in this competition.\n\n3. Some patient has 172 images and some only has 1. We want to make sure that each patient's images do not appear in multiple folds to **avoid data leakage**. Here's what we can do:\na. Make sure that the labels are stratified\nb. No patient has appeared in two folds.\n\n4. Try using **kaggle's cache** to speed up your training process.\n\n5. **Training augmentation**: Light augmentation is only needed here. Simple random left-right and top-bottom flipping will do enough good.\n\n6. **Scheduling**: Reduce learning rate by 10x every 3 epochs where the valid AUC did not improve\n\n7. Optimizer: Adam with an initial learning rate of 0.001.\n\n8. Try a 3-stage training process as suggested by @yasufuminakama or 4-stage training as suggested by @ammarali32.\n\n9. Heat map approach as hinted by @hengck23 [here](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/210312).Sounds like a good approach to narrow the position to a very fine area (Maybe won't do much help in this problem but nontheless a good tool to beware of).\n\nPlease feel free to comment below if I have missed out on anything in the list or wrongly stated something.\nHappy Learning !! ;) ",
    "1196926": "I think these are most of the key points so far. I would add that some people have taken intereste in some external datasets like chest14. Unclear exactly what to do with extra unlabeled data, maybe pretrain on it in some way or use them as additional negative samples for all of the classes, assuming that these xrays dont have catheters in them. ",
    "1211735": "any idea why resnet200d is best or why it works better than other?",
    "1222923": "Since it is a faint image with black background, do you recommend image Normalization ?\nAfter applying image normalization it looks very dark with only small portions of the chest visible.\nHave you done normalization ?",
    "1197181": "Nice notebook! Can you tell me what to do to “Make sure that the labels are stratified”?",
    "1217445": ""
  }
}