{
  "id": 207617,
  "title": "Avoid Overfitting Data",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/207617",
  "author_name": "SnowyOwl",
  "post_date": "2020-12-30T15:05:54.741000",
  "votes": 26,
  "comment_count": 5,
  "views": 0,
  "content": "<p>The data is easily over fitted to the training sets in competitions that require convolutional neural nets. For example, in the SIIM ISIC (<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/overview\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/overview</a>) competition, the top 850 scored over a 0.96, and the top 1100 scored over a 0.95. The final top score of the private leader board was a 0.9490, while the top score of the public leaderboard was a 0.9931. As matter of fact, the person in 880th place for the public leaderboard, ended up winning the competition.</p>\n<p>That competition was very similar to ours. A convolutional neural network analyzing biomedical images to identify problems. Problems arise when, while using this method, you run too many training epochs on your training data. For that competition, they were given an even larger training set, for a total of 88 thousand images. We are given less than half of that.</p>\n<p>With that in mind, here are a few ideas that may help avoid overfitting, while also focusing on the primary objective of identifying the catheter/line placement:</p>\n<ol>\n<li><p>Applying Different filters to the images (these are used by radiographers to better identify man-made implants/objects (such as catheters/lines)). A good article is: <a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3046467/\" target=\"_blank\">https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3046467/</a></p></li>\n<li><p>Some recommend supplementing the training set with mirrored images. I recommend against this (and there are multiple articles on this). Humans typically have a certain placement of body organs (heart on right-side), which need to be identified by the CNN for better understanding of Catheter and line placement. This may, however, work well with certain CNNs.</p></li>\n<li><p>Read articles about AI in radiography. There are many open-source coding projects, health articles (in the thousands) and recommendations to learn from. Here are some to get started that I found useful:<br>\nNIH: <a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6268174/\" target=\"_blank\">https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6268174/</a><br>\nTowards Data Science: <a href=\"https://towardsdatascience.com/medical-x-ray-%EF%B8%8F-image-classification-using-convolutional-neural-network-9a6d33b1c2a\" target=\"_blank\">https://towardsdatascience.com/medical-x-ray-%EF%B8%8F-image-classification-using-convolutional-neural-network-9a6d33b1c2a</a><br>\nIEEE: <a href=\"https://ieeexplore.ieee.org/document/8010614\" target=\"_blank\">https://ieeexplore.ieee.org/document/8010614</a><br>\nOther: <a href=\"https://pubs.rsna.org/doi/full/10.1148/radiol.2018180547\" target=\"_blank\">https://pubs.rsna.org/doi/full/10.1148/radiol.2018180547</a></p></li>\n</ol>\n<p>I hope this helps you. Let's get as high of a score as we can on this competition. Good luck!</p>",
  "messages": [
    {
      "id": 1132647,
      "postDate": "2020-12-30T15:05:54.740Z",
      "content": "<p>The data is easily over fitted to the training sets in competitions that require convolutional neural nets. For example, in the SIIM ISIC (<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/overview\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/overview</a>) competition, the top 850 scored over a 0.96, and the top 1100 scored over a 0.95. The final top score of the private leader board was a 0.9490, while the top score of the public leaderboard was a 0.9931. As matter of fact, the person in 880th place for the public leaderboard, ended up winning the competition.</p>\n<p>That competition was very similar to ours. A convolutional neural network analyzing biomedical images to identify problems. Problems arise when, while using this method, you run too many training epochs on your training data. For that competition, they were given an even larger training set, for a total of 88 thousand images. We are given less than half of that.</p>\n<p>With that in mind, here are a few ideas that may help avoid overfitting, while also focusing on the primary objective of identifying the catheter/line placement:</p>\n<ol>\n<li><p>Applying Different filters to the images (these are used by radiographers to better identify man-made implants/objects (such as catheters/lines)). A good article is: <a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3046467/\" target=\"_blank\">https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3046467/</a></p></li>\n<li><p>Some recommend supplementing the training set with mirrored images. I recommend against this (and there are multiple articles on this). Humans typically have a certain placement of body organs (heart on right-side), which need to be identified by the CNN for better understanding of Catheter and line placement. This may, however, work well with certain CNNs.</p></li>\n<li><p>Read articles about AI in radiography. There are many open-source coding projects, health articles (in the thousands) and recommendations to learn from. Here are some to get started that I found useful:<br>\nNIH: <a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6268174/\" target=\"_blank\">https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6268174/</a><br>\nTowards Data Science: <a href=\"https://towardsdatascience.com/medical-x-ray-%EF%B8%8F-image-classification-using-convolutional-neural-network-9a6d33b1c2a\" target=\"_blank\">https://towardsdatascience.com/medical-x-ray-%EF%B8%8F-image-classification-using-convolutional-neural-network-9a6d33b1c2a</a><br>\nIEEE: <a href=\"https://ieeexplore.ieee.org/document/8010614\" target=\"_blank\">https://ieeexplore.ieee.org/document/8010614</a><br>\nOther: <a href=\"https://pubs.rsna.org/doi/full/10.1148/radiol.2018180547\" target=\"_blank\">https://pubs.rsna.org/doi/full/10.1148/radiol.2018180547</a></p></li>\n</ol>\n<p>I hope this helps you. Let's get as high of a score as we can on this competition. Good luck!</p>",
      "rawMarkdown": "The data is easily over fitted to the training sets in competitions that require convolutional neural nets. For example, in the SIIM ISIC (https://www.kaggle.com/c/siim-isic-melanoma-classification/overview) competition, the top 850 scored over a 0.96, and the top 1100 scored over a 0.95. The final top score of the private leader board was a 0.9490, while the top score of the public leaderboard was a 0.9931. As matter of fact, the person in 880th place for the public leaderboard, ended up winning the competition.\n\nThat competition was very similar to ours. A convolutional neural network analyzing biomedical images to identify problems. Problems arise when, while using this method, you run too many training epochs on your training data. For that competition, they were given an even larger training set, for a total of 88 thousand images. We are given less than half of that.\n\nWith that in mind, here are a few ideas that may help avoid overfitting, while also focusing on the primary objective of identifying the catheter/line placement:\n\n1. \nApplying Different filters to the images (these are used by radiographers to better identify man-made implants/objects (such as catheters/lines)). A good article is: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3046467/\n\n2. \nSome recommend supplementing the training set with mirrored images. I recommend against this (and there are multiple articles on this). Humans typically have a certain placement of body organs (heart on right-side), which need to be identified by the CNN for better understanding of Catheter and line placement. This may, however, work well with certain CNNs.\n\n3. \nRead articles about AI in radiography. There are many open-source coding projects, health articles (in the thousands) and recommendations to learn from. Here are some to get started that I found useful:\nNIH: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6268174/\nTowards Data Science: https://towardsdatascience.com/medical-x-ray-%EF%B8%8F-image-classification-using-convolutional-neural-network-9a6d33b1c2a\nIEEE: https://ieeexplore.ieee.org/document/8010614\nOther: https://pubs.rsna.org/doi/full/10.1148/radiol.2018180547\n\nI hope this helps you. Let's get as high of a score as we can on this competition. Good luck!",
      "votes": 26
    },
    {
      "id": 1134945,
      "postDate": "2021-01-01T18:13:05.137Z",
      "content": "<p>Thank you to everyone that voted up my post!</p>",
      "rawMarkdown": "Thank you to everyone that voted up my post!",
      "votes": 1
    },
    {
      "id": 1137206,
      "postDate": "2021-01-03T18:53:44.437Z",
      "content": "<p>I applied a more advanced filtration technique, known as Contrast Limited Adaptive Histogram Equalization, which is medically used in radiography to enhance certain structures. You can see it <a href=\"https://www.kaggle.com/kyleberdy/advanced-image-preprocessing-with-clahe\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "I applied a more advanced filtration technique, known as Contrast Limited Adaptive Histogram Equalization, which is medically used in radiography to enhance certain structures. You can see it [here](https://www.kaggle.com/kyleberdy/advanced-image-preprocessing-with-clahe).",
      "votes": 2
    },
    {
      "id": 1227536,
      "postDate": "2021-03-05T16:19:55.850Z",
      "content": "<p>To check for overfitting you can also check the model on an independent validation dataset. I have created an Independent validation dataset, and as per guidelines of the competition am sharing the model to the public. Since labelling Test-set isnt allowed, so i came up with some smart ways to use other data.</p>\n<p>Link to Independent validation dataset - <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/223788\" target=\"_blank\">link</a><br>\nI hope it helps.</p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1194466/1996989/Ranzcr%20-%20Frame%206.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20210305%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20210305T161613Z&amp;X-Goog-Expires=172799&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=2180e9bff41c98ebe908e479b19d780d7b315ea183d11ffebd31415c82cbe1c45ab5b3a0d5632e7cd7a718f9bdaa969093acad050c2618058f7136f3d19337b2c0d15f66e420ace748c11e861cb2510ba62d7b3579ecef2c9985ec8c4ac20245d4a68d3647c17d3156bcf96fe21702acb8cf9a0c3c34cee956bdb82892696629e3068e4324c8b1e9c0af826240680adeea2c2842f3e732e420f8a69d35bd6818b68e5c62546388441f75b12c9e80487006f2ccead15737e4cecc28abd398b322362c4de0b50741cda4e7395f33c7fd9c14ebda88c7b177cdd8bce67f3c977a7086383bdbbcb655343d735df876d451454e0071bb1962148cbc9f5f643d8df30c\" alt=\"\"></p>",
      "rawMarkdown": "To check for overfitting you can also check the model on an independent validation dataset. I have created an Independent validation dataset, and as per guidelines of the competition am sharing the model to the public. Since labelling Test-set isnt allowed, so i came up with some smart ways to use other data.\n\nLink to Independent validation dataset - [link](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/223788)\nI hope it helps.\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1194466/1996989/Ranzcr%20-%20Frame%206.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20210305%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20210305T161613Z&X-Goog-Expires=172799&X-Goog-SignedHeaders=host&X-Goog-Signature=2180e9bff41c98ebe908e479b19d780d7b315ea183d11ffebd31415c82cbe1c45ab5b3a0d5632e7cd7a718f9bdaa969093acad050c2618058f7136f3d19337b2c0d15f66e420ace748c11e861cb2510ba62d7b3579ecef2c9985ec8c4ac20245d4a68d3647c17d3156bcf96fe21702acb8cf9a0c3c34cee956bdb82892696629e3068e4324c8b1e9c0af826240680adeea2c2842f3e732e420f8a69d35bd6818b68e5c62546388441f75b12c9e80487006f2ccead15737e4cecc28abd398b322362c4de0b50741cda4e7395f33c7fd9c14ebda88c7b177cdd8bce67f3c977a7086383bdbbcb655343d735df876d451454e0071bb1962148cbc9f5f643d8df30c)",
      "votes": 1,
      "replies": [
        {
          "id": 1270779,
          "postDate": "2021-04-12T02:13:11.627Z",
          "content": "<p>Interesting. I like that! I wish I had seen this comment earlier, as it is really helpful! Thank you so much for sharing!</p>",
          "rawMarkdown": "Interesting. I like that! I wish I had seen this comment earlier, as it is really helpful! Thank you so much for sharing!"
        }
      ]
    },
    {
      "id": 1227489,
      "postDate": "2021-03-05T15:31:12.150Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1134945,
      "author_name": "SnowyOwl",
      "author_url": "",
      "post_date": "2021-01-01T18:13:05.137000",
      "content": "<p>Thank you to everyone that voted up my post!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1137206,
      "author_name": "SnowyOwl",
      "author_url": "",
      "post_date": "2021-01-03T18:53:44.437000",
      "content": "<p>I applied a more advanced filtration technique, known as Contrast Limited Adaptive Histogram Equalization, which is medically used in radiography to enhance certain structures. You can see it <a href=\"https://www.kaggle.com/kyleberdy/advanced-image-preprocessing-with-clahe\" target=\"_blank\">here</a>.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1227536,
      "author_name": "Dr. Amritpal Singh",
      "author_url": "",
      "post_date": "2021-03-05T16:19:55.850000",
      "content": "<p>To check for overfitting you can also check the model on an independent validation dataset. I have created an Independent validation dataset, and as per guidelines of the competition am sharing the model to the public. Since labelling Test-set isnt allowed, so i came up with some smart ways to use other data.</p>\n<p>Link to Independent validation dataset - <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/223788\" target=\"_blank\">link</a><br>\nI hope it helps.</p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1194466/1996989/Ranzcr%20-%20Frame%206.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20210305%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20210305T161613Z&amp;X-Goog-Expires=172799&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=2180e9bff41c98ebe908e479b19d780d7b315ea183d11ffebd31415c82cbe1c45ab5b3a0d5632e7cd7a718f9bdaa969093acad050c2618058f7136f3d19337b2c0d15f66e420ace748c11e861cb2510ba62d7b3579ecef2c9985ec8c4ac20245d4a68d3647c17d3156bcf96fe21702acb8cf9a0c3c34cee956bdb82892696629e3068e4324c8b1e9c0af826240680adeea2c2842f3e732e420f8a69d35bd6818b68e5c62546388441f75b12c9e80487006f2ccead15737e4cecc28abd398b322362c4de0b50741cda4e7395f33c7fd9c14ebda88c7b177cdd8bce67f3c977a7086383bdbbcb655343d735df876d451454e0071bb1962148cbc9f5f643d8df30c\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1270779,
          "author_name": "SnowyOwl",
          "author_url": "",
          "post_date": "2021-04-12T02:13:11.627000",
          "content": "<p>Interesting. I like that! I wish I had seen this comment earlier, as it is really helpful! Thank you so much for sharing!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1227489,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-05T15:31:12.150000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1132647": "The data is easily over fitted to the training sets in competitions that require convolutional neural nets. For example, in the SIIM ISIC (https://www.kaggle.com/c/siim-isic-melanoma-classification/overview) competition, the top 850 scored over a 0.96, and the top 1100 scored over a 0.95. The final top score of the private leader board was a 0.9490, while the top score of the public leaderboard was a 0.9931. As matter of fact, the person in 880th place for the public leaderboard, ended up winning the competition.\n\nThat competition was very similar to ours. A convolutional neural network analyzing biomedical images to identify problems. Problems arise when, while using this method, you run too many training epochs on your training data. For that competition, they were given an even larger training set, for a total of 88 thousand images. We are given less than half of that.\n\nWith that in mind, here are a few ideas that may help avoid overfitting, while also focusing on the primary objective of identifying the catheter/line placement:\n\n1. \nApplying Different filters to the images (these are used by radiographers to better identify man-made implants/objects (such as catheters/lines)). A good article is: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3046467/\n\n2. \nSome recommend supplementing the training set with mirrored images. I recommend against this (and there are multiple articles on this). Humans typically have a certain placement of body organs (heart on right-side), which need to be identified by the CNN for better understanding of Catheter and line placement. This may, however, work well with certain CNNs.\n\n3. \nRead articles about AI in radiography. There are many open-source coding projects, health articles (in the thousands) and recommendations to learn from. Here are some to get started that I found useful:\nNIH: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6268174/\nTowards Data Science: https://towardsdatascience.com/medical-x-ray-%EF%B8%8F-image-classification-using-convolutional-neural-network-9a6d33b1c2a\nIEEE: https://ieeexplore.ieee.org/document/8010614\nOther: https://pubs.rsna.org/doi/full/10.1148/radiol.2018180547\n\nI hope this helps you. Let's get as high of a score as we can on this competition. Good luck!",
    "1134945": "Thank you to everyone that voted up my post!",
    "1137206": "I applied a more advanced filtration technique, known as Contrast Limited Adaptive Histogram Equalization, which is medically used in radiography to enhance certain structures. You can see it [here](https://www.kaggle.com/kyleberdy/advanced-image-preprocessing-with-clahe).",
    "1227536": "To check for overfitting you can also check the model on an independent validation dataset. I have created an Independent validation dataset, and as per guidelines of the competition am sharing the model to the public. Since labelling Test-set isnt allowed, so i came up with some smart ways to use other data.\n\nLink to Independent validation dataset - [link](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/223788)\nI hope it helps.\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1194466/1996989/Ranzcr%20-%20Frame%206.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20210305%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20210305T161613Z&X-Goog-Expires=172799&X-Goog-SignedHeaders=host&X-Goog-Signature=2180e9bff41c98ebe908e479b19d780d7b315ea183d11ffebd31415c82cbe1c45ab5b3a0d5632e7cd7a718f9bdaa969093acad050c2618058f7136f3d19337b2c0d15f66e420ace748c11e861cb2510ba62d7b3579ecef2c9985ec8c4ac20245d4a68d3647c17d3156bcf96fe21702acb8cf9a0c3c34cee956bdb82892696629e3068e4324c8b1e9c0af826240680adeea2c2842f3e732e420f8a69d35bd6818b68e5c62546388441f75b12c9e80487006f2ccead15737e4cecc28abd398b322362c4de0b50741cda4e7395f33c7fd9c14ebda88c7b177cdd8bce67f3c977a7086383bdbbcb655343d735df876d451454e0071bb1962148cbc9f5f643d8df30c)",
    "1227489": ""
  }
}