{
  "id": 203518,
  "title": "Summary of My Quick EDA Results",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/203518",
  "author_name": "Tolga",
  "post_date": "2020-12-15T15:29:38.207000",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I would like to share a <a href=\"https://www.kaggle.com/tolgadincer/ranzcr-clip-eda\" target=\"_blank\">notebook</a> in which I show my quick EDA results. A summary of my findings is as follows:</p>\n<ul>\n<li>This competition gives 30083 X-ray images and asks for the probabilities of 11 targets. The sample size is very small compared to the imagenet. So, transfer learning seems to be the way.</li>\n<li>The targets are observed in groups of 0, 1, 2, 3, 4, 5, 6. So this is a multilabel classification problem. The dataset includes 24 images that are not associated with any of the labels.</li>\n<li>The data is extremely imbalanced. The most frequent target, CVC-Normal, is observed in 71% of the data.</li>\n<li>Some patients are present in the data more than once - the most frequent patient is observed as much as 172 times. So, stratified GroupKFold splitting may be a good strategy for cross validation. <a href=\"https://www.kaggle.com/tolgadincer/iter-strat\" target=\"_blank\">Iterative-Stratification</a> does the stratification part for the multilabel data.</li>\n<li>Only part of the data has tube annotations. Majority of the images with tube annotations have 1 annotated tube. However, tube annotations are observed as much as 6 times!</li>\n<li>Annotation lengths of the tubes vary between 4 and 152.</li>\n</ul>\n<p>Furthermore,</p>\n<ul>\n<li>The competition metric is AUC. </li>\n</ul>\n<p>to be continued…</p>",
  "messages": [
    {
      "id": 1113617,
      "postDate": "2020-12-15T15:29:38.207Z",
      "content": "<p>I would like to share a <a href=\"https://www.kaggle.com/tolgadincer/ranzcr-clip-eda\" target=\"_blank\">notebook</a> in which I show my quick EDA results. A summary of my findings is as follows:</p>\n<ul>\n<li>This competition gives 30083 X-ray images and asks for the probabilities of 11 targets. The sample size is very small compared to the imagenet. So, transfer learning seems to be the way.</li>\n<li>The targets are observed in groups of 0, 1, 2, 3, 4, 5, 6. So this is a multilabel classification problem. The dataset includes 24 images that are not associated with any of the labels.</li>\n<li>The data is extremely imbalanced. The most frequent target, CVC-Normal, is observed in 71% of the data.</li>\n<li>Some patients are present in the data more than once - the most frequent patient is observed as much as 172 times. So, stratified GroupKFold splitting may be a good strategy for cross validation. <a href=\"https://www.kaggle.com/tolgadincer/iter-strat\" target=\"_blank\">Iterative-Stratification</a> does the stratification part for the multilabel data.</li>\n<li>Only part of the data has tube annotations. Majority of the images with tube annotations have 1 annotated tube. However, tube annotations are observed as much as 6 times!</li>\n<li>Annotation lengths of the tubes vary between 4 and 152.</li>\n</ul>\n<p>Furthermore,</p>\n<ul>\n<li>The competition metric is AUC. </li>\n</ul>\n<p>to be continued…</p>",
      "rawMarkdown": "I would like to share a [notebook](https://www.kaggle.com/tolgadincer/ranzcr-clip-eda) in which I show my quick EDA results. A summary of my findings is as follows:\n\n- This competition gives 30083 X-ray images and asks for the probabilities of 11 targets. The sample size is very small compared to the imagenet. So, transfer learning seems to be the way.\n- The targets are observed in groups of 0, 1, 2, 3, 4, 5, 6. So this is a multilabel classification problem. The dataset includes 24 images that are not associated with any of the labels.\n- The data is extremely imbalanced. The most frequent target, CVC-Normal, is observed in 71% of the data.\n- Some patients are present in the data more than once - the most frequent patient is observed as much as 172 times. So, stratified GroupKFold splitting may be a good strategy for cross validation. [Iterative-Stratification](https://www.kaggle.com/tolgadincer/iter-strat) does the stratification part for the multilabel data.\n- Only part of the data has tube annotations. Majority of the images with tube annotations have 1 annotated tube. However, tube annotations are observed as much as 6 times!\n- Annotation lengths of the tubes vary between 4 and 152.\n\nFurthermore,\n- The competition metric is AUC. \n\nto be continued...",
      "votes": 2
    },
    {
      "id": 1113734,
      "postDate": "2020-12-15T16:45:59.510Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 1113778,
          "postDate": "2020-12-15T17:21:03.917Z",
          "content": "<p>I guess you are concerned about the normal-abnormal combinations of the same tubes. These cases can be problematic if the image does not exhibit two tubes of the same kind. </p>\n<p>In my notebook, I see that the abnormal-normal combinations are present only for CVC tubes. I'm not an expert but I think an image can have multiple CVC tubes as it's plugged to the body below the collarbone. However, this shouldn't be the case for ETT and NGT as the former goes in from the mouth and the latter from the nose. </p>\n<p>Edit: I did some annotated plots. It seems like there are multiple tubes of the same kind with normal and abnormal conditions. One of the blue tubes in the plot below is normal and the other is abnormal.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F337744%2Fee32a5c155799d3fa29bbdc1da4080e7%2FScreen%20Shot%202020-12-15%20at%208.55.06%20PM.png?generation=1608084676817415&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "I guess you are concerned about the normal-abnormal combinations of the same tubes. These cases can be problematic if the image does not exhibit two tubes of the same kind. \n\nIn my notebook, I see that the abnormal-normal combinations are present only for CVC tubes. I'm not an expert but I think an image can have multiple CVC tubes as it's plugged to the body below the collarbone. However, this shouldn't be the case for ETT and NGT as the former goes in from the mouth and the latter from the nose. \n\nEdit: I did some annotated plots. It seems like there are multiple tubes of the same kind with normal and abnormal conditions. One of the blue tubes in the plot below is normal and the other is abnormal.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F337744%2Fee32a5c155799d3fa29bbdc1da4080e7%2FScreen%20Shot%202020-12-15%20at%208.55.06%20PM.png?generation=1608084676817415&alt=media =300x300)",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1113734,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-15T16:45:59.510000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1113778,
          "author_name": "Tolga",
          "author_url": "",
          "post_date": "2020-12-15T17:21:03.917000",
          "content": "<p>I guess you are concerned about the normal-abnormal combinations of the same tubes. These cases can be problematic if the image does not exhibit two tubes of the same kind. </p>\n<p>In my notebook, I see that the abnormal-normal combinations are present only for CVC tubes. I'm not an expert but I think an image can have multiple CVC tubes as it's plugged to the body below the collarbone. However, this shouldn't be the case for ETT and NGT as the former goes in from the mouth and the latter from the nose. </p>\n<p>Edit: I did some annotated plots. It seems like there are multiple tubes of the same kind with normal and abnormal conditions. One of the blue tubes in the plot below is normal and the other is abnormal.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F337744%2Fee32a5c155799d3fa29bbdc1da4080e7%2FScreen%20Shot%202020-12-15%20at%208.55.06%20PM.png?generation=1608084676817415&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1113617": "I would like to share a [notebook](https://www.kaggle.com/tolgadincer/ranzcr-clip-eda) in which I show my quick EDA results. A summary of my findings is as follows:\n\n- This competition gives 30083 X-ray images and asks for the probabilities of 11 targets. The sample size is very small compared to the imagenet. So, transfer learning seems to be the way.\n- The targets are observed in groups of 0, 1, 2, 3, 4, 5, 6. So this is a multilabel classification problem. The dataset includes 24 images that are not associated with any of the labels.\n- The data is extremely imbalanced. The most frequent target, CVC-Normal, is observed in 71% of the data.\n- Some patients are present in the data more than once - the most frequent patient is observed as much as 172 times. So, stratified GroupKFold splitting may be a good strategy for cross validation. [Iterative-Stratification](https://www.kaggle.com/tolgadincer/iter-strat) does the stratification part for the multilabel data.\n- Only part of the data has tube annotations. Majority of the images with tube annotations have 1 annotated tube. However, tube annotations are observed as much as 6 times!\n- Annotation lengths of the tubes vary between 4 and 152.\n\nFurthermore,\n- The competition metric is AUC. \n\nto be continued...",
    "1113734": ""
  }
}