{
  "id": 222630,
  "title": "Making sense from annotations + Starter notebook",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/222630",
  "author_name": "",
  "post_date": "2021-02-28T10:34:31.968239100Z",
  "votes": 14,
  "comment_count": 1,
  "views": 0,
  "content": "<h3>Hello!</h3>\n<p>There's a lot of topics and comments on annotations. Some users claim them to be useful for the teacher-student approach (<strong><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243\" target=\"_blank\">topic</a></strong> by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, <strong><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577\" target=\"_blank\">another topic</a></strong> with a detailed strategy and notebook links by <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>) or for segmentation approach (<strong><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/212698\" target=\"_blank\">topic</a></strong> by <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a>) while some others claim that there are no relevant improvements from these approaches for relatively big models. Some of the annotations highlighted <strong><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/210064\" target=\"_blank\">here</a></strong> are probably mislabeled. </p>\n<p>I've done a quick EDA showing that out of 9095 samples with annotations available:</p>\n<ul>\n<li>7723 samples are completely annotated i.e. have annotations for all the catheters inserted</li>\n<li>only 24 samples have incomplete annotations</li>\n<li>1349 samples have more than one catheter of the same <strong>class</strong> (not just the same <strong>type</strong>)</li>\n</ul>\n<p>Regarding the last one, it's been a <strong><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221513\" target=\"_blank\">recent topic</a></strong> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> wondering why e.g. <code>CVC - Normal</code> and <code>CVC - Abnormal</code> can be positive for the same sample (the answer is - there can be multiple CVC inserted for different purposes). It's now not so wondering to spot samples where multiple catheters have the same class, e.g.:<br>\n<img src=\"https://drive.google.com/uc?id=1uetQgQkrdSa3qVlGBarVTsbFo3vX-fjv\"></p>\n<p>Label distribution across annotated samples seems to be fine, e.g. we have only 79 <code>ETT - Abnormal</code> cases and 40 of them do have annotations.<br>\n<img src=\"https://drive.google.com/uc?id=1xmhFFRJ6K3XGfDplKAaPuS-vPiUNhnAr\"></p>\n<p>Lastly, I've made this <strong><a href=\"https://www.kaggle.com/nickuzmenkov/ranzcr-clip-making-sense-from-annotations\" target=\"_blank\">starter notebook</a></strong> for preparing dataset with images annotated via <strong>OpenCV</strong> methods in <code>.jpeg</code> and <code>.tfrec</code> format. </p>\n<p>Hope this helps someone. Happy coding!</p>",
  "messages": [
    {
      "id": "1220761",
      "postDate": "02/28/2021 10:34:31",
      "content": "<h3>Hello!</h3>\n<p>There's a lot of topics and comments on annotations. Some users claim them to be useful for the teacher-student approach (<strong><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243\" target=\"_blank\">topic</a></strong> by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, <strong><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577\" target=\"_blank\">another topic</a></strong> with a detailed strategy and notebook links by <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>) or for segmentation approach (<strong><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/212698\" target=\"_blank\">topic</a></strong> by <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a>) while some others claim that there are no relevant improvements from these approaches for relatively big models. Some of the annotations highlighted <strong><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/210064\" target=\"_blank\">here</a></strong> are probably mislabeled. </p>\n<p>I've done a quick EDA showing that out of 9095 samples with annotations available:</p>\n<ul>\n<li>7723 samples are completely annotated i.e. have annotations for all the catheters inserted</li>\n<li>only 24 samples have incomplete annotations</li>\n<li>1349 samples have more than one catheter of the same <strong>class</strong> (not just the same <strong>type</strong>)</li>\n</ul>\n<p>Regarding the last one, it's been a <strong><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221513\" target=\"_blank\">recent topic</a></strong> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> wondering why e.g. <code>CVC - Normal</code> and <code>CVC - Abnormal</code> can be positive for the same sample (the answer is - there can be multiple CVC inserted for different purposes). It's now not so wondering to spot samples where multiple catheters have the same class, e.g.:<br>\n<img src=\"https://drive.google.com/uc?id=1uetQgQkrdSa3qVlGBarVTsbFo3vX-fjv\"></p>\n<p>Label distribution across annotated samples seems to be fine, e.g. we have only 79 <code>ETT - Abnormal</code> cases and 40 of them do have annotations.<br>\n<img src=\"https://drive.google.com/uc?id=1xmhFFRJ6K3XGfDplKAaPuS-vPiUNhnAr\"></p>\n<p>Lastly, I've made this <strong><a href=\"https://www.kaggle.com/nickuzmenkov/ranzcr-clip-making-sense-from-annotations\" target=\"_blank\">starter notebook</a></strong> for preparing dataset with images annotated via <strong>OpenCV</strong> methods in <code>.jpeg</code> and <code>.tfrec</code> format. </p>\n<p>Hope this helps someone. Happy coding!</p>",
      "rawMarkdown": "### Hello!\n\nThere's a lot of topics and comments on annotations. Some users claim them to be useful for the teacher-student approach (**[topic](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243)** by @hengck23, **[another topic](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577)** with a detailed strategy and notebook links by @yasufuminakama) or for segmentation approach (**[topic](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/212698)** by @ryches) while some others claim that there are no relevant improvements from these approaches for relatively big models. Some of the annotations highlighted **[here](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/210064)** are probably mislabeled. \n\nI've done a quick EDA showing that out of 9095 samples with annotations available:\n* 7723 samples are completely annotated i.e. have annotations for all the catheters inserted\n* only 24 samples have incomplete annotations\n* 1349 samples have more than one catheter of the same **class** (not just the same **type**)\n\nRegarding the last one, it's been a **[recent topic](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221513)** by @cdeotte wondering why e.g. `CVC - Normal` and `CVC - Abnormal` can be positive for the same sample (the answer is - there can be multiple CVC inserted for different purposes). It's now not so wondering to spot samples where multiple catheters have the same class, e.g.:\n<img src='https://drive.google.com/uc?id=1uetQgQkrdSa3qVlGBarVTsbFo3vX-fjv' />\n\nLabel distribution across annotated samples seems to be fine, e.g. we have only 79 `ETT - Abnormal` cases and 40 of them do have annotations.\n<img src='https://drive.google.com/uc?id=1xmhFFRJ6K3XGfDplKAaPuS-vPiUNhnAr' />\n\nLastly, I've made this **[starter notebook](https://www.kaggle.com/nickuzmenkov/ranzcr-clip-making-sense-from-annotations)** for preparing dataset with images annotated via **OpenCV** methods in `.jpeg` and `.tfrec` format. \n\nHope this helps someone. Happy coding!",
      "votes": null
    },
    {
      "id": "1221008",
      "postDate": "02/28/2021 15:42:02",
      "content": "<p>One interesting study is looking at the relationship between the number of catheters of a given type and loss. These multi catheter scenarios appear to be hard for the model</p>",
      "rawMarkdown": "One interesting study is looking at the relationship between the number of catheters of a given type and loss. These multi catheter scenarios appear to be hard for the model",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1221008,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "02/28/2021 15:42:02",
      "content": "<p>One interesting study is looking at the relationship between the number of catheters of a given type and loss. These multi catheter scenarios appear to be hard for the model</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1220761": "### Hello!\n\nThere's a lot of topics and comments on annotations. Some users claim them to be useful for the teacher-student approach (**[topic](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243)** by @hengck23, **[another topic](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577)** with a detailed strategy and notebook links by @yasufuminakama) or for segmentation approach (**[topic](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/212698)** by @ryches) while some others claim that there are no relevant improvements from these approaches for relatively big models. Some of the annotations highlighted **[here](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/210064)** are probably mislabeled. \n\nI've done a quick EDA showing that out of 9095 samples with annotations available:\n* 7723 samples are completely annotated i.e. have annotations for all the catheters inserted\n* only 24 samples have incomplete annotations\n* 1349 samples have more than one catheter of the same **class** (not just the same **type**)\n\nRegarding the last one, it's been a **[recent topic](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221513)** by @cdeotte wondering why e.g. `CVC - Normal` and `CVC - Abnormal` can be positive for the same sample (the answer is - there can be multiple CVC inserted for different purposes). It's now not so wondering to spot samples where multiple catheters have the same class, e.g.:\n<img src='https://drive.google.com/uc?id=1uetQgQkrdSa3qVlGBarVTsbFo3vX-fjv' />\n\nLabel distribution across annotated samples seems to be fine, e.g. we have only 79 `ETT - Abnormal` cases and 40 of them do have annotations.\n<img src='https://drive.google.com/uc?id=1xmhFFRJ6K3XGfDplKAaPuS-vPiUNhnAr' />\n\nLastly, I've made this **[starter notebook](https://www.kaggle.com/nickuzmenkov/ranzcr-clip-making-sense-from-annotations)** for preparing dataset with images annotated via **OpenCV** methods in `.jpeg` and `.tfrec` format. \n\nHope this helps someone. Happy coding!",
    "1221008": "One interesting study is looking at the relationship between the number of catheters of a given type and loss. These multi catheter scenarios appear to be hard for the model"
  },
  "source": "meta"
}