{
  "id": 223788,
  "title": "Independent validation - links to dataset",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/223788",
  "author_name": "Dr. Amritpal Singh",
  "post_date": "2021-03-05T15:08:56.722000",
  "votes": 0,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hii Kagglers,<br>\nI am Dr. Amrit. Currently, there is a discussion going on discussion forms of this competition regarding over-fitting on public datasets. I have curated a validation dataset for this competition which you can use to evaluate your models. <br>\nSince labelling Test-set isnt allowed, so I came up with some smart ways to use other data.</p>\n<p>Link to the dataset - <a href=\"https://www.kaggle.com/amritpal333/ranzcrindependentvalidation2\" target=\"_blank\">https://www.kaggle.com/amritpal333/ranzcrindependentvalidation2</a></p>\n<p>As per competition guidelines, I am sharing the dataset here. (I hope I am not missing anything, if I am, please let me know below).</p>\n<p>Like if this helps you.<br>\nI hope this helps the community. May the best one win.</p>\n<p>P.S. - Some parts of the self-labeled Test subset might be wrong. If you find something, Kindly write to me <a>ap4.singh@gmail.com</a>   </p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1194466/1996989/Ranzcr%20-%20Frame%206.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20210305%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20210305T161613Z&amp;X-Goog-Expires=172799&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=2180e9bff41c98ebe908e479b19d780d7b315ea183d11ffebd31415c82cbe1c45ab5b3a0d5632e7cd7a718f9bdaa969093acad050c2618058f7136f3d19337b2c0d15f66e420ace748c11e861cb2510ba62d7b3579ecef2c9985ec8c4ac20245d4a68d3647c17d3156bcf96fe21702acb8cf9a0c3c34cee956bdb82892696629e3068e4324c8b1e9c0af826240680adeea2c2842f3e732e420f8a69d35bd6818b68e5c62546388441f75b12c9e80487006f2ccead15737e4cecc28abd398b322362c4de0b50741cda4e7395f33c7fd9c14ebda88c7b177cdd8bce67f3c977a7086383bdbbcb655343d735df876d451454e0071bb1962148cbc9f5f643d8df30c\" alt=\"\"></p>",
  "messages": [
    {
      "id": 1227474,
      "postDate": "2021-03-05T15:18:27.903Z",
      "content": "<p>Since re-labeling the test set isn't allowed, we will explore three possibilities. </p>\n<p>Description of the Dataset - </p>\n<ol>\n<li><p>3_class.csv  -  Subset of Support_devices from Chexpert dataset. (contains labels for 4 classes - ETT, NGT, CVC, Swan).<br>\n                       images in folder - Chex_pert_subset</p></li>\n<li><p>Train_subset_equally_divided.csv         A reliable label by a radiologist, with equal distribution from each class. 10-15 images per class.                                         images - Train set of Ranzcr competition</p></li>\n<li><p>Far_north+south_600.csv        Far_north and Far_south are cases that the model(one of mine) was very confident(90+ for present or below 10% meaning absent), but actually model was wrong about those predictions.              images - Train set of Ranzcr competition.  To understand this Far North and Far south concept - Read the paper - <a href=\"https://pubmed.ncbi.nlm.nih.gov/31818381/\" target=\"_blank\">Link to paper</a></p></li>\n</ol>",
      "rawMarkdown": "Since re-labeling the test set isn't allowed, we will explore three possibilities. \n \nDescription of the Dataset - \n1. 3_class.csv  -  Subset of Support_devices from Chexpert dataset. (contains labels for 4 classes - ETT, NGT, CVC, Swan).\n                           images in folder - Chex_pert_subset\n\n2.  Train_subset_equally_divided.csv         A reliable label by a radiologist, with equal distribution from each class. 10-15 images per class.                                         images - Train set of Ranzcr competition\n\n3.  Far_north+south_600.csv        Far_north and Far_south are cases that the model(one of mine) was very confident(90+ for present or below 10% meaning absent), but actually model was wrong about those predictions.              images - Train set of Ranzcr competition.  To understand this Far North and Far south concept - Read the paper - [Link to paper](https://pubmed.ncbi.nlm.nih.gov/31818381/)\n\n",
      "votes": 1
    },
    {
      "id": 1227466,
      "postDate": "2021-03-05T15:08:56.723Z",
      "content": "<p>Hii Kagglers,<br>\nI am Dr. Amrit. Currently, there is a discussion going on discussion forms of this competition regarding over-fitting on public datasets. I have curated a validation dataset for this competition which you can use to evaluate your models. <br>\nSince labelling Test-set isnt allowed, so I came up with some smart ways to use other data.</p>\n<p>Link to the dataset - <a href=\"https://www.kaggle.com/amritpal333/ranzcrindependentvalidation2\" target=\"_blank\">https://www.kaggle.com/amritpal333/ranzcrindependentvalidation2</a></p>\n<p>As per competition guidelines, I am sharing the dataset here. (I hope I am not missing anything, if I am, please let me know below).</p>\n<p>Like if this helps you.<br>\nI hope this helps the community. May the best one win.</p>\n<p>P.S. - Some parts of the self-labeled Test subset might be wrong. If you find something, Kindly write to me <a>ap4.singh@gmail.com</a>   </p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1194466/1996989/Ranzcr%20-%20Frame%206.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20210305%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20210305T161613Z&amp;X-Goog-Expires=172799&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=2180e9bff41c98ebe908e479b19d780d7b315ea183d11ffebd31415c82cbe1c45ab5b3a0d5632e7cd7a718f9bdaa969093acad050c2618058f7136f3d19337b2c0d15f66e420ace748c11e861cb2510ba62d7b3579ecef2c9985ec8c4ac20245d4a68d3647c17d3156bcf96fe21702acb8cf9a0c3c34cee956bdb82892696629e3068e4324c8b1e9c0af826240680adeea2c2842f3e732e420f8a69d35bd6818b68e5c62546388441f75b12c9e80487006f2ccead15737e4cecc28abd398b322362c4de0b50741cda4e7395f33c7fd9c14ebda88c7b177cdd8bce67f3c977a7086383bdbbcb655343d735df876d451454e0071bb1962148cbc9f5f643d8df30c\" alt=\"\"></p>",
      "rawMarkdown": "Hii Kagglers,\nI am Dr. Amrit. Currently, there is a discussion going on discussion forms of this competition regarding over-fitting on public datasets. I have curated a validation dataset for this competition which you can use to evaluate your models. \nSince labelling Test-set isnt allowed, so I came up with some smart ways to use other data.\n\nLink to the dataset - https://www.kaggle.com/amritpal333/ranzcrindependentvalidation2\n\nAs per competition guidelines, I am sharing the dataset here. (I hope I am not missing anything, if I am, please let me know below).\n\nLike if this helps you.\nI hope this helps the community. May the best one win.\n\nP.S. - Some parts of the self-labeled Test subset might be wrong. If you find something, Kindly write to me ap4.singh@gmail.com   \n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1194466/1996989/Ranzcr%20-%20Frame%206.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20210305%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20210305T161613Z&X-Goog-Expires=172799&X-Goog-SignedHeaders=host&X-Goog-Signature=2180e9bff41c98ebe908e479b19d780d7b315ea183d11ffebd31415c82cbe1c45ab5b3a0d5632e7cd7a718f9bdaa969093acad050c2618058f7136f3d19337b2c0d15f66e420ace748c11e861cb2510ba62d7b3579ecef2c9985ec8c4ac20245d4a68d3647c17d3156bcf96fe21702acb8cf9a0c3c34cee956bdb82892696629e3068e4324c8b1e9c0af826240680adeea2c2842f3e732e420f8a69d35bd6818b68e5c62546388441f75b12c9e80487006f2ccead15737e4cecc28abd398b322362c4de0b50741cda4e7395f33c7fd9c14ebda88c7b177cdd8bce67f3c977a7086383bdbbcb655343d735df876d451454e0071bb1962148cbc9f5f643d8df30c)\n"
    },
    {
      "id": 1227469,
      "postDate": "2021-03-05T15:11:34.817Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1227474,
      "author_name": "Dr. Amritpal Singh",
      "author_url": "",
      "post_date": "2021-03-05T15:18:27.903000",
      "content": "<p>Since re-labeling the test set isn't allowed, we will explore three possibilities. </p>\n<p>Description of the Dataset - </p>\n<ol>\n<li><p>3_class.csv  -  Subset of Support_devices from Chexpert dataset. (contains labels for 4 classes - ETT, NGT, CVC, Swan).<br>\n                       images in folder - Chex_pert_subset</p></li>\n<li><p>Train_subset_equally_divided.csv         A reliable label by a radiologist, with equal distribution from each class. 10-15 images per class.                                         images - Train set of Ranzcr competition</p></li>\n<li><p>Far_north+south_600.csv        Far_north and Far_south are cases that the model(one of mine) was very confident(90+ for present or below 10% meaning absent), but actually model was wrong about those predictions.              images - Train set of Ranzcr competition.  To understand this Far North and Far south concept - Read the paper - <a href=\"https://pubmed.ncbi.nlm.nih.gov/31818381/\" target=\"_blank\">Link to paper</a></p></li>\n</ol>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1227469,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-05T15:11:34.817000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1227474": "Since re-labeling the test set isn't allowed, we will explore three possibilities. \n \nDescription of the Dataset - \n1. 3_class.csv  -  Subset of Support_devices from Chexpert dataset. (contains labels for 4 classes - ETT, NGT, CVC, Swan).\n                           images in folder - Chex_pert_subset\n\n2.  Train_subset_equally_divided.csv         A reliable label by a radiologist, with equal distribution from each class. 10-15 images per class.                                         images - Train set of Ranzcr competition\n\n3.  Far_north+south_600.csv        Far_north and Far_south are cases that the model(one of mine) was very confident(90+ for present or below 10% meaning absent), but actually model was wrong about those predictions.              images - Train set of Ranzcr competition.  To understand this Far North and Far south concept - Read the paper - [Link to paper](https://pubmed.ncbi.nlm.nih.gov/31818381/)\n\n",
    "1227466": "Hii Kagglers,\nI am Dr. Amrit. Currently, there is a discussion going on discussion forms of this competition regarding over-fitting on public datasets. I have curated a validation dataset for this competition which you can use to evaluate your models. \nSince labelling Test-set isnt allowed, so I came up with some smart ways to use other data.\n\nLink to the dataset - https://www.kaggle.com/amritpal333/ranzcrindependentvalidation2\n\nAs per competition guidelines, I am sharing the dataset here. (I hope I am not missing anything, if I am, please let me know below).\n\nLike if this helps you.\nI hope this helps the community. May the best one win.\n\nP.S. - Some parts of the self-labeled Test subset might be wrong. If you find something, Kindly write to me ap4.singh@gmail.com   \n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1194466/1996989/Ranzcr%20-%20Frame%206.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20210305%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20210305T161613Z&X-Goog-Expires=172799&X-Goog-SignedHeaders=host&X-Goog-Signature=2180e9bff41c98ebe908e479b19d780d7b315ea183d11ffebd31415c82cbe1c45ab5b3a0d5632e7cd7a718f9bdaa969093acad050c2618058f7136f3d19337b2c0d15f66e420ace748c11e861cb2510ba62d7b3579ecef2c9985ec8c4ac20245d4a68d3647c17d3156bcf96fe21702acb8cf9a0c3c34cee956bdb82892696629e3068e4324c8b1e9c0af826240680adeea2c2842f3e732e420f8a69d35bd6818b68e5c62546388441f75b12c9e80487006f2ccead15737e4cecc28abd398b322362c4de0b50741cda4e7395f33c7fd9c14ebda88c7b177cdd8bce67f3c977a7086383bdbbcb655343d735df876d451454e0071bb1962148cbc9f5f643d8df30c)\n",
    "1227469": ""
  }
}