{
  "id": 223794,
  "title": "How to make a Independent validation dataset? -Steps!",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/223794",
  "author_name": "Dr. Amritpal Singh",
  "post_date": "2021-03-05T16:00:13.769000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>There are several ways to prevent over-fitting, but to be sure that the model isn't really over-fitting, we should always test the model on an independent validation dataset.</p>\n<p>But creating a validation dataset can be difficult for several reasons. Labeling data is difficult, expensive, and requires field expertise.(at least in the case of medical images)<br>\nSo what if you are the field expert and you are creating the independent validation dataset, We still need to keep certain things in mind -</p>\n<ol>\n<li><p><strong>Class representation</strong> - Test dataset should have almost equal/or at least significant representation from each class.</p></li>\n<li><p><strong>Data distribution</strong> - To really check for generalization of the model, the test dataset should be of different distribution than training data. For example - in medical imaging - make sure you pick data from different machines, diff views(AP/PA), different geographical distribution. </p></li>\n<li><p><strong>Make a Hard Test for model</strong> - Preferable the dataset should have hand-picked difficult cases by the radiologist, that are common and the model is very likely to make false predictions; or the cases where the consequences of False predictions are too high.</p></li>\n<li><p><strong>Far North and Far South cases</strong> -This is a very interesting idea that I read in <a href=\"https://pubmed.ncbi.nlm.nih.gov/31818381/\" target=\"_blank\">this</a> paper.   <br>\nFar North (high probability false positives)<br>\nFar South (low probability false negatives) <br>\nNow the problem with this approach is that it requires an labeled dataset, which means</p></li>\n</ol>\n<ul>\n<li>You label some dataset manually</li>\n<li>or you save some data from the training dataset(Train 80, Val 15, unseen data - 5 splits)<ul>\n<li>or You can at times, run the model on train dataset(80%) and see where does the model fail in a big way.</li></ul></li>\n</ul>\n<p>An example of one such data is <a href=\"https://www.kaggle.com/amritpal333/ranzcrindependentvalidation2\" target=\"_blank\">here</a> that I created for RANZCR competition.</p>",
  "messages": [
    {
      "id": 1227515,
      "postDate": "2021-03-05T16:00:13.770Z",
      "content": "<p>There are several ways to prevent over-fitting, but to be sure that the model isn't really over-fitting, we should always test the model on an independent validation dataset.</p>\n<p>But creating a validation dataset can be difficult for several reasons. Labeling data is difficult, expensive, and requires field expertise.(at least in the case of medical images)<br>\nSo what if you are the field expert and you are creating the independent validation dataset, We still need to keep certain things in mind -</p>\n<ol>\n<li><p><strong>Class representation</strong> - Test dataset should have almost equal/or at least significant representation from each class.</p></li>\n<li><p><strong>Data distribution</strong> - To really check for generalization of the model, the test dataset should be of different distribution than training data. For example - in medical imaging - make sure you pick data from different machines, diff views(AP/PA), different geographical distribution. </p></li>\n<li><p><strong>Make a Hard Test for model</strong> - Preferable the dataset should have hand-picked difficult cases by the radiologist, that are common and the model is very likely to make false predictions; or the cases where the consequences of False predictions are too high.</p></li>\n<li><p><strong>Far North and Far South cases</strong> -This is a very interesting idea that I read in <a href=\"https://pubmed.ncbi.nlm.nih.gov/31818381/\" target=\"_blank\">this</a> paper.   <br>\nFar North (high probability false positives)<br>\nFar South (low probability false negatives) <br>\nNow the problem with this approach is that it requires an labeled dataset, which means</p></li>\n</ol>\n<ul>\n<li>You label some dataset manually</li>\n<li>or you save some data from the training dataset(Train 80, Val 15, unseen data - 5 splits)<ul>\n<li>or You can at times, run the model on train dataset(80%) and see where does the model fail in a big way.</li></ul></li>\n</ul>\n<p>An example of one such data is <a href=\"https://www.kaggle.com/amritpal333/ranzcrindependentvalidation2\" target=\"_blank\">here</a> that I created for RANZCR competition.</p>",
      "rawMarkdown": "There are several ways to prevent over-fitting, but to be sure that the model isn't really over-fitting, we should always test the model on an independent validation dataset.\n\nBut creating a validation dataset can be difficult for several reasons. Labeling data is difficult, expensive, and requires field expertise.(at least in the case of medical images)\nSo what if you are the field expert and you are creating the independent validation dataset, We still need to keep certain things in mind -\n\n1. **Class representation** - Test dataset should have almost equal/or at least significant representation from each class.\n\n2. **Data distribution** - To really check for generalization of the model, the test dataset should be of different distribution than training data. For example - in medical imaging - make sure you pick data from different machines, diff views(AP/PA), different geographical distribution. \n\n3. **Make a Hard Test for model** - Preferable the dataset should have hand-picked difficult cases by the radiologist, that are common and the model is very likely to make false predictions; or the cases where the consequences of False predictions are too high.\n\n4. **Far North and Far South cases** -This is a very interesting idea that I read in [this](https://pubmed.ncbi.nlm.nih.gov/31818381/) paper.   \nFar North (high probability false positives)\n Far South (low probability false negatives) \nNow the problem with this approach is that it requires an labeled dataset, which means\n- You label some dataset manually\n- or you save some data from the training dataset(Train 80, Val 15, unseen data - 5 splits)\n - or You can at times, run the model on train dataset(80%) and see where does the model fail in a big way.\n\nAn example of one such data is [here](https://www.kaggle.com/amritpal333/ranzcrindependentvalidation2) that I created for RANZCR competition.",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1227515": "There are several ways to prevent over-fitting, but to be sure that the model isn't really over-fitting, we should always test the model on an independent validation dataset.\n\nBut creating a validation dataset can be difficult for several reasons. Labeling data is difficult, expensive, and requires field expertise.(at least in the case of medical images)\nSo what if you are the field expert and you are creating the independent validation dataset, We still need to keep certain things in mind -\n\n1. **Class representation** - Test dataset should have almost equal/or at least significant representation from each class.\n\n2. **Data distribution** - To really check for generalization of the model, the test dataset should be of different distribution than training data. For example - in medical imaging - make sure you pick data from different machines, diff views(AP/PA), different geographical distribution. \n\n3. **Make a Hard Test for model** - Preferable the dataset should have hand-picked difficult cases by the radiologist, that are common and the model is very likely to make false predictions; or the cases where the consequences of False predictions are too high.\n\n4. **Far North and Far South cases** -This is a very interesting idea that I read in [this](https://pubmed.ncbi.nlm.nih.gov/31818381/) paper.   \nFar North (high probability false positives)\n Far South (low probability false negatives) \nNow the problem with this approach is that it requires an labeled dataset, which means\n- You label some dataset manually\n- or you save some data from the training dataset(Train 80, Val 15, unseen data - 5 splits)\n - or You can at times, run the model on train dataset(80%) and see where does the model fail in a big way.\n\nAn example of one such data is [here](https://www.kaggle.com/amritpal333/ranzcrindependentvalidation2) that I created for RANZCR competition."
  }
}