{
  "id": 202651,
  "title": "Create TF Record + Stratified Sampling",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/202651",
  "author_name": "Aditya Baurai",
  "post_date": "2020-12-11T08:07:50.117000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p>As we all have noticed that there is a significant bias in the class label distribution, in our training dataset :</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2599569%2Fc8de93dc2e4eea126d3654144aa82113%2Flabellll.png?generation=1607673525163097&amp;alt=media\" alt=\"\"></p>\n<p>Hence, sampling our dataset for validation becomes important(<em>there isn't a validation set provided</em>), and a random sample won't do effectively because of the aforementioned bias.</p>\n<p><strong>Enter : Stratified Sampling</strong></p>\n<p>This method, which is a form of random sampling, consists of dividing the entire population being studied into different subgroups(<strong>here, according to the classes of diseases</strong>), so that an individual can belong to only one class(the singular). Once the groups have been defined, in order to create a sample, we select individuals by applying a sampling method to each of the groups separately.</p>\n<p><strong>This notebook samples the data using stratified sampling, followed by the creation of TRAINING, VALIDATION &amp; TESTING TFRECORDS</strong>. : </p>\n<p><strong><a href=\"https://www.kaggle.com/fireheart7/cassava-create-tfrecords\" target=\"_blank\">https://www.kaggle.com/fireheart7/cassava-create-tfrecords</a></strong></p>\n<p>A 256 X 256 record prepared(without any image preprocessing) is available at : </p>\n<p><strong><a href=\"https://www.kaggle.com/fireheart7/cassava-tfrecords-version-neptune\" target=\"_blank\">https://www.kaggle.com/fireheart7/cassava-tfrecords-version-neptune</a></strong> </p>",
  "messages": [
    {
      "id": 1108981,
      "postDate": "2020-12-11T08:07:50.117Z",
      "content": "<p>As we all have noticed that there is a significant bias in the class label distribution, in our training dataset :</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2599569%2Fc8de93dc2e4eea126d3654144aa82113%2Flabellll.png?generation=1607673525163097&amp;alt=media\" alt=\"\"></p>\n<p>Hence, sampling our dataset for validation becomes important(<em>there isn't a validation set provided</em>), and a random sample won't do effectively because of the aforementioned bias.</p>\n<p><strong>Enter : Stratified Sampling</strong></p>\n<p>This method, which is a form of random sampling, consists of dividing the entire population being studied into different subgroups(<strong>here, according to the classes of diseases</strong>), so that an individual can belong to only one class(the singular). Once the groups have been defined, in order to create a sample, we select individuals by applying a sampling method to each of the groups separately.</p>\n<p><strong>This notebook samples the data using stratified sampling, followed by the creation of TRAINING, VALIDATION &amp; TESTING TFRECORDS</strong>. : </p>\n<p><strong><a href=\"https://www.kaggle.com/fireheart7/cassava-create-tfrecords\" target=\"_blank\">https://www.kaggle.com/fireheart7/cassava-create-tfrecords</a></strong></p>\n<p>A 256 X 256 record prepared(without any image preprocessing) is available at : </p>\n<p><strong><a href=\"https://www.kaggle.com/fireheart7/cassava-tfrecords-version-neptune\" target=\"_blank\">https://www.kaggle.com/fireheart7/cassava-tfrecords-version-neptune</a></strong> </p>",
      "rawMarkdown": "As we all have noticed that there is a significant bias in the class label distribution, in our training dataset :\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2599569%2Fc8de93dc2e4eea126d3654144aa82113%2Flabellll.png?generation=1607673525163097&alt=media)\n\nHence, sampling our dataset for validation becomes important(*there isn't a validation set provided*), and a random sample won't do effectively because of the aforementioned bias.\n\n**Enter : Stratified Sampling**\n\nThis method, which is a form of random sampling, consists of dividing the entire population being studied into different subgroups(**here, according to the classes of diseases**), so that an individual can belong to only one class(the singular). Once the groups have been defined, in order to create a sample, we select individuals by applying a sampling method to each of the groups separately.\n\n**This notebook samples the data using stratified sampling, followed by the creation of TRAINING, VALIDATION & TESTING TFRECORDS**. : \n\n**https://www.kaggle.com/fireheart7/cassava-create-tfrecords**\n\nA 256 X 256 record prepared(without any image preprocessing) is available at : \n\n**https://www.kaggle.com/fireheart7/cassava-tfrecords-version-neptune** ",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1108981": "As we all have noticed that there is a significant bias in the class label distribution, in our training dataset :\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2599569%2Fc8de93dc2e4eea126d3654144aa82113%2Flabellll.png?generation=1607673525163097&alt=media)\n\nHence, sampling our dataset for validation becomes important(*there isn't a validation set provided*), and a random sample won't do effectively because of the aforementioned bias.\n\n**Enter : Stratified Sampling**\n\nThis method, which is a form of random sampling, consists of dividing the entire population being studied into different subgroups(**here, according to the classes of diseases**), so that an individual can belong to only one class(the singular). Once the groups have been defined, in order to create a sample, we select individuals by applying a sampling method to each of the groups separately.\n\n**This notebook samples the data using stratified sampling, followed by the creation of TRAINING, VALIDATION & TESTING TFRECORDS**. : \n\n**https://www.kaggle.com/fireheart7/cassava-create-tfrecords**\n\nA 256 X 256 record prepared(without any image preprocessing) is available at : \n\n**https://www.kaggle.com/fireheart7/cassava-tfrecords-version-neptune** "
  }
}