{
  "id": 203500,
  "title": "MEGA TFRec Dataset - Cassava 256 X 256",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/203500",
  "author_name": "",
  "post_date": "2020-12-15T14:42:34.256038900Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi all!</p>\n<p>I prepared this dataset few days back. For each training image, I have applied random augmentations such as coarse dropout, mirror flip, and so on.. and cooked this mega TFRec dataset. </p>\n<p>As we very well know, low bias problems can always be resolved by increasing the size of our dataset. As I tried it myself on very few starting models, I found the validation and training accuracy increasing by 3%. </p>\n<p>This contains : </p>\n<ul>\n<li><strong>180,000+ training images - .tfrec format</strong></li>\n<li><strong>3000+ validation images - .tfrec format</strong></li>\n<li><strong>1 test image - the one provided by the competition itself</strong></li>\n</ul>\n<p>All images ae rescaled to 256 X 256 format for use. Also, the structure of .tfrec files during creation was :</p>\n<hr>\n<p><strong>Training &amp; Validation</strong></p>\n<p>feature = {<br>\n        \"image\" : bytes_features(feature_list[0]),<br>\n        \"image_name\" : bytes_features(feature_list[1]),<br>\n        \"label\" : int64_features(feature_list[2])<br>\n    }</p>\n<hr>\n<p><strong>Testing</strong> : <br>\nfeature = {<br>\n        \"image\" : bytes_features(feature_list[0]),<br>\n        \"image_name\" : bytes_features(feature_list[1])<br>\n    }</p>\n<hr>\n<p>(<em>Note the feature name is <strong>label</strong> in order to access it. I think in the official tfrecords of the competition it was class or something.</em>)</p>\n<h1><strong>LINK</strong> :</h1>\n<h2><strong><a href=\"https://www.kaggle.com/fireheart7/augmented-cassava-tfrec\" target=\"_blank\">https://www.kaggle.com/fireheart7/augmented-cassava-tfrec</a></strong></h2>\n<h1>Guide(Quite Standard) Book With This Dataset :</h1>\n<h2><strong><a href=\"https://www.kaggle.com/fireheart7/guide-mega-tfrec-dataset?scriptVersionId=49407011\" target=\"_blank\">https://www.kaggle.com/fireheart7/guide-mega-tfrec-dataset?scriptVersionId=49407011</a></strong></h2>\n<p>Hope this will be of aid in this competition! All the very best!</p>",
  "messages": [
    {
      "id": "1113547",
      "postDate": "12/15/2020 14:42:34",
      "content": "<p>Hi all!</p>\n<p>I prepared this dataset few days back. For each training image, I have applied random augmentations such as coarse dropout, mirror flip, and so on.. and cooked this mega TFRec dataset. </p>\n<p>As we very well know, low bias problems can always be resolved by increasing the size of our dataset. As I tried it myself on very few starting models, I found the validation and training accuracy increasing by 3%. </p>\n<p>This contains : </p>\n<ul>\n<li><strong>180,000+ training images - .tfrec format</strong></li>\n<li><strong>3000+ validation images - .tfrec format</strong></li>\n<li><strong>1 test image - the one provided by the competition itself</strong></li>\n</ul>\n<p>All images ae rescaled to 256 X 256 format for use. Also, the structure of .tfrec files during creation was :</p>\n<hr>\n<p><strong>Training &amp; Validation</strong></p>\n<p>feature = {<br>\n        \"image\" : bytes_features(feature_list[0]),<br>\n        \"image_name\" : bytes_features(feature_list[1]),<br>\n        \"label\" : int64_features(feature_list[2])<br>\n    }</p>\n<hr>\n<p><strong>Testing</strong> : <br>\nfeature = {<br>\n        \"image\" : bytes_features(feature_list[0]),<br>\n        \"image_name\" : bytes_features(feature_list[1])<br>\n    }</p>\n<hr>\n<p>(<em>Note the feature name is <strong>label</strong> in order to access it. I think in the official tfrecords of the competition it was class or something.</em>)</p>\n<h1><strong>LINK</strong> :</h1>\n<h2><strong><a href=\"https://www.kaggle.com/fireheart7/augmented-cassava-tfrec\" target=\"_blank\">https://www.kaggle.com/fireheart7/augmented-cassava-tfrec</a></strong></h2>\n<h1>Guide(Quite Standard) Book With This Dataset :</h1>\n<h2><strong><a href=\"https://www.kaggle.com/fireheart7/guide-mega-tfrec-dataset?scriptVersionId=49407011\" target=\"_blank\">https://www.kaggle.com/fireheart7/guide-mega-tfrec-dataset?scriptVersionId=49407011</a></strong></h2>\n<p>Hope this will be of aid in this competition! All the very best!</p>",
      "rawMarkdown": "Hi all!\n\nI prepared this dataset few days back. For each training image, I have applied random augmentations such as coarse dropout, mirror flip, and so on.. and cooked this mega TFRec dataset. \n\nAs we very well know, low bias problems can always be resolved by increasing the size of our dataset. As I tried it myself on very few starting models, I found the validation and training accuracy increasing by 3%. \n\nThis contains : \n* **180,000+ training images - .tfrec format**\n* **3000+ validation images - .tfrec format**\n* **1 test image - the one provided by the competition itself**\n\nAll images ae rescaled to 256 X 256 format for use. Also, the structure of .tfrec files during creation was :\n**************************************************************\n**Training & Validation**\n\nfeature = {\n        \"image\" : bytes_features(feature_list[0]),\n        \"image_name\" : bytes_features(feature_list[1]),\n        \"label\" : int64_features(feature_list[2])\n    }\n**************************************************************\n**Testing** : \nfeature = {\n        \"image\" : bytes_features(feature_list[0]),\n        \"image_name\" : bytes_features(feature_list[1])\n    }\n**************************************************************\n(*Note the feature name is **label** in order to access it. I think in the official tfrecords of the competition it was class or something.*)\n    \n# **LINK** : \n\n## **https://www.kaggle.com/fireheart7/augmented-cassava-tfrec**\n\n# Guide(Quite Standard) Book With This Dataset : \n\n## **https://www.kaggle.com/fireheart7/guide-mega-tfrec-dataset?scriptVersionId=49407011**\nHope this will be of aid in this competition! All the very best!",
      "votes": null
    },
    {
      "id": "1116809",
      "postDate": "12/17/2020 13:38:25",
      "content": "<p>Nice work although I think you will get better score on LB if you go with 512x512 (Atleast that's what I observed in the Discussion)</p>",
      "rawMarkdown": "Nice work although I think you will get better score on LB if you go with 512x512 (Atleast that's what I observed in the Discussion)",
      "votes": null
    },
    {
      "id": "1122239",
      "postDate": "12/22/2020 09:31:24",
      "content": "<p>oh ok! will try!! ain't 512 real big for memory ??</p>",
      "rawMarkdown": "oh ok! will try!! ain't 512 real big for memory ??",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1116809,
      "author_name": "harshsdw",
      "author_url": "",
      "post_date": "12/17/2020 13:38:25",
      "content": "<p>Nice work although I think you will get better score on LB if you go with 512x512 (Atleast that's what I observed in the Discussion)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1122239,
          "author_name": "fireheart7",
          "author_url": "",
          "post_date": "12/22/2020 09:31:24",
          "content": "<p>oh ok! will try!! ain't 512 real big for memory ??</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1113547": "Hi all!\n\nI prepared this dataset few days back. For each training image, I have applied random augmentations such as coarse dropout, mirror flip, and so on.. and cooked this mega TFRec dataset. \n\nAs we very well know, low bias problems can always be resolved by increasing the size of our dataset. As I tried it myself on very few starting models, I found the validation and training accuracy increasing by 3%. \n\nThis contains : \n* **180,000+ training images - .tfrec format**\n* **3000+ validation images - .tfrec format**\n* **1 test image - the one provided by the competition itself**\n\nAll images ae rescaled to 256 X 256 format for use. Also, the structure of .tfrec files during creation was :\n**************************************************************\n**Training & Validation**\n\nfeature = {\n        \"image\" : bytes_features(feature_list[0]),\n        \"image_name\" : bytes_features(feature_list[1]),\n        \"label\" : int64_features(feature_list[2])\n    }\n**************************************************************\n**Testing** : \nfeature = {\n        \"image\" : bytes_features(feature_list[0]),\n        \"image_name\" : bytes_features(feature_list[1])\n    }\n**************************************************************\n(*Note the feature name is **label** in order to access it. I think in the official tfrecords of the competition it was class or something.*)\n    \n# **LINK** : \n\n## **https://www.kaggle.com/fireheart7/augmented-cassava-tfrec**\n\n# Guide(Quite Standard) Book With This Dataset : \n\n## **https://www.kaggle.com/fireheart7/guide-mega-tfrec-dataset?scriptVersionId=49407011**\nHope this will be of aid in this competition! All the very best!",
    "1116809": "Nice work although I think you will get better score on LB if you go with 512x512 (Atleast that's what I observed in the Discussion)",
    "1122239": "oh ok! will try!! ain't 512 real big for memory ??"
  },
  "source": "meta"
}