{
  "id": 215490,
  "title": "For rising in LB - Old Competition Cassava Training Data in TF format",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/215490",
  "author_name": "Vicky Goyal",
  "post_date": "2021-01-30T06:33:02.443000",
  "votes": 5,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I have converted old cassava training data in RGB format in tf file format. The height and width is 512 * 512 pixels. The format I used is exactly the same as new competition. The dataset can be used to rise more on the lb.</p>\n<p>To combine both old and new dataset use:</p>\n<pre><code>from kaggle_datasets import KaggleDatasets\nimport tensorflow as tf\nGCS_PATH = KaggleDatasets().get_gcs_path('oldcassavadata') \nG_PATH = KaggleDatasets().get_gcs_path('cassava-leaf-disease-classification') \n\nALL_TFRECS=tf.io.gfile.glob(GCS_PATH + '/*.tfrec') +tf.io.gfile.glob(G_PATH + '/train_tfrecords/*.tfrec')\nprint(ALL_TFRECS)\n</code></pre>\n<p>To get the dataset click <a href=\"https://www.kaggle.com/vickygoyal/oldcassavadata\" target=\"_blank\">here</a></p>",
  "messages": [
    {
      "id": 1177234,
      "postDate": "2021-01-30T06:33:02.443Z",
      "content": "<p>I have converted old cassava training data in RGB format in tf file format. The height and width is 512 * 512 pixels. The format I used is exactly the same as new competition. The dataset can be used to rise more on the lb.</p>\n<p>To combine both old and new dataset use:</p>\n<pre><code>from kaggle_datasets import KaggleDatasets\nimport tensorflow as tf\nGCS_PATH = KaggleDatasets().get_gcs_path('oldcassavadata') \nG_PATH = KaggleDatasets().get_gcs_path('cassava-leaf-disease-classification') \n\nALL_TFRECS=tf.io.gfile.glob(GCS_PATH + '/*.tfrec') +tf.io.gfile.glob(G_PATH + '/train_tfrecords/*.tfrec')\nprint(ALL_TFRECS)\n</code></pre>\n<p>To get the dataset click <a href=\"https://www.kaggle.com/vickygoyal/oldcassavadata\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "I have converted old cassava training data in RGB format in tf file format. The height and width is 512 * 512 pixels. The format I used is exactly the same as new competition. The dataset can be used to rise more on the lb.\n\nTo combine both old and new dataset use:\n```\nfrom kaggle_datasets import KaggleDatasets\nimport tensorflow as tf\nGCS_PATH = KaggleDatasets().get_gcs_path('oldcassavadata') \nG_PATH = KaggleDatasets().get_gcs_path('cassava-leaf-disease-classification') \n\nALL_TFRECS=tf.io.gfile.glob(GCS_PATH + '/*.tfrec') +tf.io.gfile.glob(G_PATH + '/train_tfrecords/*.tfrec')\nprint(ALL_TFRECS)\n```\n\nTo get the dataset click [here](https://www.kaggle.com/vickygoyal/oldcassavadata)",
      "votes": 5
    },
    {
      "id": 1177345,
      "postDate": "2021-01-30T08:11:59.170Z",
      "content": "<p>Thanks for sharing!</p>\n<p>Are the duplicates removed?</p>",
      "rawMarkdown": "Thanks for sharing!\n\nAre the duplicates removed?",
      "votes": 1
    },
    {
      "id": 1185792,
      "postDate": "2021-02-04T11:16:12.363Z",
      "content": "<p>Thanks for the data,  some basic question:<br>\nI want to ask, how can I use kaggle_datasets  outside the kaggle kernel , locally? Can someone point me to the link where I can install the exact same module.<br>\nThank you.</p>",
      "rawMarkdown": "\nThanks for the data,  some basic question:\n\nI want to ask, how can I use kaggle_datasets  outside the kaggle kernel , locally? Can someone point me to the link where I can install the exact same module.\nThank you.",
      "replies": [
        {
          "id": 1185809,
          "postDate": "2021-02-04T11:26:56.140Z",
          "content": "<p>You can download the data and notebook to try it on local but doing locally will take may be from many  hours to days to train depending on your local gpu. You cannot use tpu locally. </p>",
          "rawMarkdown": "You can download the data and notebook to try it on local but doing locally will take may be from many  hours to days to train depending on your local gpu. You cannot use tpu locally. "
        }
      ]
    },
    {
      "id": 1182262,
      "postDate": "2021-02-02T11:40:05.780Z",
      "content": "<p>How i use this dataset</p>",
      "rawMarkdown": "How i use this dataset"
    },
    {
      "id": 1177360,
      "postDate": "2021-01-30T08:25:05.357Z",
      "content": "<p>No, duplicates if any are still present. The duplicates still needs to be removed.</p>",
      "rawMarkdown": "No, duplicates if any are still present. The duplicates still needs to be removed.",
      "replies": [
        {
          "id": 1177454,
          "postDate": "2021-01-30T09:41:21.183Z",
          "content": "<p>If my training is doing heavy augmentation do I really have to remove duplicates - after all they will be \"changed\" due to the augmentation?</p>",
          "rawMarkdown": "If my training is doing heavy augmentation do I really have to remove duplicates - after all they will be \"changed\" due to the augmentation?"
        },
        {
          "id": 1177478,
          "postDate": "2021-01-30T09:59:45.390Z",
          "content": "<p>Yes duplicates will be changed but the pattern will remain more or less same due to heavy augmentation. The problem is if we have same image in CV as in training it gives false hope because we get high CV because training is already done on it. The model has already learned the pattern during training and we already know the answer and hence high CV.</p>",
          "rawMarkdown": "Yes duplicates will be changed but the pattern will remain more or less same due to heavy augmentation. The problem is if we have same image in CV as in training it gives false hope because we get high CV because training is already done on it. The model has already learned the pattern during training and we already know the answer and hence high CV.",
          "votes": 1
        },
        {
          "id": 1177568,
          "postDate": "2021-01-30T11:34:53.480Z",
          "content": "<p>Agree that duplicates that end up in validation and train will increase cv.  When I know duplicates might exist I augment validation also to slow down the fake news of better cv.   </p>",
          "rawMarkdown": "Agree that duplicates that end up in validation and train will increase cv.  When I know duplicates might exist I augment validation also to slow down the fake news of better cv.   \n\n\n"
        },
        {
          "id": 1178465,
          "postDate": "2021-01-30T22:25:43.070Z",
          "content": "<p>Good idea. Will give it a try to see how it works in this competition.</p>",
          "rawMarkdown": "Good idea. Will give it a try to see how it works in this competition."
        }
      ]
    },
    {
      "id": 1185791,
      "postDate": "2021-02-04T11:16:12.363Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1177345,
      "author_name": "Nikita Kuzmenkov",
      "author_url": "",
      "post_date": "2021-01-30T08:11:59.170000",
      "content": "<p>Thanks for sharing!</p>\n<p>Are the duplicates removed?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1185792,
      "author_name": "mane.stoimchev",
      "author_url": "",
      "post_date": "2021-02-04T11:16:12.363000",
      "content": "<p>Thanks for the data,  some basic question:<br>\nI want to ask, how can I use kaggle_datasets  outside the kaggle kernel , locally? Can someone point me to the link where I can install the exact same module.<br>\nThank you.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1185809,
          "author_name": "Vicky Goyal",
          "author_url": "",
          "post_date": "2021-02-04T11:26:56.140000",
          "content": "<p>You can download the data and notebook to try it on local but doing locally will take may be from many  hours to days to train depending on your local gpu. You cannot use tpu locally. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1182262,
      "author_name": "JingYun Zeng",
      "author_url": "",
      "post_date": "2021-02-02T11:40:05.780000",
      "content": "<p>How i use this dataset</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1177360,
      "author_name": "Vicky Goyal",
      "author_url": "",
      "post_date": "2021-01-30T08:25:05.357000",
      "content": "<p>No, duplicates if any are still present. The duplicates still needs to be removed.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1177454,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2021-01-30T09:41:21.183000",
          "content": "<p>If my training is doing heavy augmentation do I really have to remove duplicates - after all they will be \"changed\" due to the augmentation?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1177478,
          "author_name": "Vicky Goyal",
          "author_url": "",
          "post_date": "2021-01-30T09:59:45.390000",
          "content": "<p>Yes duplicates will be changed but the pattern will remain more or less same due to heavy augmentation. The problem is if we have same image in CV as in training it gives false hope because we get high CV because training is already done on it. The model has already learned the pattern during training and we already know the answer and hence high CV.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1177568,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2021-01-30T11:34:53.480000",
          "content": "<p>Agree that duplicates that end up in validation and train will increase cv.  When I know duplicates might exist I augment validation also to slow down the fake news of better cv.   </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1178465,
          "author_name": "Vicky Goyal",
          "author_url": "",
          "post_date": "2021-01-30T22:25:43.070000",
          "content": "<p>Good idea. Will give it a try to see how it works in this competition.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1185791,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-04T11:16:12.363000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1177234": "I have converted old cassava training data in RGB format in tf file format. The height and width is 512 * 512 pixels. The format I used is exactly the same as new competition. The dataset can be used to rise more on the lb.\n\nTo combine both old and new dataset use:\n```\nfrom kaggle_datasets import KaggleDatasets\nimport tensorflow as tf\nGCS_PATH = KaggleDatasets().get_gcs_path('oldcassavadata') \nG_PATH = KaggleDatasets().get_gcs_path('cassava-leaf-disease-classification') \n\nALL_TFRECS=tf.io.gfile.glob(GCS_PATH + '/*.tfrec') +tf.io.gfile.glob(G_PATH + '/train_tfrecords/*.tfrec')\nprint(ALL_TFRECS)\n```\n\nTo get the dataset click [here](https://www.kaggle.com/vickygoyal/oldcassavadata)",
    "1177345": "Thanks for sharing!\n\nAre the duplicates removed?",
    "1185792": "\nThanks for the data,  some basic question:\n\nI want to ask, how can I use kaggle_datasets  outside the kaggle kernel , locally? Can someone point me to the link where I can install the exact same module.\nThank you.",
    "1182262": "How i use this dataset",
    "1177360": "No, duplicates if any are still present. The duplicates still needs to be removed.",
    "1185791": ""
  }
}