{
  "id": 216676,
  "title": "Merged old and new dataset with extraimages (in image and TFrecord format)",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/216676",
  "author_name": "",
  "post_date": "2021-02-03T16:29:09.994055100Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I have merged Current and old data (from <a href=\"https://www.kaggle.com/c/cassava-disease\" target=\"_blank\">2019</a> competition) and converted it to Tfrecords and original images.</p>\n<p>The dataset contains folder for training images from both competitions and unlabelled extraimages from 2019 competition's extraimages and test images set. They are available in both image and tfrecords format with same storage manner as current (extraimages only has 'image' field).</p>\n<p>You can find the dataset <a href=\"https://www.kaggle.com/srg9000/cassava-plant-disease-merged-20192020\" target=\"_blank\">here</a>. </p>\n<p>The extraimages can be a part of pretraining approaches such as siamese network based or autoencoder based. I have created a kernel for autoencoder based pretraining which you can find <a href=\"https://www.kaggle.com/srg9000/autoencder-pretraining-with-merged-dataset\" target=\"_blank\">here</a> </p>\n<p>Do let me know your thoughts/suggestions.</p>",
  "messages": [
    {
      "id": "1184633",
      "postDate": "02/03/2021 16:29:09",
      "content": "<p>I have merged Current and old data (from <a href=\"https://www.kaggle.com/c/cassava-disease\" target=\"_blank\">2019</a> competition) and converted it to Tfrecords and original images.</p>\n<p>The dataset contains folder for training images from both competitions and unlabelled extraimages from 2019 competition's extraimages and test images set. They are available in both image and tfrecords format with same storage manner as current (extraimages only has 'image' field).</p>\n<p>You can find the dataset <a href=\"https://www.kaggle.com/srg9000/cassava-plant-disease-merged-20192020\" target=\"_blank\">here</a>. </p>\n<p>The extraimages can be a part of pretraining approaches such as siamese network based or autoencoder based. I have created a kernel for autoencoder based pretraining which you can find <a href=\"https://www.kaggle.com/srg9000/autoencder-pretraining-with-merged-dataset\" target=\"_blank\">here</a> </p>\n<p>Do let me know your thoughts/suggestions.</p>",
      "rawMarkdown": "I have merged Current and old data (from [2019](https://www.kaggle.com/c/cassava-disease) competition) and converted it to Tfrecords and original images.\n\nThe dataset contains folder for training images from both competitions and unlabelled extraimages from 2019 competition's extraimages and test images set. They are available in both image and tfrecords format with same storage manner as current (extraimages only has 'image' field).\n\nYou can find the dataset [here](https://www.kaggle.com/srg9000/cassava-plant-disease-merged-20192020). \n\nThe extraimages can be a part of pretraining approaches such as siamese network based or autoencoder based. I have created a kernel for autoencoder based pretraining which you can find [here](https://www.kaggle.com/srg9000/autoencder-pretraining-with-merged-dataset) \n\nDo let me know your thoughts/suggestions.",
      "votes": null
    },
    {
      "id": "1185701",
      "postDate": "02/04/2021 09:50:23",
      "content": "<p>Interesting…Will there be duplicates?</p>",
      "rawMarkdown": "Interesting...Will there be duplicates?",
      "votes": null
    },
    {
      "id": "1186194",
      "postDate": "02/04/2021 16:34:56",
      "content": "<p>Yes, I haven't done anything to remove them as that will be taken care of by augmentation and class weighing while training. I should have checked for those, and will update the dataset if there are significant amount of duplicates</p>",
      "rawMarkdown": "Yes, I haven't done anything to remove them as that will be taken care of by augmentation and class weighing while training. I should have checked for those, and will update the dataset if there are significant amount of duplicates",
      "votes": null
    },
    {
      "id": "1186687",
      "postDate": "02/05/2021 02:01:40",
      "content": "<p>Thanks a lot for the reply!!! Appreciate it!</p>",
      "rawMarkdown": "Thanks a lot for the reply!!! Appreciate it!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1185701,
      "author_name": "nelsonwongisme",
      "author_url": "",
      "post_date": "02/04/2021 09:50:23",
      "content": "<p>Interesting…Will there be duplicates?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1186194,
          "author_name": "srg9000",
          "author_url": "",
          "post_date": "02/04/2021 16:34:56",
          "content": "<p>Yes, I haven't done anything to remove them as that will be taken care of by augmentation and class weighing while training. I should have checked for those, and will update the dataset if there are significant amount of duplicates</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1186687,
          "author_name": "nelsonwongisme",
          "author_url": "",
          "post_date": "02/05/2021 02:01:40",
          "content": "<p>Thanks a lot for the reply!!! Appreciate it!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1184633": "I have merged Current and old data (from [2019](https://www.kaggle.com/c/cassava-disease) competition) and converted it to Tfrecords and original images.\n\nThe dataset contains folder for training images from both competitions and unlabelled extraimages from 2019 competition's extraimages and test images set. They are available in both image and tfrecords format with same storage manner as current (extraimages only has 'image' field).\n\nYou can find the dataset [here](https://www.kaggle.com/srg9000/cassava-plant-disease-merged-20192020). \n\nThe extraimages can be a part of pretraining approaches such as siamese network based or autoencoder based. I have created a kernel for autoencoder based pretraining which you can find [here](https://www.kaggle.com/srg9000/autoencder-pretraining-with-merged-dataset) \n\nDo let me know your thoughts/suggestions.",
    "1185701": "Interesting...Will there be duplicates?",
    "1186194": "Yes, I haven't done anything to remove them as that will be taken care of by augmentation and class weighing while training. I should have checked for those, and will update the dataset if there are significant amount of duplicates",
    "1186687": "Thanks a lot for the reply!!! Appreciate it!"
  },
  "source": "meta"
}