{
  "id": 199131,
  "title": "Merged(2019 && 2020)-Duplicates removed TFrecords - 512 x 512",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/199131",
  "author_name": "",
  "post_date": "2020-11-24T15:07:08.030958Z",
  "votes": 53,
  "comment_count": 16,
  "views": 0,
  "content": "<p>We created a dataset merging both training sets of 2019 and 2020 which is available <a href=\"https://www.kaggle.com/kingofarmy/cassavapreprocessed?select=merged_train_tfrecs\" target=\"_blank\">here</a></p>\n<p>A dataset containing samples generated by GAN will be released soon…</p>",
  "messages": [
    {
      "id": "1089505",
      "postDate": "11/24/2020 15:07:08",
      "content": "<p>We created a dataset merging both training sets of 2019 and 2020 which is available <a href=\"https://www.kaggle.com/kingofarmy/cassavapreprocessed?select=merged_train_tfrecs\" target=\"_blank\">here</a></p>\n<p>A dataset containing samples generated by GAN will be released soon…</p>",
      "rawMarkdown": "We created a dataset merging both training sets of 2019 and 2020 which is available [here](https://www.kaggle.com/kingofarmy/cassavapreprocessed?select=merged_train_tfrecs)\n\nA dataset containing samples generated by GAN will be released soon...",
      "votes": null
    },
    {
      "id": "1089539",
      "postDate": "11/24/2020 15:40:29",
      "content": "<p>Nice efforts! How did you remove duplicates?</p>",
      "rawMarkdown": "Nice efforts! How did you remove duplicates?",
      "votes": null
    },
    {
      "id": "1089615",
      "postDate": "11/24/2020 16:26:21",
      "content": "<p>used python package called imagededup</p>",
      "rawMarkdown": "used python package called imagededup",
      "votes": null
    },
    {
      "id": "1090769",
      "postDate": "11/25/2020 15:19:01",
      "content": "<p>Thank you for your sharing.<br>\nIs it possible to share jpg files?<br>\nMoreover, did you also merge the extra data on 2019?</p>",
      "rawMarkdown": "Thank you for your sharing.\nIs it possible to share jpg files?\nMoreover, did you also merge the extra data on 2019?",
      "votes": null
    },
    {
      "id": "1090805",
      "postDate": "11/25/2020 15:42:30",
      "content": "<p>Yeah sure will add that too. I did not add the extra data.</p>",
      "rawMarkdown": "Yeah sure will add that too. I did not add the extra data.",
      "votes": null
    },
    {
      "id": "1090891",
      "postDate": "11/25/2020 16:42:33",
      "content": "<p>Thank you!<br>\nLooking forward to it.</p>",
      "rawMarkdown": "Thank you!\nLooking forward to it.",
      "votes": null
    },
    {
      "id": "1091442",
      "postDate": "11/26/2020 03:05:21",
      "content": "<p>I used the 2019 jpeg data with reference to the csv file you created.<br>\nLocal CV improved from 0.873 to 0.876. Many thanks!</p>",
      "rawMarkdown": "I used the 2019 jpeg data with reference to the csv file you created.\nLocal CV improved from 0.873 to 0.876. Many thanks!",
      "votes": null
    },
    {
      "id": "1091452",
      "postDate": "11/26/2020 03:15:28",
      "content": "<p>Happy to hear this</p>",
      "rawMarkdown": "Happy to hear this",
      "votes": null
    },
    {
      "id": "1093598",
      "postDate": "11/27/2020 21:13:36",
      "content": "<p>Nice work . I am planning to use this new dataset. Anyhow you could rename the files in the competition format to include the number of records in each tfrecord file like this <br>\nld_train00-<strong>1338</strong>.tfrec <br>\nagain thanks a lot for the merged data set </p>",
      "rawMarkdown": "Nice work . I am planning to use this new dataset. Anyhow you could rename the files in the competition format to include the number of records in each tfrecord file like this \nld_train00-**1338**.tfrec \nagain thanks a lot for the merged data set",
      "votes": null
    },
    {
      "id": "1095263",
      "postDate": "11/29/2020 12:18:14",
      "content": "<p>Thanks for sharing! This is a great effort.</p>",
      "rawMarkdown": "Thanks for sharing! This is a great effort.",
      "votes": null
    },
    {
      "id": "1098995",
      "postDate": "12/02/2020 02:20:00",
      "content": "<p>Thank you so much </p>",
      "rawMarkdown": "Thank you so much",
      "votes": null
    },
    {
      "id": "1101610",
      "postDate": "12/04/2020 04:20:23",
      "content": "<p>Has anyone tried using the tfrecords in this dataset. I am seeing a lot of latency <br>\n<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/201278\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/201278</a></p>",
      "rawMarkdown": "Has anyone tried using the tfrecords in this dataset. I am seeing a lot of latency \nhttps://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/201278",
      "votes": null
    },
    {
      "id": "1108643",
      "postDate": "12/10/2020 21:20:44",
      "content": "<p><a href=\"https://www.kaggle.com/kingofarmy\" target=\"_blank\">@kingofarmy</a> Thanks for sharing this dataset. Does that mean there is only roughtly 5 000 images from 2019 ? I saw that the merged data has 27k.</p>",
      "rawMarkdown": "kingofarmy Thanks for sharing this dataset. Does that mean there is only roughtly 5 000 images from 2019 ? I saw that the merged data has 27k.",
      "votes": null
    },
    {
      "id": "1132316",
      "postDate": "12/30/2020 09:59:12",
      "content": "<p>How to use your dataset if I use pytorch, not tensorflow ? thx</p>",
      "rawMarkdown": "How to use your dataset if I use pytorch, not tensorflow ? thx",
      "votes": null
    },
    {
      "id": "1147146",
      "postDate": "01/10/2021 10:06:02",
      "content": "<p>I havent tried it with pytorch yet. may be this could help you <a href=\"https://github.com/vahidk/tfrecord\" target=\"_blank\">https://github.com/vahidk/tfrecord</a></p>",
      "rawMarkdown": "I havent tried it with pytorch yet. may be this could help you https://github.com/vahidk/tfrecord",
      "votes": null
    },
    {
      "id": "1147149",
      "postDate": "01/10/2021 10:06:44",
      "content": "<p>The latency may be due to it isn't in the competitions format</p>",
      "rawMarkdown": "The latency may be due to it isn't in the competitions format",
      "votes": null
    },
    {
      "id": "1160903",
      "postDate": "01/20/2021 07:52:36",
      "content": "<p>Hey, I tried using imagededup, but I am getting the following error for all the images !! Are you aware of this? <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2029256%2F39e75371924f1fe6c1bcdec1c88b5340%2FScreen%20Shot%202021-01-20%20at%201.21.34%20PM.png?generation=1611129152228310&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hey, I tried using imagededup, but I am getting the following error for all the images !! Are you aware of this? ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2029256%2F39e75371924f1fe6c1bcdec1c88b5340%2FScreen%20Shot%202021-01-20%20at%201.21.34%20PM.png?generation=1611129152228310&alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1089539,
      "author_name": "kaushal2896",
      "author_url": "",
      "post_date": "11/24/2020 15:40:29",
      "content": "<p>Nice efforts! How did you remove duplicates?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1089615,
          "author_name": "kingofarmy",
          "author_url": "",
          "post_date": "11/24/2020 16:26:21",
          "content": "<p>used python package called imagededup</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1160903,
          "author_name": "ks2019",
          "author_url": "",
          "post_date": "01/20/2021 07:52:36",
          "content": "<p>Hey, I tried using imagededup, but I am getting the following error for all the images !! Are you aware of this? <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2029256%2F39e75371924f1fe6c1bcdec1c88b5340%2FScreen%20Shot%202021-01-20%20at%201.21.34%20PM.png?generation=1611129152228310&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1090769,
      "author_name": "kasim0226",
      "author_url": "",
      "post_date": "11/25/2020 15:19:01",
      "content": "<p>Thank you for your sharing.<br>\nIs it possible to share jpg files?<br>\nMoreover, did you also merge the extra data on 2019?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1090805,
          "author_name": "kingofarmy",
          "author_url": "",
          "post_date": "11/25/2020 15:42:30",
          "content": "<p>Yeah sure will add that too. I did not add the extra data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1090891,
          "author_name": "kasim0226",
          "author_url": "",
          "post_date": "11/25/2020 16:42:33",
          "content": "<p>Thank you!<br>\nLooking forward to it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1091442,
      "author_name": "yosukeyama",
      "author_url": "",
      "post_date": "11/26/2020 03:05:21",
      "content": "<p>I used the 2019 jpeg data with reference to the csv file you created.<br>\nLocal CV improved from 0.873 to 0.876. Many thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1091452,
          "author_name": "kingofarmy",
          "author_url": "",
          "post_date": "11/26/2020 03:15:28",
          "content": "<p>Happy to hear this</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1093598,
      "author_name": "venkat555",
      "author_url": "",
      "post_date": "11/27/2020 21:13:36",
      "content": "<p>Nice work . I am planning to use this new dataset. Anyhow you could rename the files in the competition format to include the number of records in each tfrecord file like this <br>\nld_train00-<strong>1338</strong>.tfrec <br>\nagain thanks a lot for the merged data set </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1095263,
      "author_name": "ar2017",
      "author_url": "",
      "post_date": "11/29/2020 12:18:14",
      "content": "<p>Thanks for sharing! This is a great effort.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1098995,
          "author_name": "kingofarmy",
          "author_url": "",
          "post_date": "12/02/2020 02:20:00",
          "content": "<p>Thank you so much </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1101610,
      "author_name": "venkat555",
      "author_url": "",
      "post_date": "12/04/2020 04:20:23",
      "content": "<p>Has anyone tried using the tfrecords in this dataset. I am seeing a lot of latency <br>\n<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/201278\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/201278</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1147149,
          "author_name": "kingofarmy",
          "author_url": "",
          "post_date": "01/10/2021 10:06:44",
          "content": "<p>The latency may be due to it isn't in the competitions format</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1108643,
      "author_name": "ludovick",
      "author_url": "",
      "post_date": "12/10/2020 21:20:44",
      "content": "<p><a href=\"https://www.kaggle.com/kingofarmy\" target=\"_blank\">@kingofarmy</a> Thanks for sharing this dataset. Does that mean there is only roughtly 5 000 images from 2019 ? I saw that the merged data has 27k.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1132316,
      "author_name": "clwwlc",
      "author_url": "",
      "post_date": "12/30/2020 09:59:12",
      "content": "<p>How to use your dataset if I use pytorch, not tensorflow ? thx</p>",
      "votes": null,
      "replies": [
        {
          "id": 1147146,
          "author_name": "kingofarmy",
          "author_url": "",
          "post_date": "01/10/2021 10:06:02",
          "content": "<p>I havent tried it with pytorch yet. may be this could help you <a href=\"https://github.com/vahidk/tfrecord\" target=\"_blank\">https://github.com/vahidk/tfrecord</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1089505": "We created a dataset merging both training sets of 2019 and 2020 which is available [here](https://www.kaggle.com/kingofarmy/cassavapreprocessed?select=merged_train_tfrecs)\n\nA dataset containing samples generated by GAN will be released soon...",
    "1089539": "Nice efforts! How did you remove duplicates?",
    "1089615": "used python package called imagededup",
    "1090769": "Thank you for your sharing.\nIs it possible to share jpg files?\nMoreover, did you also merge the extra data on 2019?",
    "1090805": "Yeah sure will add that too. I did not add the extra data.",
    "1090891": "Thank you!\nLooking forward to it.",
    "1091442": "I used the 2019 jpeg data with reference to the csv file you created.\nLocal CV improved from 0.873 to 0.876. Many thanks!",
    "1091452": "Happy to hear this",
    "1093598": "Nice work . I am planning to use this new dataset. Anyhow you could rename the files in the competition format to include the number of records in each tfrecord file like this \nld_train00-**1338**.tfrec \nagain thanks a lot for the merged data set",
    "1095263": "Thanks for sharing! This is a great effort.",
    "1098995": "Thank you so much",
    "1101610": "Has anyone tried using the tfrecords in this dataset. I am seeing a lot of latency \nhttps://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/201278",
    "1108643": "kingofarmy Thanks for sharing this dataset. Does that mean there is only roughtly 5 000 images from 2019 ? I saw that the merged data has 27k.",
    "1132316": "How to use your dataset if I use pytorch, not tensorflow ? thx",
    "1147146": "I havent tried it with pytorch yet. may be this could help you https://github.com/vahidk/tfrecord",
    "1147149": "The latency may be due to it isn't in the competitions format",
    "1160903": "Hey, I tried using imagededup, but I am getting the following error for all the images !! Are you aware of this? ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2029256%2F39e75371924f1fe6c1bcdec1c88b5340%2FScreen%20Shot%202021-01-20%20at%201.21.34%20PM.png?generation=1611129152228310&alt=media)"
  },
  "source": "meta"
}