{
  "id": 200201,
  "title": "Merged JPEG Dataset [Duplicates Removed]",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/200201",
  "author_name": "Tahsin Mostafiz",
  "post_date": "2020-11-29T12:04:08.989000",
  "votes": 70,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I have created a dataset combining images without resizing from the 2020 and 2019 competitions <a href=\"https://www.kaggle.com/tahsin/cassava-leaf-disease-merged\" target=\"_blank\">here</a>. I found 722 duplicate images using <a href=\"https://www.kaggle.com/zzy990106/duplicate-images-in-two-competitions\" target=\"_blank\">this</a> notebook and removed them. The <code>merged.csv</code> contains all images, corresponding labels, and source(2019/2020). </p>",
  "messages": [
    {
      "id": 1095254,
      "postDate": "2020-11-29T12:04:08.990Z",
      "content": "<p>I have created a dataset combining images without resizing from the 2020 and 2019 competitions <a href=\"https://www.kaggle.com/tahsin/cassava-leaf-disease-merged\" target=\"_blank\">here</a>. I found 722 duplicate images using <a href=\"https://www.kaggle.com/zzy990106/duplicate-images-in-two-competitions\" target=\"_blank\">this</a> notebook and removed them. The <code>merged.csv</code> contains all images, corresponding labels, and source(2019/2020). </p>",
      "rawMarkdown": "I have created a dataset combining images without resizing from the 2020 and 2019 competitions [here](https://www.kaggle.com/tahsin/cassava-leaf-disease-merged). I found 722 duplicate images using [this](https://www.kaggle.com/zzy990106/duplicate-images-in-two-competitions) notebook and removed them. The `merged.csv` contains all images, corresponding labels, and source(2019/2020). ",
      "votes": 70
    },
    {
      "id": 1097831,
      "postDate": "2020-12-01T09:18:34.673Z",
      "content": "<p>Cool, thank you!<br>\nI wonder why the organizers didn't include all previous images in the current dataset. 😯</p>",
      "rawMarkdown": "Cool, thank you!\nI wonder why the organizers didn't include all previous images in the current dataset. 😯",
      "votes": 1
    },
    {
      "id": 1123426,
      "postDate": "2020-12-23T08:05:57.200Z",
      "content": "<p>Anyone's LB improved using the dataset?</p>",
      "rawMarkdown": "Anyone's LB improved using the dataset?\n",
      "replies": [
        {
          "id": 1124016,
          "postDate": "2020-12-23T16:29:01.710Z",
          "content": "<p>Mine did not, my CV didnt change either.</p>",
          "rawMarkdown": "Mine did not, my CV didnt change either."
        },
        {
          "id": 1124604,
          "postDate": "2020-12-24T04:41:38.690Z",
          "content": "<p>To be honest, it is really strange; generally adding more data makes the model more generalize; I think the noise in labels is really effecting it greatly.</p>",
          "rawMarkdown": "To be honest, it is really strange; generally adding more data makes the model more generalize; I think the noise in labels is really effecting it greatly."
        }
      ]
    },
    {
      "id": 1095394,
      "postDate": "2020-11-29T15:07:53.803Z",
      "content": "<p>I tried using previous dataset but it reduced my LB by 0.001. I didn't remove duplicates… That might be the reason.</p>",
      "rawMarkdown": "I tried using previous dataset but it reduced my LB by 0.001. I didn't remove duplicates... That might be the reason.",
      "replies": [
        {
          "id": 1095407,
          "postDate": "2020-11-29T15:18:14.937Z",
          "content": "<p>You can try my dataset. Also, you can easily switch between 2019 and 2020 data now as I've added a source for each image.</p>",
          "rawMarkdown": "You can try my dataset. Also, you can easily switch between 2019 and 2020 data now as I've added a source for each image."
        },
        {
          "id": 1095418,
          "postDate": "2020-11-29T15:29:24.593Z",
          "content": "<p>Did it improve your score?</p>",
          "rawMarkdown": "Did it improve your score?"
        },
        {
          "id": 1095450,
          "postDate": "2020-11-29T15:50:16.813Z",
          "content": "<p>I haven't finished training on this dataset yet.</p>",
          "rawMarkdown": "I haven't finished training on this dataset yet."
        },
        {
          "id": 1095534,
          "postDate": "2020-11-29T17:43:43.210Z",
          "content": "<p>For me it did not change CV , my CV is the same.</p>",
          "rawMarkdown": "For me it did not change CV , my CV is the same.",
          "votes": 2
        },
        {
          "id": 1097302,
          "postDate": "2020-12-01T02:26:58.720Z",
          "content": "<p>Is there any improvement after using this data?</p>",
          "rawMarkdown": "Is there any improvement after using this data?"
        },
        {
          "id": 1101188,
          "postDate": "2020-12-03T17:41:48.807Z",
          "content": "<p>I think since the merged data is still unbalanced, and your model still couldn't generalize properly. I can be wrong. <a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> </p>",
          "rawMarkdown": "I think since the merged data is still unbalanced, and your model still couldn't generalize properly. I can be wrong. @aliabdin1 "
        },
        {
          "id": 1101366,
          "postDate": "2020-12-03T21:06:58.990Z",
          "content": "<p>I also tought about that, so plainly using the merged dataset did not improve my CV at all. It also did not improve my LB. </p>",
          "rawMarkdown": "I also tought about that, so plainly using the merged dataset did not improve my CV at all. It also did not improve my LB. "
        },
        {
          "id": 1101495,
          "postDate": "2020-12-04T00:14:03.330Z",
          "content": "<p><a href=\"https://www.kaggle.com/kaushal2896\" target=\"_blank\">@kaushal2896</a> <a href=\"https://www.kaggle.com/sjtuyxc\" target=\"_blank\">@sjtuyxc</a> the merged data improved my CV but somehow my LB dropped. </p>",
          "rawMarkdown": "@kaushal2896 @sjtuyxc the merged data improved my CV but somehow my LB dropped. "
        },
        {
          "id": 1101568,
          "postDate": "2020-12-04T03:02:19.443Z",
          "content": "<p>Me too！My LB dropped.</p>",
          "rawMarkdown": "Me too！My LB dropped."
        },
        {
          "id": 1101732,
          "postDate": "2020-12-04T07:22:18.543Z",
          "content": "<p>Is it possible that the old data has too much label noise?</p>",
          "rawMarkdown": "Is it possible that the old data has too much label noise?"
        },
        {
          "id": 1102003,
          "postDate": "2020-12-04T13:33:23.553Z",
          "content": "<p>yes, it can be;  but let us go with the old kaggle saying \"trust your CV more \".</p>",
          "rawMarkdown": "yes, it can be;  but let us go with the old kaggle saying \"trust your CV more \".",
          "votes": 1
        },
        {
          "id": 1124335,
          "postDate": "2020-12-23T20:47:21.080Z",
          "content": "<p>I haven't tried including the 2019 data, but isn't it true that if you do, you should not include it in the validation set, otherwise you're validating against data that may be too different from the test data? That may explain why it's possible the 2019 data could improve the  CV score but lower the LB score.</p>",
          "rawMarkdown": "I haven't tried including the 2019 data, but isn't it true that if you do, you should not include it in the validation set, otherwise you're validating against data that may be too different from the test data? That may explain why it's possible the 2019 data could improve the  CV score but lower the LB score.",
          "votes": 6
        }
      ]
    },
    {
      "id": 1101181,
      "postDate": "2020-12-03T17:36:30.767Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1095302,
      "postDate": "2020-11-29T13:10:17.350Z",
      "content": "<p>Thank you, thats very useful.</p>",
      "rawMarkdown": "Thank you, thats very useful.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1097831,
      "author_name": "Sergey Zlobin",
      "author_url": "",
      "post_date": "2020-12-01T09:18:34.673000",
      "content": "<p>Cool, thank you!<br>\nI wonder why the organizers didn't include all previous images in the current dataset. 😯</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1123426,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2020-12-23T08:05:57.200000",
      "content": "<p>Anyone's LB improved using the dataset?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1124016,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-12-23T16:29:01.710000",
          "content": "<p>Mine did not, my CV didnt change either.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1124604,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2020-12-24T04:41:38.690000",
          "content": "<p>To be honest, it is really strange; generally adding more data makes the model more generalize; I think the noise in labels is really effecting it greatly.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1095394,
      "author_name": "Kaushal Shah",
      "author_url": "",
      "post_date": "2020-11-29T15:07:53.803000",
      "content": "<p>I tried using previous dataset but it reduced my LB by 0.001. I didn't remove duplicates… That might be the reason.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1095407,
          "author_name": "Tahsin Mostafiz",
          "author_url": "",
          "post_date": "2020-11-29T15:18:14.937000",
          "content": "<p>You can try my dataset. Also, you can easily switch between 2019 and 2020 data now as I've added a source for each image.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1095418,
          "author_name": "Kaushal Shah",
          "author_url": "",
          "post_date": "2020-11-29T15:29:24.593000",
          "content": "<p>Did it improve your score?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1095450,
          "author_name": "Tahsin Mostafiz",
          "author_url": "",
          "post_date": "2020-11-29T15:50:16.813000",
          "content": "<p>I haven't finished training on this dataset yet.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1095534,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-11-29T17:43:43.210000",
          "content": "<p>For me it did not change CV , my CV is the same.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1097302,
          "author_name": "sky_fly",
          "author_url": "",
          "post_date": "2020-12-01T02:26:58.720000",
          "content": "<p>Is there any improvement after using this data?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1101188,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2020-12-03T17:41:48.807000",
          "content": "<p>I think since the merged data is still unbalanced, and your model still couldn't generalize properly. I can be wrong. <a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1101366,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-12-03T21:06:58.990000",
          "content": "<p>I also tought about that, so plainly using the merged dataset did not improve my CV at all. It also did not improve my LB. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1101495,
          "author_name": "Tahsin Mostafiz",
          "author_url": "",
          "post_date": "2020-12-04T00:14:03.330000",
          "content": "<p><a href=\"https://www.kaggle.com/kaushal2896\" target=\"_blank\">@kaushal2896</a> <a href=\"https://www.kaggle.com/sjtuyxc\" target=\"_blank\">@sjtuyxc</a> the merged data improved my CV but somehow my LB dropped. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1101568,
          "author_name": "sky_fly",
          "author_url": "",
          "post_date": "2020-12-04T03:02:19.443000",
          "content": "<p>Me too！My LB dropped.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1101732,
          "author_name": "Tahsin Mostafiz",
          "author_url": "",
          "post_date": "2020-12-04T07:22:18.543000",
          "content": "<p>Is it possible that the old data has too much label noise?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1102003,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2020-12-04T13:33:23.553000",
          "content": "<p>yes, it can be;  but let us go with the old kaggle saying \"trust your CV more \".</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1124335,
          "author_name": "David J. Slate",
          "author_url": "",
          "post_date": "2020-12-23T20:47:21.080000",
          "content": "<p>I haven't tried including the 2019 data, but isn't it true that if you do, you should not include it in the validation set, otherwise you're validating against data that may be too different from the test data? That may explain why it's possible the 2019 data could improve the  CV score but lower the LB score.</p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 1101181,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-03T17:36:30.767000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1095302,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2020-11-29T13:10:17.350000",
      "content": "<p>Thank you, thats very useful.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1095254": "I have created a dataset combining images without resizing from the 2020 and 2019 competitions [here](https://www.kaggle.com/tahsin/cassava-leaf-disease-merged). I found 722 duplicate images using [this](https://www.kaggle.com/zzy990106/duplicate-images-in-two-competitions) notebook and removed them. The `merged.csv` contains all images, corresponding labels, and source(2019/2020). ",
    "1097831": "Cool, thank you!\nI wonder why the organizers didn't include all previous images in the current dataset. 😯",
    "1123426": "Anyone's LB improved using the dataset?\n",
    "1095394": "I tried using previous dataset but it reduced my LB by 0.001. I didn't remove duplicates... That might be the reason.",
    "1101181": "",
    "1095302": "Thank you, thats very useful."
  }
}