{
  "id": 214489,
  "title": "Merging 2019 competitions dataset by excluding label 3 (majority records)",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/214489",
  "author_name": "",
  "post_date": "2021-01-26T20:22:19.127270600Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I have read in the different discussion (though people have a different perspective) about merging the 2019 dataset in this competition.  Is this a good idea to merge only low label count data from the 2019 dataset rather than all data. In this process, there would be a little increase in data for those labels which has fewer records. I am thinking to start training with this strategy including <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/212347\" target=\"_blank\">epoch thresholding</a> and other different augmentations but looking for the best advice if anyone has done it. Thanks.</p>",
  "messages": [
    {
      "id": "1171437",
      "postDate": "01/26/2021 20:22:19",
      "content": "<p>I have read in the different discussion (though people have a different perspective) about merging the 2019 dataset in this competition.  Is this a good idea to merge only low label count data from the 2019 dataset rather than all data. In this process, there would be a little increase in data for those labels which has fewer records. I am thinking to start training with this strategy including <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/212347\" target=\"_blank\">epoch thresholding</a> and other different augmentations but looking for the best advice if anyone has done it. Thanks.</p>",
      "rawMarkdown": "I have read in the different discussion (though people have a different perspective) about merging the 2019 dataset in this competition.  Is this a good idea to merge only low label count data from the 2019 dataset rather than all data. In this process, there would be a little increase in data for those labels which has fewer records. I am thinking to start training with this strategy including [epoch thresholding](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/212347) and other different augmentations but looking for the best advice if anyone has done it. Thanks.",
      "votes": null
    },
    {
      "id": "1171638",
      "postDate": "01/27/2021 02:24:06",
      "content": "<p>IMHO, there is only one way to find out -- try it! </p>\n<p>(I do not see any fundamental problems with what you are suggesting.)</p>",
      "rawMarkdown": "IMHO, there is only one way to find out -- try it! \n\n(I do not see any fundamental problems with what you are suggesting.)",
      "votes": null
    },
    {
      "id": "1171769",
      "postDate": "01/27/2021 05:11:09",
      "content": "<p>I have done that but local cv didn't improve, even decreased a little. I also tried to add all 2019 data and local cv decreased too，Maybe the 2019 data is more noiser than 2020?😅</p>",
      "rawMarkdown": "I have done that but local cv didn't improve, even decreased a little. I also tried to add all 2019 data and local cv decreased too，Maybe the 2019 data is more noiser than 2020?😅",
      "votes": null
    },
    {
      "id": "1171875",
      "postDate": "01/27/2021 06:46:14",
      "content": "<p>I tried 2019+2020 dataset , but cv and pb decreased . I guess it has something to do with the image size of the 2019 dataset. some images in the 2019 dataset cannot be directly cropped to 512x512. finding a suitable RESIZE method may make 2019 dataset works</p>",
      "rawMarkdown": "I tried 2019+2020 dataset , but cv and pb decreased . I guess it has something to do with the image size of the 2019 dataset. some images in the 2019 dataset cannot be directly cropped to 512x512. finding a suitable RESIZE method may make 2019 dataset works",
      "votes": null
    },
    {
      "id": "1171912",
      "postDate": "01/27/2021 07:04:50",
      "content": "<p>I tried 2019+2020 dataset ,sigle model resnext101_0.898-0.900</p>",
      "rawMarkdown": "I tried 2019+2020 dataset ,sigle model resnext101_0.898-0.900",
      "votes": null
    },
    {
      "id": "1172031",
      "postDate": "01/27/2021 08:20:56",
      "content": "<p><a href=\"https://www.kaggle.com/dbwlalagaga\" target=\"_blank\">@dbwlalagaga</a> That‘s a great insight! I'll check my augmentation and give it a try. Thanks a lot</p>",
      "rawMarkdown": "dbwlalagaga That‘s a great insight! I'll check my augmentation and give it a try. Thanks a lot",
      "votes": null
    },
    {
      "id": "1172303",
      "postDate": "01/27/2021 10:43:29",
      "content": "<p>I did it with a small tweak. First, resize all images to the desired dimension and then crop the required sizes. I did not face any cropping issues after that. </p>",
      "rawMarkdown": "I did it with a small tweak. First, resize all images to the desired dimension and then crop the required sizes. I did not face any cropping issues after that.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1171638,
      "author_name": "graf10a",
      "author_url": "",
      "post_date": "01/27/2021 02:24:06",
      "content": "<p>IMHO, there is only one way to find out -- try it! </p>\n<p>(I do not see any fundamental problems with what you are suggesting.)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1171769,
      "author_name": "kewang777",
      "author_url": "",
      "post_date": "01/27/2021 05:11:09",
      "content": "<p>I have done that but local cv didn't improve, even decreased a little. I also tried to add all 2019 data and local cv decreased too，Maybe the 2019 data is more noiser than 2020?😅</p>",
      "votes": null,
      "replies": [
        {
          "id": 1171875,
          "author_name": "dbwlalagaga",
          "author_url": "",
          "post_date": "01/27/2021 06:46:14",
          "content": "<p>I tried 2019+2020 dataset , but cv and pb decreased . I guess it has something to do with the image size of the 2019 dataset. some images in the 2019 dataset cannot be directly cropped to 512x512. finding a suitable RESIZE method may make 2019 dataset works</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1172031,
          "author_name": "kewang777",
          "author_url": "",
          "post_date": "01/27/2021 08:20:56",
          "content": "<p><a href=\"https://www.kaggle.com/dbwlalagaga\" target=\"_blank\">@dbwlalagaga</a> That‘s a great insight! I'll check my augmentation and give it a try. Thanks a lot</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1172303,
          "author_name": "saurabh2mishra",
          "author_url": "",
          "post_date": "01/27/2021 10:43:29",
          "content": "<p>I did it with a small tweak. First, resize all images to the desired dimension and then crop the required sizes. I did not face any cropping issues after that. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1171912,
      "author_name": "darknesszx",
      "author_url": "",
      "post_date": "01/27/2021 07:04:50",
      "content": "<p>I tried 2019+2020 dataset ,sigle model resnext101_0.898-0.900</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1171437": "I have read in the different discussion (though people have a different perspective) about merging the 2019 dataset in this competition.  Is this a good idea to merge only low label count data from the 2019 dataset rather than all data. In this process, there would be a little increase in data for those labels which has fewer records. I am thinking to start training with this strategy including [epoch thresholding](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/212347) and other different augmentations but looking for the best advice if anyone has done it. Thanks.",
    "1171638": "IMHO, there is only one way to find out -- try it! \n\n(I do not see any fundamental problems with what you are suggesting.)",
    "1171769": "I have done that but local cv didn't improve, even decreased a little. I also tried to add all 2019 data and local cv decreased too，Maybe the 2019 data is more noiser than 2020?😅",
    "1171875": "I tried 2019+2020 dataset , but cv and pb decreased . I guess it has something to do with the image size of the 2019 dataset. some images in the 2019 dataset cannot be directly cropped to 512x512. finding a suitable RESIZE method may make 2019 dataset works",
    "1171912": "I tried 2019+2020 dataset ,sigle model resnext101_0.898-0.900",
    "1172031": "dbwlalagaga That‘s a great insight! I'll check my augmentation and give it a try. Thanks a lot",
    "1172303": "I did it with a small tweak. First, resize all images to the desired dimension and then crop the required sizes. I did not face any cropping issues after that."
  },
  "source": "meta"
}