{
  "id": 212533,
  "title": "Adding data from previous Cassava competition",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/212533",
  "author_name": "",
  "post_date": "2021-01-19T09:06:06.598717600Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Did anyone try increasing the size of training data by adding data from the previous <a href=\"https://www.kaggle.com/c/cassava-disease\" target=\"_blank\">Cassava Disease Classification Competition</a>? <br>\nIf yes, what was the affect on the accuracy results of train set and validation set?</p>",
  "messages": [
    {
      "id": "1159449",
      "postDate": "01/19/2021 09:06:06",
      "content": "<p>Did anyone try increasing the size of training data by adding data from the previous <a href=\"https://www.kaggle.com/c/cassava-disease\" target=\"_blank\">Cassava Disease Classification Competition</a>? <br>\nIf yes, what was the affect on the accuracy results of train set and validation set?</p>",
      "rawMarkdown": "Did anyone try increasing the size of training data by adding data from the previous [Cassava Disease Classification Competition](https://www.kaggle.com/c/cassava-disease)? \nIf yes, what was the affect on the accuracy results of train set and validation set?",
      "votes": null
    },
    {
      "id": "1159490",
      "postDate": "01/19/2021 09:58:39",
      "content": "<p>The previous competition data has been used, but I didn’t see a significant increase in LB effect, and some people have also confirmed this.</p>",
      "rawMarkdown": "The previous competition data has been used, but I didn’t see a significant increase in LB effect, and some people have also confirmed this.",
      "votes": null
    },
    {
      "id": "1160198",
      "postDate": "01/19/2021 18:21:30",
      "content": "<p>I agree, I experienced the same. </p>",
      "rawMarkdown": "I agree, I experienced the same.",
      "votes": null
    },
    {
      "id": "1161896",
      "postDate": "01/20/2021 20:01:20",
      "content": "<p>Running an extended test now (4 different models on 4 machines).  Hope the models done by tomorrow but they might not all be completed training for another couple of days.</p>\n<p>I got to 43K total images with the following.</p>\n<ol>\n<li>Train set 2020 </li>\n<li>Train set 2019</li>\n<li>Extra images and test set 2019 - I created the \"truth\" by running an assembly of 12 keras models - several of the top scoring 2020 shared models and my top CV local models.  I believe the 2019 winner used a similar method to increase images for training using the extras.</li>\n</ol>\n<p>Other posts have indicated that 2020 contains several hundred duplicates from 2019.  One post suggested that some of the duplicates were from the extra images with the truth generated by the 2019 winning model.  2019 had number of different image sizes suggesting different cameras, etc.   I resized all the 2019 images to match the 800x600 size of 2020.</p>\n<p>43K is too many images to run in a single Kaggle kernel with enough epochs to see a difference (IMO).  A couple of the configurations involved some up sampling which really adds to the training time with 129K worth of images.</p>\n<p>I am using heavy augmentation in training.  </p>\n<p>I will probably add these four models to the assembly and establish a new truth for the extra images and the 2019 test set if the results look at all promising and run all 4 one more time.</p>",
      "rawMarkdown": "Running an extended test now (4 different models on 4 machines).  Hope the models done by tomorrow but they might not all be completed training for another couple of days.\n\nI got to 43K total images with the following.\n1.  Train set 2020 \n2.  Train set 2019\n3.  Extra images and test set 2019 - I created the \"truth\" by running an assembly of 12 keras models - several of the top scoring 2020 shared models and my top CV local models.  I believe the 2019 winner used a similar method to increase images for training using the extras.\n\nOther posts have indicated that 2020 contains several hundred duplicates from 2019.  One post suggested that some of the duplicates were from the extra images with the truth generated by the 2019 winning model.  2019 had number of different image sizes suggesting different cameras, etc.   I resized all the 2019 images to match the 800x600 size of 2020.\n\n43K is too many images to run in a single Kaggle kernel with enough epochs to see a difference (IMO).  A couple of the configurations involved some up sampling which really adds to the training time with 129K worth of images.\n\nI am using heavy augmentation in training.  \n\nI will probably add these four models to the assembly and establish a new truth for the extra images and the 2019 test set if the results look at all promising and run all 4 one more time.",
      "votes": null
    },
    {
      "id": "1164334",
      "postDate": "01/22/2021 09:59:49",
      "content": "<p>That all sounds really intriguing. Have your model trained by now? Would love to know the effect of adding those images on training part and CV part as well. If model has not completed its training by now, kindly let us know the results when it's done. Thanks.<br>\n<a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> </p>",
      "rawMarkdown": "That all sounds really intriguing. Have your model trained by now? Would love to know the effect of adding those images on training part and CV part as well. If model has not completed its training by now, kindly let us know the results when it's done. Thanks.\n@pcjimmmy",
      "votes": null
    },
    {
      "id": "1164638",
      "postDate": "01/22/2021 13:47:36",
      "content": "<p>First of 4 models completed training yesterday evening.  This model basic efficient B0 with a few additional layers.  </p>\n<p>Training accuracy was about 7% better than the same model with the base 2020 train set.  This was not a surprise as about 1/2 of the images had the Truth generated by the 12 model assembly.  If 12 models predict the truth one would expect that a model of sufficient size should always learn and have improved training accuracy.</p>\n<p>The validation cv was also better but I did nothing special to keep the 2019 model estimated truth images out.  Again with 1/2 the images having model predicted truth it makes sense that local cv should improve if you include those images.  I was doing a class stratified 5 fold.   Pretty sure it's beyond my skill set to generate a 5 fold and keep that half out.   So my future attempts will be simple validation set generated from the 2020 train set.  Also 5 fold with 43K takes too many hours of training time.</p>\n<p>The leaderboard submission was around 1/2 percent worse.  Based on this first model, I am expecting that the other three I have running will prove to be equally duds.</p>\n<p>Need to think about the nature of the 2019 images.  The winning accuracy values in 2019 I think are around 2% higher than our current LB top score.  I think I have seen mostly posts indicating no improvement when adding the base 2019 train set - somehow that fails to make sense in my mind.</p>\n<p>Will probably continue to play around for the remainder of this competition with ways to use the entire 2019 data set.  In the last day or two a kernel was shared using a threshold for the 2019 extra images to be used.  That will be my next effort once the remaining 3 models are completed.  Will report back if the results are any different.</p>\n<p><a href=\"https://www.kaggle.com/electro/pseudo-labeling-extra-images-from-2019\" target=\"_blank\">https://www.kaggle.com/electro/pseudo-labeling-extra-images-from-2019</a></p>",
      "rawMarkdown": "First of 4 models completed training yesterday evening.  This model basic efficient B0 with a few additional layers.  \n\nTraining accuracy was about 7% better than the same model with the base 2020 train set.  This was not a surprise as about 1/2 of the images had the Truth generated by the 12 model assembly.  If 12 models predict the truth one would expect that a model of sufficient size should always learn and have improved training accuracy.\n\nThe validation cv was also better but I did nothing special to keep the 2019 model estimated truth images out.  Again with 1/2 the images having model predicted truth it makes sense that local cv should improve if you include those images.  I was doing a class stratified 5 fold.   Pretty sure it's beyond my skill set to generate a 5 fold and keep that half out.   So my future attempts will be simple validation set generated from the 2020 train set.  Also 5 fold with 43K takes too many hours of training time.\n\nThe leaderboard submission was around 1/2 percent worse.  Based on this first model, I am expecting that the other three I have running will prove to be equally duds.\n\nNeed to think about the nature of the 2019 images.  The winning accuracy values in 2019 I think are around 2% higher than our current LB top score.  I think I have seen mostly posts indicating no improvement when adding the base 2019 train set - somehow that fails to make sense in my mind.\n\nWill probably continue to play around for the remainder of this competition with ways to use the entire 2019 data set.  In the last day or two a kernel was shared using a threshold for the 2019 extra images to be used.  That will be my next effort once the remaining 3 models are completed.  Will report back if the results are any different.\n\nhttps://www.kaggle.com/electro/pseudo-labeling-extra-images-from-2019",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1159490,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "01/19/2021 09:58:39",
      "content": "<p>The previous competition data has been used, but I didn’t see a significant increase in LB effect, and some people have also confirmed this.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1160198,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "01/19/2021 18:21:30",
          "content": "<p>I agree, I experienced the same. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1161896,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "01/20/2021 20:01:20",
      "content": "<p>Running an extended test now (4 different models on 4 machines).  Hope the models done by tomorrow but they might not all be completed training for another couple of days.</p>\n<p>I got to 43K total images with the following.</p>\n<ol>\n<li>Train set 2020 </li>\n<li>Train set 2019</li>\n<li>Extra images and test set 2019 - I created the \"truth\" by running an assembly of 12 keras models - several of the top scoring 2020 shared models and my top CV local models.  I believe the 2019 winner used a similar method to increase images for training using the extras.</li>\n</ol>\n<p>Other posts have indicated that 2020 contains several hundred duplicates from 2019.  One post suggested that some of the duplicates were from the extra images with the truth generated by the 2019 winning model.  2019 had number of different image sizes suggesting different cameras, etc.   I resized all the 2019 images to match the 800x600 size of 2020.</p>\n<p>43K is too many images to run in a single Kaggle kernel with enough epochs to see a difference (IMO).  A couple of the configurations involved some up sampling which really adds to the training time with 129K worth of images.</p>\n<p>I am using heavy augmentation in training.  </p>\n<p>I will probably add these four models to the assembly and establish a new truth for the extra images and the 2019 test set if the results look at all promising and run all 4 one more time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1164334,
          "author_name": "hammaadali",
          "author_url": "",
          "post_date": "01/22/2021 09:59:49",
          "content": "<p>That all sounds really intriguing. Have your model trained by now? Would love to know the effect of adding those images on training part and CV part as well. If model has not completed its training by now, kindly let us know the results when it's done. Thanks.<br>\n<a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1164638,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/22/2021 13:47:36",
          "content": "<p>First of 4 models completed training yesterday evening.  This model basic efficient B0 with a few additional layers.  </p>\n<p>Training accuracy was about 7% better than the same model with the base 2020 train set.  This was not a surprise as about 1/2 of the images had the Truth generated by the 12 model assembly.  If 12 models predict the truth one would expect that a model of sufficient size should always learn and have improved training accuracy.</p>\n<p>The validation cv was also better but I did nothing special to keep the 2019 model estimated truth images out.  Again with 1/2 the images having model predicted truth it makes sense that local cv should improve if you include those images.  I was doing a class stratified 5 fold.   Pretty sure it's beyond my skill set to generate a 5 fold and keep that half out.   So my future attempts will be simple validation set generated from the 2020 train set.  Also 5 fold with 43K takes too many hours of training time.</p>\n<p>The leaderboard submission was around 1/2 percent worse.  Based on this first model, I am expecting that the other three I have running will prove to be equally duds.</p>\n<p>Need to think about the nature of the 2019 images.  The winning accuracy values in 2019 I think are around 2% higher than our current LB top score.  I think I have seen mostly posts indicating no improvement when adding the base 2019 train set - somehow that fails to make sense in my mind.</p>\n<p>Will probably continue to play around for the remainder of this competition with ways to use the entire 2019 data set.  In the last day or two a kernel was shared using a threshold for the 2019 extra images to be used.  That will be my next effort once the remaining 3 models are completed.  Will report back if the results are any different.</p>\n<p><a href=\"https://www.kaggle.com/electro/pseudo-labeling-extra-images-from-2019\" target=\"_blank\">https://www.kaggle.com/electro/pseudo-labeling-extra-images-from-2019</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1159449": "Did anyone try increasing the size of training data by adding data from the previous [Cassava Disease Classification Competition](https://www.kaggle.com/c/cassava-disease)? \nIf yes, what was the affect on the accuracy results of train set and validation set?",
    "1159490": "The previous competition data has been used, but I didn’t see a significant increase in LB effect, and some people have also confirmed this.",
    "1160198": "I agree, I experienced the same.",
    "1161896": "Running an extended test now (4 different models on 4 machines).  Hope the models done by tomorrow but they might not all be completed training for another couple of days.\n\nI got to 43K total images with the following.\n1.  Train set 2020 \n2.  Train set 2019\n3.  Extra images and test set 2019 - I created the \"truth\" by running an assembly of 12 keras models - several of the top scoring 2020 shared models and my top CV local models.  I believe the 2019 winner used a similar method to increase images for training using the extras.\n\nOther posts have indicated that 2020 contains several hundred duplicates from 2019.  One post suggested that some of the duplicates were from the extra images with the truth generated by the 2019 winning model.  2019 had number of different image sizes suggesting different cameras, etc.   I resized all the 2019 images to match the 800x600 size of 2020.\n\n43K is too many images to run in a single Kaggle kernel with enough epochs to see a difference (IMO).  A couple of the configurations involved some up sampling which really adds to the training time with 129K worth of images.\n\nI am using heavy augmentation in training.  \n\nI will probably add these four models to the assembly and establish a new truth for the extra images and the 2019 test set if the results look at all promising and run all 4 one more time.",
    "1164334": "That all sounds really intriguing. Have your model trained by now? Would love to know the effect of adding those images on training part and CV part as well. If model has not completed its training by now, kindly let us know the results when it's done. Thanks.\n@pcjimmmy",
    "1164638": "First of 4 models completed training yesterday evening.  This model basic efficient B0 with a few additional layers.  \n\nTraining accuracy was about 7% better than the same model with the base 2020 train set.  This was not a surprise as about 1/2 of the images had the Truth generated by the 12 model assembly.  If 12 models predict the truth one would expect that a model of sufficient size should always learn and have improved training accuracy.\n\nThe validation cv was also better but I did nothing special to keep the 2019 model estimated truth images out.  Again with 1/2 the images having model predicted truth it makes sense that local cv should improve if you include those images.  I was doing a class stratified 5 fold.   Pretty sure it's beyond my skill set to generate a 5 fold and keep that half out.   So my future attempts will be simple validation set generated from the 2020 train set.  Also 5 fold with 43K takes too many hours of training time.\n\nThe leaderboard submission was around 1/2 percent worse.  Based on this first model, I am expecting that the other three I have running will prove to be equally duds.\n\nNeed to think about the nature of the 2019 images.  The winning accuracy values in 2019 I think are around 2% higher than our current LB top score.  I think I have seen mostly posts indicating no improvement when adding the base 2019 train set - somehow that fails to make sense in my mind.\n\nWill probably continue to play around for the remainder of this competition with ways to use the entire 2019 data set.  In the last day or two a kernel was shared using a threshold for the 2019 extra images to be used.  That will be my next effort once the remaining 3 models are completed.  Will report back if the results are any different.\n\nhttps://www.kaggle.com/electro/pseudo-labeling-extra-images-from-2019"
  },
  "source": "meta"
}