{
  "id": 201699,
  "title": "Rethinking CV Strategy : StratifiedGroup-Kfold",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/201699",
  "author_name": "Mr_KnowNothing",
  "post_date": "2020-12-06T09:35:18.695000",
  "votes": 13,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hello everyone , After my EDA and complete Data Understanding , its now time for me to move to modelling , and the first step is on deciding reliable CV . I have seen people using StratifiedKfold and its fairly reasonable thing to do given the data . </p>\n<p>But in my previous notebook <a href=\"https://www.kaggle.com/tanulsingh077/how-to-become-leaf-doctor-with-deep-learning\" target=\"_blank\">here</a> I showed that there are similarity in Images across different Labels and also some mislabels present . This idea is the outcome of that same fact</p>\n<p>I argue that given the fact , images corresponding to different labels have similarities among them , dataset have mislabeled images , it would be better to use GroupStratified-Kfold .</p>\n<p>Now the question is from where do we get the groups? We get the groups by clustering the images </p>\n<p><a href=\"https://www.kaggle.com/tanulsingh077/rethinking-cv-strategy-with-visualizations\" target=\"_blank\">Here</a> In this notebook I analyze the clusters and this CV strategy . It would be great to know community's thought on that</p>",
  "messages": [
    {
      "id": 1103788,
      "postDate": "2020-12-06T09:35:18.697Z",
      "content": "<p>Hello everyone , After my EDA and complete Data Understanding , its now time for me to move to modelling , and the first step is on deciding reliable CV . I have seen people using StratifiedKfold and its fairly reasonable thing to do given the data . </p>\n<p>But in my previous notebook <a href=\"https://www.kaggle.com/tanulsingh077/how-to-become-leaf-doctor-with-deep-learning\" target=\"_blank\">here</a> I showed that there are similarity in Images across different Labels and also some mislabels present . This idea is the outcome of that same fact</p>\n<p>I argue that given the fact , images corresponding to different labels have similarities among them , dataset have mislabeled images , it would be better to use GroupStratified-Kfold .</p>\n<p>Now the question is from where do we get the groups? We get the groups by clustering the images </p>\n<p><a href=\"https://www.kaggle.com/tanulsingh077/rethinking-cv-strategy-with-visualizations\" target=\"_blank\">Here</a> In this notebook I analyze the clusters and this CV strategy . It would be great to know community's thought on that</p>",
      "rawMarkdown": "Hello everyone , After my EDA and complete Data Understanding , its now time for me to move to modelling , and the first step is on deciding reliable CV . I have seen people using StratifiedKfold and its fairly reasonable thing to do given the data . \n\nBut in my previous notebook [here](https://www.kaggle.com/tanulsingh077/how-to-become-leaf-doctor-with-deep-learning) I showed that there are similarity in Images across different Labels and also some mislabels present . This idea is the outcome of that same fact\n\nI argue that given the fact , images corresponding to different labels have similarities among them , dataset have mislabeled images , it would be better to use GroupStratified-Kfold .\n\nNow the question is from where do we get the groups? We get the groups by clustering the images \n\n[Here](https://www.kaggle.com/tanulsingh077/rethinking-cv-strategy-with-visualizations) In this notebook I analyze the clusters and this CV strategy . It would be great to know community's thought on that",
      "votes": 13
    },
    {
      "id": 1104323,
      "postDate": "2020-12-06T20:23:39.870Z",
      "content": "<p>Great job on the analysis, but I'm not sure this approach will work here - the test data might have same issues around mislabelling (and why wouldn't it?).</p>",
      "rawMarkdown": "Great job on the analysis, but I'm not sure this approach will work here - the test data might have same issues around mislabelling (and why wouldn't it?).",
      "votes": 1,
      "replies": [
        {
          "id": 1104555,
          "postDate": "2020-12-07T04:13:19.477Z",
          "content": "<p>Thanks for the comment <a href=\"https://www.kaggle.com/konradb\" target=\"_blank\">@konradb</a>, I agree that the test data might contain the same mislabels but doing group stratified kfold makes sure that most of the mislabelled examples lie in the same group and thus the model could perform better doesn't it? </p>",
          "rawMarkdown": "Thanks for the comment @konradb, I agree that the test data might contain the same mislabels but doing group stratified kfold makes sure that most of the mislabelled examples lie in the same group and thus the model could perform better doesn't it? "
        },
        {
          "id": 1106519,
          "postDate": "2020-12-08T22:41:17.937Z",
          "content": "<blockquote>\n  <p>doing group stratified kfold makes sure that most of the mislabelled examples lie in the same group and thus the model could perform better doesn't it?</p>\n</blockquote>\n<p>That is the place where you might be wrong, there are times when noise in one place can actually have model to learn more from the noise and not the true images. Do I really have to explain why? Isnt that obvious given what data we have?</p>",
          "rawMarkdown": ">doing group stratified kfold makes sure that most of the mislabelled examples lie in the same group and thus the model could perform better doesn't it?\n\nThat is the place where you might be wrong, there are times when noise in one place can actually have model to learn more from the noise and not the true images. Do I really have to explain why? Isnt that obvious given what data we have?",
          "votes": -1
        },
        {
          "id": 1106880,
          "postDate": "2020-12-09T07:40:21.597Z",
          "content": "<p>I just understand that you say 'we can get better model from mislabeled (noise) data'.</p>",
          "rawMarkdown": "I just understand that you say 'we can get better model from mislabeled (noise) data'."
        },
        {
          "id": 1106908,
          "postDate": "2020-12-09T08:02:33.160Z",
          "content": "<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> What I meant was if all the noise is in the group , then the model might learn better that even if these are noisy , they belong to the x class and not the y ones. <br>\nAnd again I am not sure , I just wanted to put this idea forward to know everyone's thought</p>",
          "rawMarkdown": "@harshitsheoran What I meant was if all the noise is in the group , then the model might learn better that even if these are noisy , they belong to the x class and not the y ones. \nAnd again I am not sure , I just wanted to put this idea forward to know everyone's thought"
        },
        {
          "id": 1106931,
          "postDate": "2020-12-09T08:25:08.483Z",
          "content": "<p><a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> </p>\n<p>What I said is if noise was in a group, the chances of model learning better is pretty slim.</p>\n<p>Explanation:-</p>\n<p>We use augmentations like shift scale rotate and randomresizedcrop, there will be images with more leaves (more leaves means more places to classify features) in some images and less in others, if that number of leaves just decreases in true data by randomness which is very likely considering there is more data in true so more images will be low number of leaves.</p>\n<p>That will likely to cause model learn features from one class in another. You can try it out but it is highly likely to be the case in this dataset.</p>\n<p>Also, <a href=\"https://www.kaggle.com/jinkyh\" target=\"_blank\">@jinkyh</a> </p>\n<p>That is why the model can actually learn wrong from noise, I did not say it will get better as learning wrong things from noise will actually make model worse.</p>",
          "rawMarkdown": "@tanulsingh077 \n\nWhat I said is if noise was in a group, the chances of model learning better is pretty slim.\n\nExplanation:-\n\nWe use augmentations like shift scale rotate and randomresizedcrop, there will be images with more leaves (more leaves means more places to classify features) in some images and less in others, if that number of leaves just decreases in true data by randomness which is very likely considering there is more data in true so more images will be low number of leaves.\n\nThat will likely to cause model learn features from one class in another. You can try it out but it is highly likely to be the case in this dataset.\n\nAlso, @jinkyh \n\nThat is why the model can actually learn wrong from noise, I did not say it will get better as learning wrong things from noise will actually make model worse.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1104504,
      "postDate": "2020-12-07T02:13:26.840Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1104323,
      "author_name": "Konrad Banachewicz",
      "author_url": "",
      "post_date": "2020-12-06T20:23:39.870000",
      "content": "<p>Great job on the analysis, but I'm not sure this approach will work here - the test data might have same issues around mislabelling (and why wouldn't it?).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1104555,
          "author_name": "Mr_KnowNothing",
          "author_url": "",
          "post_date": "2020-12-07T04:13:19.477000",
          "content": "<p>Thanks for the comment <a href=\"https://www.kaggle.com/konradb\" target=\"_blank\">@konradb</a>, I agree that the test data might contain the same mislabels but doing group stratified kfold makes sure that most of the mislabelled examples lie in the same group and thus the model could perform better doesn't it? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1106519,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-12-08T22:41:17.937000",
          "content": "<blockquote>\n  <p>doing group stratified kfold makes sure that most of the mislabelled examples lie in the same group and thus the model could perform better doesn't it?</p>\n</blockquote>\n<p>That is the place where you might be wrong, there are times when noise in one place can actually have model to learn more from the noise and not the true images. Do I really have to explain why? Isnt that obvious given what data we have?</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1106880,
          "author_name": "Kim Yeong Hyeon",
          "author_url": "",
          "post_date": "2020-12-09T07:40:21.597000",
          "content": "<p>I just understand that you say 'we can get better model from mislabeled (noise) data'.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1106908,
          "author_name": "Mr_KnowNothing",
          "author_url": "",
          "post_date": "2020-12-09T08:02:33.160000",
          "content": "<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> What I meant was if all the noise is in the group , then the model might learn better that even if these are noisy , they belong to the x class and not the y ones. <br>\nAnd again I am not sure , I just wanted to put this idea forward to know everyone's thought</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1106931,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-12-09T08:25:08.483000",
          "content": "<p><a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> </p>\n<p>What I said is if noise was in a group, the chances of model learning better is pretty slim.</p>\n<p>Explanation:-</p>\n<p>We use augmentations like shift scale rotate and randomresizedcrop, there will be images with more leaves (more leaves means more places to classify features) in some images and less in others, if that number of leaves just decreases in true data by randomness which is very likely considering there is more data in true so more images will be low number of leaves.</p>\n<p>That will likely to cause model learn features from one class in another. You can try it out but it is highly likely to be the case in this dataset.</p>\n<p>Also, <a href=\"https://www.kaggle.com/jinkyh\" target=\"_blank\">@jinkyh</a> </p>\n<p>That is why the model can actually learn wrong from noise, I did not say it will get better as learning wrong things from noise will actually make model worse.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1104504,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-07T02:13:26.840000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1103788": "Hello everyone , After my EDA and complete Data Understanding , its now time for me to move to modelling , and the first step is on deciding reliable CV . I have seen people using StratifiedKfold and its fairly reasonable thing to do given the data . \n\nBut in my previous notebook [here](https://www.kaggle.com/tanulsingh077/how-to-become-leaf-doctor-with-deep-learning) I showed that there are similarity in Images across different Labels and also some mislabels present . This idea is the outcome of that same fact\n\nI argue that given the fact , images corresponding to different labels have similarities among them , dataset have mislabeled images , it would be better to use GroupStratified-Kfold .\n\nNow the question is from where do we get the groups? We get the groups by clustering the images \n\n[Here](https://www.kaggle.com/tanulsingh077/rethinking-cv-strategy-with-visualizations) In this notebook I analyze the clusters and this CV strategy . It would be great to know community's thought on that",
    "1104323": "Great job on the analysis, but I'm not sure this approach will work here - the test data might have same issues around mislabelling (and why wouldn't it?).",
    "1104504": ""
  }
}