{
  "id": 70917,
  "title": "Help a newbie - How to balance train and test sets properly?",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/70917",
  "author_name": "",
  "post_date": "2018-11-08T11:17:42.095414800Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello there :)</p>\n\n<p>In this competition, more than in others, balancing the train and test sets evenly seems very important due to the scarcity of some classes.</p>\n\n<p>So I ask the experienced users:</p>\n\n<blockquote>\n  <p>How do you balance your datasets properly?</p>\n</blockquote>",
  "messages": [
    {
      "id": "417494",
      "postDate": "11/08/2018 11:17:42",
      "content": "<p>Hello there :)</p>\n\n<p>In this competition, more than in others, balancing the train and test sets evenly seems very important due to the scarcity of some classes.</p>\n\n<p>So I ask the experienced users:</p>\n\n<blockquote>\n  <p>How do you balance your datasets properly?</p>\n</blockquote>",
      "rawMarkdown": "Hello there :)\n\nIn this competition, more than in others, balancing the train and test sets evenly seems very important due to the scarcity of some classes.\n\nSo I ask the experienced users:\n\n&gt; How do you balance your datasets properly?",
      "votes": null
    },
    {
      "id": "418388",
      "postDate": "11/09/2018 20:25:51",
      "content": "<p>I'm using 5 fold CV with no stratification and at the moment not training for all folds.</p>",
      "rawMarkdown": "I'm using 5 fold CV with no stratification and at the moment not training for all folds.",
      "votes": null
    },
    {
      "id": "424787",
      "postDate": "11/20/2018 16:49:57",
      "content": "<p>Is there a reason for not using stratification ? </p>",
      "rawMarkdown": "Is there a reason for not using stratification ?",
      "votes": null
    },
    {
      "id": "424843",
      "postDate": "11/20/2018 18:48:42",
      "content": "<p>I tend to favour folds that are different and feel more comfortable if my model improves on all folds. If all your folds are stratified in the same manner you are not gaining much information in terms of assessing out of sample performance (what you care about) by assessing on k-similar folds. Note for k-fold the training data sharing between any 2 models is k-2 folds and so it's likely we'll reach similar models in weight space - you're just getting a more robust/higher entropy assessment of out of sample model performance in my view.</p>",
      "rawMarkdown": "I tend to favour folds that are different and feel more comfortable if my model improves on all folds. If all your folds are stratified in the same manner you are not gaining much information in terms of assessing out of sample performance (what you care about) by assessing on k-similar folds. Note for k-fold the training data sharing between any 2 models is k-2 folds and so it's likely we'll reach similar models in weight space - you're just getting a more robust/higher entropy assessment of out of sample model performance in my view.",
      "votes": null
    },
    {
      "id": "424930",
      "postDate": "11/20/2018 22:28:02",
      "content": "<p>Without stratification, the std of the loss will get higher, right? I do not quite understand the reason of getting a higher std instead of a lower one. Or maybe un-stratified models perform better after ensembling? Is that the reason?\n<a href=\"/maw501\">@maw501</a></p>",
      "rawMarkdown": "Without stratification, the std of the loss will get higher, right? I do not quite understand the reason of getting a higher std instead of a lower one. Or maybe un-stratified models perform better after ensembling? Is that the reason?\n@maw501",
      "votes": null
    },
    {
      "id": "425162",
      "postDate": "11/21/2018 08:09:48",
      "content": "<p>The point of your CV is to give you (as much as possible) an unbiased estimate of out of sample performance. In the example of 5-fold you have a choice to use 5 validation sets that are all stratified in the same manner or to use validation sets with more variation. My view is that the latter gives more robust estimates of true out of sample performance and this is generally a reason for favouring a higher entropy validation scheme.</p>\n\n<p>And just to clarify a common misunderstanding: in 5-fold any two models share 3 of the same folds for training out of 4 i.e. 75% of the same data and it's usually the case the models are of similar quality. The fact your validation loss is worse doesn't mean that model is necessarily worse - it is usually due to variation in the validation sets. And getting a sense of how volatile your out of sample performance is is key to generalizing well.</p>",
      "rawMarkdown": "The point of your CV is to give you (as much as possible) an unbiased estimate of out of sample performance. In the example of 5-fold you have a choice to use 5 validation sets that are all stratified in the same manner or to use validation sets with more variation. My view is that the latter gives more robust estimates of true out of sample performance and this is generally a reason for favouring a higher entropy validation scheme.\n\nAnd just to clarify a common misunderstanding: in 5-fold any two models share 3 of the same folds for training out of 4 i.e. 75% of the same data and it's usually the case the models are of similar quality. The fact your validation loss is worse doesn't mean that model is necessarily worse - it is usually due to variation in the validation sets. And getting a sense of how volatile your out of sample performance is is key to generalizing well.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 418388,
      "author_name": "maw501",
      "author_url": "",
      "post_date": "11/09/2018 20:25:51",
      "content": "<p>I'm using 5 fold CV with no stratification and at the moment not training for all folds.</p>",
      "votes": null,
      "replies": [
        {
          "id": 424787,
          "author_name": "manuscrits",
          "author_url": "",
          "post_date": "11/20/2018 16:49:57",
          "content": "<p>Is there a reason for not using stratification ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 424843,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "11/20/2018 18:48:42",
          "content": "<p>I tend to favour folds that are different and feel more comfortable if my model improves on all folds. If all your folds are stratified in the same manner you are not gaining much information in terms of assessing out of sample performance (what you care about) by assessing on k-similar folds. Note for k-fold the training data sharing between any 2 models is k-2 folds and so it's likely we'll reach similar models in weight space - you're just getting a more robust/higher entropy assessment of out of sample model performance in my view.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 424930,
          "author_name": "kokecacao",
          "author_url": "",
          "post_date": "11/20/2018 22:28:02",
          "content": "<p>Without stratification, the std of the loss will get higher, right? I do not quite understand the reason of getting a higher std instead of a lower one. Or maybe un-stratified models perform better after ensembling? Is that the reason?\n<a href=\"/maw501\">@maw501</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 425162,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "11/21/2018 08:09:48",
          "content": "<p>The point of your CV is to give you (as much as possible) an unbiased estimate of out of sample performance. In the example of 5-fold you have a choice to use 5 validation sets that are all stratified in the same manner or to use validation sets with more variation. My view is that the latter gives more robust estimates of true out of sample performance and this is generally a reason for favouring a higher entropy validation scheme.</p>\n\n<p>And just to clarify a common misunderstanding: in 5-fold any two models share 3 of the same folds for training out of 4 i.e. 75% of the same data and it's usually the case the models are of similar quality. The fact your validation loss is worse doesn't mean that model is necessarily worse - it is usually due to variation in the validation sets. And getting a sense of how volatile your out of sample performance is is key to generalizing well.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "417494": "Hello there :)\n\nIn this competition, more than in others, balancing the train and test sets evenly seems very important due to the scarcity of some classes.\n\nSo I ask the experienced users:\n\n&gt; How do you balance your datasets properly?",
    "418388": "I'm using 5 fold CV with no stratification and at the moment not training for all folds.",
    "424787": "Is there a reason for not using stratification ?",
    "424843": "I tend to favour folds that are different and feel more comfortable if my model improves on all folds. If all your folds are stratified in the same manner you are not gaining much information in terms of assessing out of sample performance (what you care about) by assessing on k-similar folds. Note for k-fold the training data sharing between any 2 models is k-2 folds and so it's likely we'll reach similar models in weight space - you're just getting a more robust/higher entropy assessment of out of sample model performance in my view.",
    "424930": "Without stratification, the std of the loss will get higher, right? I do not quite understand the reason of getting a higher std instead of a lower one. Or maybe un-stratified models perform better after ensembling? Is that the reason?\n@maw501",
    "425162": "The point of your CV is to give you (as much as possible) an unbiased estimate of out of sample performance. In the example of 5-fold you have a choice to use 5 validation sets that are all stratified in the same manner or to use validation sets with more variation. My view is that the latter gives more robust estimates of true out of sample performance and this is generally a reason for favouring a higher entropy validation scheme.\n\nAnd just to clarify a common misunderstanding: in 5-fold any two models share 3 of the same folds for training out of 4 i.e. 75% of the same data and it's usually the case the models are of similar quality. The fact your validation loss is worse doesn't mean that model is necessarily worse - it is usually due to variation in the validation sets. And getting a sense of how volatile your out of sample performance is is key to generalizing well."
  },
  "source": "meta"
}