{
  "id": 80661,
  "title": "Use of full training set vs class rebalancing vs class rebalancing + augmentations?",
  "url": "/competitions/histopathologic-cancer-detection/discussion/80661",
  "author_name": "",
  "post_date": "2019-02-15T08:01:26.358315100Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi, I've noticed that some Kernels use all the images (and even add augmentations), while others choose to rebalance the distributions of the classes so that there is an even split between tumours and non-tumours for both the training set and validation set. </p>\n\n<p>Could people please comment why there would be an improved performance for using a balanced set? I would have thought that more training data = better performance. Or you could use augmentations to bulk up whichever class is smaller so that you have an even distribution between the classes. </p>",
  "messages": [
    {
      "id": "472009",
      "postDate": "02/15/2019 08:01:26",
      "content": "<p>Hi, I've noticed that some Kernels use all the images (and even add augmentations), while others choose to rebalance the distributions of the classes so that there is an even split between tumours and non-tumours for both the training set and validation set. </p>\n\n<p>Could people please comment why there would be an improved performance for using a balanced set? I would have thought that more training data = better performance. Or you could use augmentations to bulk up whichever class is smaller so that you have an even distribution between the classes. </p>",
      "rawMarkdown": "Hi, I've noticed that some Kernels use all the images (and even add augmentations), while others choose to rebalance the distributions of the classes so that there is an even split between tumours and non-tumours for both the training set and validation set. \n\nCould people please comment why there would be an improved performance for using a balanced set? I would have thought that more training data = better performance. Or you could use augmentations to bulk up whichever class is smaller so that you have an even distribution between the classes.",
      "votes": null
    },
    {
      "id": "472273",
      "postDate": "02/15/2019 16:04:42",
      "content": "<p>Hi,\nI am not sure about the other approaches but for me using the balanced data did not lead to better performance. </p>",
      "rawMarkdown": "Hi,\nI am not sure about the other approaches but for me using the balanced data did not lead to better performance.",
      "votes": null
    },
    {
      "id": "473422",
      "postDate": "02/18/2019 01:22:12",
      "content": "<p>It depends on how one rebalances the dataset... Preferrably, one should not touch the validation set and keep it intact...</p>\n\n<p>But it anyway seems that the dataset is pretty well-balanced to begin with so I don't think there would be a need to worry about rebalancing</p>",
      "rawMarkdown": "It depends on how one rebalances the dataset... Preferrably, one should not touch the validation set and keep it intact...\n\nBut it anyway seems that the dataset is pretty well-balanced to begin with so I don't think there would be a need to worry about rebalancing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 472273,
      "author_name": "ipateam",
      "author_url": "",
      "post_date": "02/15/2019 16:04:42",
      "content": "<p>Hi,\nI am not sure about the other approaches but for me using the balanced data did not lead to better performance. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 473422,
      "author_name": "tanlikesmath",
      "author_url": "",
      "post_date": "02/18/2019 01:22:12",
      "content": "<p>It depends on how one rebalances the dataset... Preferrably, one should not touch the validation set and keep it intact...</p>\n\n<p>But it anyway seems that the dataset is pretty well-balanced to begin with so I don't think there would be a need to worry about rebalancing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "472009": "Hi, I've noticed that some Kernels use all the images (and even add augmentations), while others choose to rebalance the distributions of the classes so that there is an even split between tumours and non-tumours for both the training set and validation set. \n\nCould people please comment why there would be an improved performance for using a balanced set? I would have thought that more training data = better performance. Or you could use augmentations to bulk up whichever class is smaller so that you have an even distribution between the classes.",
    "472273": "Hi,\nI am not sure about the other approaches but for me using the balanced data did not lead to better performance.",
    "473422": "It depends on how one rebalances the dataset... Preferrably, one should not touch the validation set and keep it intact...\n\nBut it anyway seems that the dataset is pretty well-balanced to begin with so I don't think there would be a need to worry about rebalancing"
  },
  "source": "meta"
}