{
  "id": 102826,
  "title": "How properly do class balancing?",
  "url": "/competitions/aptos2019-blindness-detection/discussion/102826",
  "author_name": "",
  "post_date": "2019-08-05T11:00:07.589690300Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I tried to equalize sizes of all classes just by sampling images of small classes with augmentation. But it made my score worse. How properly do class balancing? Maybe it's necessary to save proportion between classes or something like it?</p>",
  "messages": [
    {
      "id": "592478",
      "postDate": "08/05/2019 11:00:07",
      "content": "<p>I tried to equalize sizes of all classes just by sampling images of small classes with augmentation. But it made my score worse. How properly do class balancing? Maybe it's necessary to save proportion between classes or something like it?</p>",
      "rawMarkdown": "I tried to equalize sizes of all classes just by sampling images of small classes with augmentation. But it made my score worse. How properly do class balancing? Maybe it's necessary to save proportion between classes or something like it?",
      "votes": null
    },
    {
      "id": "592554",
      "postDate": "08/05/2019 13:21:42",
      "content": "<p>Augmentation is just an image but shown in different ways. Imagine a wave and you just shift it or slightly modified it, without drastic transformation it basically just the same image. With enough training time, model will be used to it and those images are just as good as duplicates.\nWith dataset so imbalance for example 27k 0-label images and 1k 4-label images, up-sampling up to 27k 4-labels means each picture has 27 duplicates! So the model will overtrain on those small classes. And we can all agree duplicates are not good for training.\nAlso not using transformation on 0-label images means if model encounters a new 0-label images but slightly transformed, it will fail to predict.</p>",
      "rawMarkdown": "Augmentation is just an image but shown in different ways. Imagine a wave and you just shift it or slightly modified it, without drastic transformation it basically just the same image. With enough training time, model will be used to it and those images are just as good as duplicates.\nWith dataset so imbalance for example 27k 0-label images and 1k 4-label images, up-sampling up to 27k 4-labels means each picture has 27 duplicates! So the model will overtrain on those small classes. And we can all agree duplicates are not good for training.\nAlso not using transformation on 0-label images means if model encounters a new 0-label images but slightly transformed, it will fail to predict.",
      "votes": null
    },
    {
      "id": "592929",
      "postDate": "08/06/2019 01:34:21",
      "content": "<p>So you mean that the only way to do class balancing properly is to sample minor classes from previous data competition?</p>",
      "rawMarkdown": "So you mean that the only way to do class balancing properly is to sample minor classes from previous data competition?",
      "votes": null
    },
    {
      "id": "592938",
      "postDate": "08/06/2019 02:06:38",
      "content": "<p>No, what I meant it is not effective to sample minor classes to balance the data. If it makes your score worse then don't do sampling at all. I personally don't do class balancing.</p>",
      "rawMarkdown": "No, what I meant it is not effective to sample minor classes to balance the data. If it makes your score worse then don't do sampling at all. I personally don't do class balancing.",
      "votes": null
    },
    {
      "id": "593011",
      "postDate": "08/06/2019 05:26:08",
      "content": "<p>No need to do class balancing, the old data is enough and introduces enough variance into the system that even if you dont use dropout your model should generalize better(that is if your model is small enough, for me this has worked on models up to 20M parameters)</p>",
      "rawMarkdown": "No need to do class balancing, the old data is enough and introduces enough variance into the system that even if you dont use dropout your model should generalize better(that is if your model is small enough, for me this has worked on models up to 20M parameters)",
      "votes": null
    },
    {
      "id": "594029",
      "postDate": "08/07/2019 13:09:05",
      "content": "<p>Such small parameters make it to prevent overfitting?</p>",
      "rawMarkdown": "Such small parameters make it to prevent overfitting?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 592554,
      "author_name": "quandapro",
      "author_url": "",
      "post_date": "08/05/2019 13:21:42",
      "content": "<p>Augmentation is just an image but shown in different ways. Imagine a wave and you just shift it or slightly modified it, without drastic transformation it basically just the same image. With enough training time, model will be used to it and those images are just as good as duplicates.\nWith dataset so imbalance for example 27k 0-label images and 1k 4-label images, up-sampling up to 27k 4-labels means each picture has 27 duplicates! So the model will overtrain on those small classes. And we can all agree duplicates are not good for training.\nAlso not using transformation on 0-label images means if model encounters a new 0-label images but slightly transformed, it will fail to predict.</p>",
      "votes": null,
      "replies": [
        {
          "id": 592929,
          "author_name": "talalaev",
          "author_url": "",
          "post_date": "08/06/2019 01:34:21",
          "content": "<p>So you mean that the only way to do class balancing properly is to sample minor classes from previous data competition?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592938,
          "author_name": "quandapro",
          "author_url": "",
          "post_date": "08/06/2019 02:06:38",
          "content": "<p>No, what I meant it is not effective to sample minor classes to balance the data. If it makes your score worse then don't do sampling at all. I personally don't do class balancing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 593011,
      "author_name": "sidhanthholalkere",
      "author_url": "",
      "post_date": "08/06/2019 05:26:08",
      "content": "<p>No need to do class balancing, the old data is enough and introduces enough variance into the system that even if you dont use dropout your model should generalize better(that is if your model is small enough, for me this has worked on models up to 20M parameters)</p>",
      "votes": null,
      "replies": [
        {
          "id": 594029,
          "author_name": "a18974761777",
          "author_url": "",
          "post_date": "08/07/2019 13:09:05",
          "content": "<p>Such small parameters make it to prevent overfitting?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "592478": "I tried to equalize sizes of all classes just by sampling images of small classes with augmentation. But it made my score worse. How properly do class balancing? Maybe it's necessary to save proportion between classes or something like it?",
    "592554": "Augmentation is just an image but shown in different ways. Imagine a wave and you just shift it or slightly modified it, without drastic transformation it basically just the same image. With enough training time, model will be used to it and those images are just as good as duplicates.\nWith dataset so imbalance for example 27k 0-label images and 1k 4-label images, up-sampling up to 27k 4-labels means each picture has 27 duplicates! So the model will overtrain on those small classes. And we can all agree duplicates are not good for training.\nAlso not using transformation on 0-label images means if model encounters a new 0-label images but slightly transformed, it will fail to predict.",
    "592929": "So you mean that the only way to do class balancing properly is to sample minor classes from previous data competition?",
    "592938": "No, what I meant it is not effective to sample minor classes to balance the data. If it makes your score worse then don't do sampling at all. I personally don't do class balancing.",
    "593011": "No need to do class balancing, the old data is enough and introduces enough variance into the system that even if you dont use dropout your model should generalize better(that is if your model is small enough, for me this has worked on models up to 20M parameters)",
    "594029": "Such small parameters make it to prevent overfitting?"
  },
  "source": "meta"
}