{
  "id": 99079,
  "title": "Data Balancing?",
  "url": "/competitions/aptos2019-blindness-detection/discussion/99079",
  "author_name": "",
  "post_date": "2019-07-08T17:18:42.322849900Z",
  "votes": -1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Did anyone try balancing the dataset such that the model sees enough examples of every label?\nOne way to do this could be to use data from <a href=\"https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized\">here</a> and combine them with the training set such that each labels have enough examples</p>",
  "messages": [
    {
      "id": "570708",
      "postDate": "07/08/2019 17:18:42",
      "content": "<p>Did anyone try balancing the dataset such that the model sees enough examples of every label?\nOne way to do this could be to use data from <a href=\"https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized\">here</a> and combine them with the training set such that each labels have enough examples</p>",
      "rawMarkdown": "Did anyone try balancing the dataset such that the model sees enough examples of every label?\nOne way to do this could be to use data from [here](https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized) and combine them with the training set such that each labels have enough examples",
      "votes": null
    },
    {
      "id": "570794",
      "postDate": "07/08/2019 19:07:45",
      "content": "<p>Hello <a href=\"/bibek777\">@bibek777</a> , my understanding from your question is - you want to know how to address the 'Imbalanced Data' problem.</p>\n\n<p>If so , indeed , imbalanced data is a major challenge for ML and has to be addressed before applying any ML model to an imbalanced data.</p>\n\n<p>There are many ways to deal with this challenge -</p>\n\n<pre><code>  a) Under-Sampling : Undersampling can be defined as removing some observations of the majority class.\n\n  b) Over-Sampling : Oversampling can be defined as adding more copies of the minority class. \n\n  c) Synthetic sample generation \n</code></pre>\n\n<p>Please refer to the following link for further details -\n<a href=\"https://towardsdatascience.com/methods-for-dealing-with-imbalanced-data-5b761be45a18\">https://towardsdatascience.com/methods-for-dealing-with-imbalanced-data-5b761be45a18</a></p>",
      "rawMarkdown": "Hello @bibek777 , my understanding from your question is - you want to know how to address the 'Imbalanced Data' problem.\n\nIf so , indeed , imbalanced data is a major challenge for ML and has to be addressed before applying any ML model to an imbalanced data.\n\nThere are many ways to deal with this challenge -\n\n      a) Under-Sampling : Undersampling can be defined as removing some observations of the majority class.\n\n      b) Over-Sampling : Oversampling can be defined as adding more copies of the minority class. \n\n      c) Synthetic sample generation \n\nPlease refer to the following link for further details -\nhttps://towardsdatascience.com/methods-for-dealing-with-imbalanced-data-5b761be45a18",
      "votes": null
    },
    {
      "id": "570958",
      "postDate": "07/09/2019 01:22:21",
      "content": "<p>Those three ways are really nice way to deal with data imbalance problem but what I am proposing is different from the three ways that u mentioned. I want to add new samples from external data to our training data so that we have enough examples of each label.</p>",
      "rawMarkdown": "Those three ways are really nice way to deal with data imbalance problem but what I am proposing is different from the three ways that u mentioned. I want to add new samples from external data to our training data so that we have enough examples of each label.",
      "votes": null
    },
    {
      "id": "571049",
      "postDate": "07/09/2019 05:32:29",
      "content": "<p>oh k ...got you ...</p>",
      "rawMarkdown": "oh k ...got you ...",
      "votes": null
    },
    {
      "id": "573520",
      "postDate": "07/12/2019 11:10:18",
      "content": "<p>Did you tried this and does this helped?</p>",
      "rawMarkdown": "Did you tried this and does this helped?",
      "votes": null
    },
    {
      "id": "574174",
      "postDate": "07/13/2019 10:16:26",
      "content": "<p>I am working on it...will update once finished</p>",
      "rawMarkdown": "I am working on it...will update once finished",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 570794,
      "author_name": "prmohanty",
      "author_url": "",
      "post_date": "07/08/2019 19:07:45",
      "content": "<p>Hello <a href=\"/bibek777\">@bibek777</a> , my understanding from your question is - you want to know how to address the 'Imbalanced Data' problem.</p>\n\n<p>If so , indeed , imbalanced data is a major challenge for ML and has to be addressed before applying any ML model to an imbalanced data.</p>\n\n<p>There are many ways to deal with this challenge -</p>\n\n<pre><code>  a) Under-Sampling : Undersampling can be defined as removing some observations of the majority class.\n\n  b) Over-Sampling : Oversampling can be defined as adding more copies of the minority class. \n\n  c) Synthetic sample generation \n</code></pre>\n\n<p>Please refer to the following link for further details -\n<a href=\"https://towardsdatascience.com/methods-for-dealing-with-imbalanced-data-5b761be45a18\">https://towardsdatascience.com/methods-for-dealing-with-imbalanced-data-5b761be45a18</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 570958,
          "author_name": "bibek777",
          "author_url": "",
          "post_date": "07/09/2019 01:22:21",
          "content": "<p>Those three ways are really nice way to deal with data imbalance problem but what I am proposing is different from the three ways that u mentioned. I want to add new samples from external data to our training data so that we have enough examples of each label.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 571049,
          "author_name": "prmohanty",
          "author_url": "",
          "post_date": "07/09/2019 05:32:29",
          "content": "<p>oh k ...got you ...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 573520,
      "author_name": "nitin29",
      "author_url": "",
      "post_date": "07/12/2019 11:10:18",
      "content": "<p>Did you tried this and does this helped?</p>",
      "votes": null,
      "replies": [
        {
          "id": 574174,
          "author_name": "bibek777",
          "author_url": "",
          "post_date": "07/13/2019 10:16:26",
          "content": "<p>I am working on it...will update once finished</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "570708": "Did anyone try balancing the dataset such that the model sees enough examples of every label?\nOne way to do this could be to use data from [here](https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized) and combine them with the training set such that each labels have enough examples",
    "570794": "Hello @bibek777 , my understanding from your question is - you want to know how to address the 'Imbalanced Data' problem.\n\nIf so , indeed , imbalanced data is a major challenge for ML and has to be addressed before applying any ML model to an imbalanced data.\n\nThere are many ways to deal with this challenge -\n\n      a) Under-Sampling : Undersampling can be defined as removing some observations of the majority class.\n\n      b) Over-Sampling : Oversampling can be defined as adding more copies of the minority class. \n\n      c) Synthetic sample generation \n\nPlease refer to the following link for further details -\nhttps://towardsdatascience.com/methods-for-dealing-with-imbalanced-data-5b761be45a18",
    "570958": "Those three ways are really nice way to deal with data imbalance problem but what I am proposing is different from the three ways that u mentioned. I want to add new samples from external data to our training data so that we have enough examples of each label.",
    "571049": "oh k ...got you ...",
    "573520": "Did you tried this and does this helped?",
    "574174": "I am working on it...will update once finished"
  },
  "source": "meta"
}