{
  "id": 156614,
  "title": "Should we use class_weight to handle class imbalance ??",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/156614",
  "author_name": "",
  "post_date": "2020-06-06T23:03:41.504150900Z",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>In a course, it was mentioned that for imbalanced dataset model get biased towards the larger class. Using class_weight is one of the ways to handle this imbalance. But in this competition, people hardly used class_weight. Is there any reason behind it??</p>",
  "messages": [
    {
      "id": "876687",
      "postDate": "06/06/2020 23:03:41",
      "content": "<p>In a course, it was mentioned that for imbalanced dataset model get biased towards the larger class. Using class_weight is one of the ways to handle this imbalance. But in this competition, people hardly used class_weight. Is there any reason behind it??</p>",
      "rawMarkdown": "In a course, it was mentioned that for imbalanced dataset model get biased towards the larger class. Using class_weight is one of the ways to handle this imbalance. But in this competition, people hardly used class_weight. Is there any reason behind it??",
      "votes": null
    },
    {
      "id": "876737",
      "postDate": "06/07/2020 01:59:23",
      "content": "<p>We should try it and see if it increases our CV/LB. That's the best way to find out. </p>\n\n<p>Usually i use <code>class_weight</code> to match the train distribution to the test distribution. For example, let's say that the train distribution has 1% malignant and the test distribution has 10%. If you train a model on 1%, then when you predict on test, it will think that the overall population has 1% and not make enough malignant predictions. But if you increase the weight during training by a factor of 10. Then when you predict the test data it will more accurately model the population.</p>\n\n<p>In many Kaggle comps when the train distribution is the same as the test distribution, increasing <code>class_weight</code> hasn't helped me. In this comp, I have not tried it yet.</p>",
      "rawMarkdown": "We should try it and see if it increases our CV/LB. That's the best way to find out. \n\nUsually i use `class_weight` to match the train distribution to the test distribution. For example, let's say that the train distribution has 1% malignant and the test distribution has 10%. If you train a model on 1%, then when you predict on test, it will think that the overall population has 1% and not make enough malignant predictions. But if you increase the weight during training by a factor of 10. Then when you predict the test data it will more accurately model the population.\n\nIn many Kaggle comps when the train distribution is the same as the test distribution, increasing `class_weight` hasn't helped me. In this comp, I have not tried it yet.",
      "votes": null
    },
    {
      "id": "876864",
      "postDate": "06/07/2020 05:31:45",
      "content": "<p>[Updated] Weight for training (1.76% malignant) vs test (<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215\">2.37%</a>).</p>",
      "rawMarkdown": "[Updated] Weight for training (1.76% malignant) vs test ([2.37%](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215)).",
      "votes": null
    },
    {
      "id": "952679",
      "postDate": "07/31/2020 06:27:00",
      "content": "<p>So the weights are helpful if we have a similar distribution between the test and train.\nSo when we use the Triple stratified data set then we might get a high validation score as the validation set f a fold has a similar distribution to the train set of the same fold, but it cannot be relied upon as the test data might have 10% malignant data?\nI wish to try the weights  along with the awsome data set by <a href=\"/cdeotte\">@cdeotte</a>  of 4000 malignant images but what should be the ratio of the weights to begin with(your suggestion and insights ) and how to create a reliable CV</p>",
      "rawMarkdown": "So the weights are helpful if we have a similar distribution between the test and train.\nSo when we use the Triple stratified data set then we might get a high validation score as the validation set f a fold has a similar distribution to the train set of the same fold, but it cannot be relied upon as the test data might have 10% malignant data?\nI wish to try the weights  along with the awsome data set by @cdeotte  of 4000 malignant images but what should be the ratio of the weights to begin with(your suggestion and insights ) and how to create a reliable CV",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 876737,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/07/2020 01:59:23",
      "content": "<p>We should try it and see if it increases our CV/LB. That's the best way to find out. </p>\n\n<p>Usually i use <code>class_weight</code> to match the train distribution to the test distribution. For example, let's say that the train distribution has 1% malignant and the test distribution has 10%. If you train a model on 1%, then when you predict on test, it will think that the overall population has 1% and not make enough malignant predictions. But if you increase the weight during training by a factor of 10. Then when you predict the test data it will more accurately model the population.</p>\n\n<p>In many Kaggle comps when the train distribution is the same as the test distribution, increasing <code>class_weight</code> hasn't helped me. In this comp, I have not tried it yet.</p>",
      "votes": null,
      "replies": [
        {
          "id": 876864,
          "author_name": "sirishks",
          "author_url": "",
          "post_date": "06/07/2020 05:31:45",
          "content": "<p>[Updated] Weight for training (1.76% malignant) vs test (<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215\">2.37%</a>).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 952679,
          "author_name": "adigamer970",
          "author_url": "",
          "post_date": "07/31/2020 06:27:00",
          "content": "<p>So the weights are helpful if we have a similar distribution between the test and train.\nSo when we use the Triple stratified data set then we might get a high validation score as the validation set f a fold has a similar distribution to the train set of the same fold, but it cannot be relied upon as the test data might have 10% malignant data?\nI wish to try the weights  along with the awsome data set by <a href=\"/cdeotte\">@cdeotte</a>  of 4000 malignant images but what should be the ratio of the weights to begin with(your suggestion and insights ) and how to create a reliable CV</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "876687": "In a course, it was mentioned that for imbalanced dataset model get biased towards the larger class. Using class_weight is one of the ways to handle this imbalance. But in this competition, people hardly used class_weight. Is there any reason behind it??",
    "876737": "We should try it and see if it increases our CV/LB. That's the best way to find out. \n\nUsually i use `class_weight` to match the train distribution to the test distribution. For example, let's say that the train distribution has 1% malignant and the test distribution has 10%. If you train a model on 1%, then when you predict on test, it will think that the overall population has 1% and not make enough malignant predictions. But if you increase the weight during training by a factor of 10. Then when you predict the test data it will more accurately model the population.\n\nIn many Kaggle comps when the train distribution is the same as the test distribution, increasing `class_weight` hasn't helped me. In this comp, I have not tried it yet.",
    "876864": "[Updated] Weight for training (1.76% malignant) vs test ([2.37%](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215)).",
    "952679": "So the weights are helpful if we have a similar distribution between the test and train.\nSo when we use the Triple stratified data set then we might get a high validation score as the validation set f a fold has a similar distribution to the train set of the same fold, but it cannot be relied upon as the test data might have 10% malignant data?\nI wish to try the weights  along with the awsome data set by @cdeotte  of 4000 malignant images but what should be the ratio of the weights to begin with(your suggestion and insights ) and how to create a reliable CV"
  },
  "source": "meta"
}