{
  "id": 537169,
  "title": " How to deal with the Imbalanced Dataset",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/537169",
  "author_name": "",
  "post_date": "2024-10-01T17:10:00.008917800Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I find there are only a few samples which labels are '3' . I want to use SMOTE to generate some data to avoid this,but there are so many 'nan' in features so that I can't use it directly.. So how to deal with this problem?</p>",
  "messages": [
    {
      "id": "3004284",
      "postDate": "10/01/2024 17:10:00",
      "content": "<p>I find there are only a few samples which labels are '3' . I want to use SMOTE to generate some data to avoid this,but there are so many 'nan' in features so that I can't use it directly.. So how to deal with this problem?</p>",
      "rawMarkdown": "I find there are only a few samples which labels are '3' . I want to use SMOTE to generate some data to avoid this,but there are so many 'nan' in features so that I can't use it directly.. So how to deal with this problem?",
      "votes": null
    },
    {
      "id": "3004319",
      "postDate": "10/01/2024 17:40:10",
      "content": "<p>Class Weight and Decision Treshold can be a good way to lead with this type of problem, you can find more about it in these links:</p>\n<p><a href=\"https://medium.com/analytics-vidhya/manipulating-class-weights-and-decision-threshold-cb7d8d9433a4\" target=\"_blank\">https://medium.com/analytics-vidhya/manipulating-class-weights-and-decision-threshold-cb7d8d9433a4</a></p>\n<p><a href=\"https://www.analyticsvidhya.com/blog/2020/10/improve-class-imbalance-class-weights/\" target=\"_blank\">https://www.analyticsvidhya.com/blog/2020/10/improve-class-imbalance-class-weights/</a></p>\n<p><a href=\"https://stackoverflow.com/questions/30972029/how-does-the-class-weight-parameter-in-scikit-learn-work\" target=\"_blank\">https://stackoverflow.com/questions/30972029/how-does-the-class-weight-parameter-in-scikit-learn-work</a></p>\n<p><a href=\"https://scikit-learn.org/stable/modules/classification_threshold.html\" target=\"_blank\">https://scikit-learn.org/stable/modules/classification_threshold.html</a></p>",
      "rawMarkdown": "Class Weight and Decision Treshold can be a good way to lead with this type of problem, you can find more about it in these links:\n\nhttps://medium.com/analytics-vidhya/manipulating-class-weights-and-decision-threshold-cb7d8d9433a4\n\nhttps://www.analyticsvidhya.com/blog/2020/10/improve-class-imbalance-class-weights/\n\nhttps://stackoverflow.com/questions/30972029/how-does-the-class-weight-parameter-in-scikit-learn-work\n\nhttps://scikit-learn.org/stable/modules/classification_threshold.html",
      "votes": null
    },
    {
      "id": "3038255",
      "postDate": "11/06/2024 19:00:27",
      "content": "<p><a href=\"https://www.kaggle.com/yashi003\" target=\"_blank\">@yashi003</a> you can take advantage of the mblearn library that has all advanced approaches of sampling ..</p>",
      "rawMarkdown": "yashi003 you can take advantage of the mblearn library that has all advanced approaches of sampling ..",
      "votes": null
    },
    {
      "id": "3038863",
      "postDate": "11/07/2024 13:02:47",
      "content": "<p>I tried SMOTE, but it actually made performance worse.In this competition, the amount of information in the training data is insufficient, so I think that it is difficult to completely solve the problem of imbalanced data. <a href=\"https://www.kaggle.com/yashi003\" target=\"_blank\">@yashi003</a></p>",
      "rawMarkdown": "I tried SMOTE, but it actually made performance worse.In this competition, the amount of information in the training data is insufficient, so I think that it is difficult to completely solve the problem of imbalanced data. @yashi003",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3004319,
      "author_name": "jvctor",
      "author_url": "",
      "post_date": "10/01/2024 17:40:10",
      "content": "<p>Class Weight and Decision Treshold can be a good way to lead with this type of problem, you can find more about it in these links:</p>\n<p><a href=\"https://medium.com/analytics-vidhya/manipulating-class-weights-and-decision-threshold-cb7d8d9433a4\" target=\"_blank\">https://medium.com/analytics-vidhya/manipulating-class-weights-and-decision-threshold-cb7d8d9433a4</a></p>\n<p><a href=\"https://www.analyticsvidhya.com/blog/2020/10/improve-class-imbalance-class-weights/\" target=\"_blank\">https://www.analyticsvidhya.com/blog/2020/10/improve-class-imbalance-class-weights/</a></p>\n<p><a href=\"https://stackoverflow.com/questions/30972029/how-does-the-class-weight-parameter-in-scikit-learn-work\" target=\"_blank\">https://stackoverflow.com/questions/30972029/how-does-the-class-weight-parameter-in-scikit-learn-work</a></p>\n<p><a href=\"https://scikit-learn.org/stable/modules/classification_threshold.html\" target=\"_blank\">https://scikit-learn.org/stable/modules/classification_threshold.html</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3038255,
      "author_name": "saidkoussi",
      "author_url": "",
      "post_date": "11/06/2024 19:00:27",
      "content": "<p><a href=\"https://www.kaggle.com/yashi003\" target=\"_blank\">@yashi003</a> you can take advantage of the mblearn library that has all advanced approaches of sampling ..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3038863,
      "author_name": "yuogawa",
      "author_url": "",
      "post_date": "11/07/2024 13:02:47",
      "content": "<p>I tried SMOTE, but it actually made performance worse.In this competition, the amount of information in the training data is insufficient, so I think that it is difficult to completely solve the problem of imbalanced data. <a href=\"https://www.kaggle.com/yashi003\" target=\"_blank\">@yashi003</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3004284": "I find there are only a few samples which labels are '3' . I want to use SMOTE to generate some data to avoid this,but there are so many 'nan' in features so that I can't use it directly.. So how to deal with this problem?",
    "3004319": "Class Weight and Decision Treshold can be a good way to lead with this type of problem, you can find more about it in these links:\n\nhttps://medium.com/analytics-vidhya/manipulating-class-weights-and-decision-threshold-cb7d8d9433a4\n\nhttps://www.analyticsvidhya.com/blog/2020/10/improve-class-imbalance-class-weights/\n\nhttps://stackoverflow.com/questions/30972029/how-does-the-class-weight-parameter-in-scikit-learn-work\n\nhttps://scikit-learn.org/stable/modules/classification_threshold.html",
    "3038255": "yashi003 you can take advantage of the mblearn library that has all advanced approaches of sampling ..",
    "3038863": "I tried SMOTE, but it actually made performance worse.In this competition, the amount of information in the training data is insufficient, so I think that it is difficult to completely solve the problem of imbalanced data. @yashi003"
  },
  "source": "meta"
}