{
  "id": 52999,
  "title": "Data Imbalance",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/52999",
  "author_name": "",
  "post_date": "2018-03-26T05:09:05.417184400Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I was wondering how everyone was handling the data imbalance. Specifically, if you are using sampling techniques, how/what are you implementing?</p>\n\n<p>Also, there are some good notebooks on different strategies. I'll highlight my favorite:</p>\n\n<p><a href=\"https://www.kaggle.com/kailex/talkingdata-eda-and-class-imbalance\">https://www.kaggle.com/kailex/talkingdata-eda-and-class-imbalance</a></p>\n\n<p>I usually like to use SMOTE. However, I haven't tried it here yet. I have tried using various values for <code>scale_pos_weight</code> in several LGBM and XGBoost models with decent results. What have you seen? </p>",
  "messages": [
    {
      "id": "303390",
      "postDate": "03/26/2018 05:09:05",
      "content": "<p>I was wondering how everyone was handling the data imbalance. Specifically, if you are using sampling techniques, how/what are you implementing?</p>\n\n<p>Also, there are some good notebooks on different strategies. I'll highlight my favorite:</p>\n\n<p><a href=\"https://www.kaggle.com/kailex/talkingdata-eda-and-class-imbalance\">https://www.kaggle.com/kailex/talkingdata-eda-and-class-imbalance</a></p>\n\n<p>I usually like to use SMOTE. However, I haven't tried it here yet. I have tried using various values for <code>scale_pos_weight</code> in several LGBM and XGBoost models with decent results. What have you seen? </p>",
      "rawMarkdown": "I was wondering how everyone was handling the data imbalance. Specifically, if you are using sampling techniques, how/what are you implementing?\n\nAlso, there are some good notebooks on different strategies. I'll highlight my favorite:\n\nhttps://www.kaggle.com/kailex/talkingdata-eda-and-class-imbalance\n\nI usually like to use SMOTE. However, I haven't tried it here yet. I have tried using various values for `scale_pos_weight` in several LGBM and XGBoost models with decent results. What have you seen?",
      "votes": null
    },
    {
      "id": "303402",
      "postDate": "03/26/2018 05:42:31",
      "content": "<p>I'm using scale_pos_weight so far, it helps.</p>",
      "rawMarkdown": "I'm using scale_pos_weight so far, it helps.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 303402,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "03/26/2018 05:42:31",
      "content": "<p>I'm using scale_pos_weight so far, it helps.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "303390": "I was wondering how everyone was handling the data imbalance. Specifically, if you are using sampling techniques, how/what are you implementing?\n\nAlso, there are some good notebooks on different strategies. I'll highlight my favorite:\n\nhttps://www.kaggle.com/kailex/talkingdata-eda-and-class-imbalance\n\nI usually like to use SMOTE. However, I haven't tried it here yet. I have tried using various values for `scale_pos_weight` in several LGBM and XGBoost models with decent results. What have you seen?",
    "303402": "I'm using scale_pos_weight so far, it helps."
  },
  "source": "meta"
}