{
  "id": 357924,
  "title": "Ways to handle imbalance dataset",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/357924",
  "author_name": "Dr. Alvinleenh",
  "post_date": "2022-10-06T06:18:54.638000",
  "votes": 9,
  "comment_count": 0,
  "views": 0,
  "content": "<p><strong>CREDIT</strong><br>\nThanks <a href=\"https://www.kaggle.com/aatiffraz\" target=\"_blank\">@aatiffraz</a> <a href=\"https://www.kaggle.com/gazu468\" target=\"_blank\">@gazu468</a> <a href=\"https://www.kaggle.com/infrarosso\" target=\"_blank\">@infrarosso</a> for bringing up the class imbalance issue from TPS Oct's dataset</p>\n<p>Based on initial EDA on train_0.csv,</p>\n<ul>\n<li>Team A has the ratio for no_score:score of 16:1</li>\n<li>Team B has the ratio for no_score:score of 17:1<br>\nAnd the probability of scoring a goal is much more lower, which is hard to predict.</li>\n</ul>\n<p>Detailed analysis by Aatif Fraz <a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/357241\" target=\"_blank\">here</a></p>\n<p>Here are some suggestion from the following notebook, but we do not have some example on how do we implement this.</p>\n<p><strong>Suggestions by Infrared</strong> <a href=\"https://www.kaggle.com/code/infrarosso/tps-oct-2022-eda-lgbm-model-ensemble#Targets-Distribution-(imbalanced)\" target=\"_blank\">here</a></p>\n<ul>\n<li>Use class weight option in ML model (if supported)</li>\n<li>Imbalanced-learn tool which handle imbalanced classes here</li>\n<li>Use ensemble/bagging techniques to reduce effect of imbalanced</li>\n<li>Validate the model with Stratified cross-validation or train/test split (Implemented in Infrared's notebook)</li>\n</ul>\n<p><strong>Suggestions by Gaju Ahmed</strong> <a href=\"https://www.kaggle.com/code/gazu468/tps-oct-22-simple-eda-and-xgboost#Class-Imbalance\" target=\"_blank\">here</a></p>\n<ul>\n<li>Random under-sampling or over_sampling</li>\n<li>SMOTE (Synthetic Minority Oversampling Technique)</li>\n<li>Change performance metric (Confusion Matrix, Precision, Recall, F1 Score, AUROC)</li>\n</ul>\n<p>Therefore, I have created a notebook to test out following method </p>\n<p><strong>Example code</strong> <a href=\"https://www.kaggle.com/code/alvinleenh/tpsoct22-imbalance-dataset-with-4-techniques\" target=\"_blank\">here</a></p>\n<ul>\n<li>Stratified cross validation (must have)</li>\n<li>Class weight in ML model</li>\n<li>Bagging with subsample</li>\n<li>Undersampling (tested on event_time, may not be useful)</li>\n</ul>\n<p>Please upvote 👍 or leave a comment if you wish to share other techniques/notebooks to handle class imbalance. <br>\nI will include your suggestion in this topic as reference. Thank you!</p>\n<p>Good luck!</p>",
  "messages": [
    {
      "id": 1974232,
      "postDate": "2022-10-06T06:18:54.640Z",
      "content": "<p><strong>CREDIT</strong><br>\nThanks <a href=\"https://www.kaggle.com/aatiffraz\" target=\"_blank\">@aatiffraz</a> <a href=\"https://www.kaggle.com/gazu468\" target=\"_blank\">@gazu468</a> <a href=\"https://www.kaggle.com/infrarosso\" target=\"_blank\">@infrarosso</a> for bringing up the class imbalance issue from TPS Oct's dataset</p>\n<p>Based on initial EDA on train_0.csv,</p>\n<ul>\n<li>Team A has the ratio for no_score:score of 16:1</li>\n<li>Team B has the ratio for no_score:score of 17:1<br>\nAnd the probability of scoring a goal is much more lower, which is hard to predict.</li>\n</ul>\n<p>Detailed analysis by Aatif Fraz <a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/357241\" target=\"_blank\">here</a></p>\n<p>Here are some suggestion from the following notebook, but we do not have some example on how do we implement this.</p>\n<p><strong>Suggestions by Infrared</strong> <a href=\"https://www.kaggle.com/code/infrarosso/tps-oct-2022-eda-lgbm-model-ensemble#Targets-Distribution-(imbalanced)\" target=\"_blank\">here</a></p>\n<ul>\n<li>Use class weight option in ML model (if supported)</li>\n<li>Imbalanced-learn tool which handle imbalanced classes here</li>\n<li>Use ensemble/bagging techniques to reduce effect of imbalanced</li>\n<li>Validate the model with Stratified cross-validation or train/test split (Implemented in Infrared's notebook)</li>\n</ul>\n<p><strong>Suggestions by Gaju Ahmed</strong> <a href=\"https://www.kaggle.com/code/gazu468/tps-oct-22-simple-eda-and-xgboost#Class-Imbalance\" target=\"_blank\">here</a></p>\n<ul>\n<li>Random under-sampling or over_sampling</li>\n<li>SMOTE (Synthetic Minority Oversampling Technique)</li>\n<li>Change performance metric (Confusion Matrix, Precision, Recall, F1 Score, AUROC)</li>\n</ul>\n<p>Therefore, I have created a notebook to test out following method </p>\n<p><strong>Example code</strong> <a href=\"https://www.kaggle.com/code/alvinleenh/tpsoct22-imbalance-dataset-with-4-techniques\" target=\"_blank\">here</a></p>\n<ul>\n<li>Stratified cross validation (must have)</li>\n<li>Class weight in ML model</li>\n<li>Bagging with subsample</li>\n<li>Undersampling (tested on event_time, may not be useful)</li>\n</ul>\n<p>Please upvote 👍 or leave a comment if you wish to share other techniques/notebooks to handle class imbalance. <br>\nI will include your suggestion in this topic as reference. Thank you!</p>\n<p>Good luck!</p>",
      "rawMarkdown": "**CREDIT**\nThanks @aatiffraz @gazu468 @infrarosso for bringing up the class imbalance issue from TPS Oct's dataset\n\nBased on initial EDA on train_0.csv,\n- Team A has the ratio for no_score:score of 16:1\n- Team B has the ratio for no_score:score of 17:1\nAnd the probability of scoring a goal is much more lower, which is hard to predict.\n\nDetailed analysis by Aatif Fraz [here](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/357241)\n\n\nHere are some suggestion from the following notebook, but we do not have some example on how do we implement this.\n\n**Suggestions by Infrared** [here](https://www.kaggle.com/code/infrarosso/tps-oct-2022-eda-lgbm-model-ensemble#Targets-Distribution-(imbalanced))\n- Use class weight option in ML model (if supported)\n- Imbalanced-learn tool which handle imbalanced classes here\n- Use ensemble/bagging techniques to reduce effect of imbalanced\n- Validate the model with Stratified cross-validation or train/test split (Implemented in Infrared's notebook)\n\n**Suggestions by Gaju Ahmed** [here](https://www.kaggle.com/code/gazu468/tps-oct-22-simple-eda-and-xgboost#Class-Imbalance)\n- Random under-sampling or over_sampling\n- SMOTE (Synthetic Minority Oversampling Technique)\n- Change performance metric (Confusion Matrix, Precision, Recall, F1 Score, AUROC)\n\nTherefore, I have created a notebook to test out following method \n\n**Example code** [here](https://www.kaggle.com/code/alvinleenh/tpsoct22-imbalance-dataset-with-4-techniques)\n- Stratified cross validation (must have)\n- Class weight in ML model\n- Bagging with subsample\n- Undersampling (tested on event_time, may not be useful)\n\nPlease upvote 👍 or leave a comment if you wish to share other techniques/notebooks to handle class imbalance. \nI will include your suggestion in this topic as reference. Thank you!\n\n\nGood luck!",
      "votes": 9
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1974232": "**CREDIT**\nThanks @aatiffraz @gazu468 @infrarosso for bringing up the class imbalance issue from TPS Oct's dataset\n\nBased on initial EDA on train_0.csv,\n- Team A has the ratio for no_score:score of 16:1\n- Team B has the ratio for no_score:score of 17:1\nAnd the probability of scoring a goal is much more lower, which is hard to predict.\n\nDetailed analysis by Aatif Fraz [here](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/357241)\n\n\nHere are some suggestion from the following notebook, but we do not have some example on how do we implement this.\n\n**Suggestions by Infrared** [here](https://www.kaggle.com/code/infrarosso/tps-oct-2022-eda-lgbm-model-ensemble#Targets-Distribution-(imbalanced))\n- Use class weight option in ML model (if supported)\n- Imbalanced-learn tool which handle imbalanced classes here\n- Use ensemble/bagging techniques to reduce effect of imbalanced\n- Validate the model with Stratified cross-validation or train/test split (Implemented in Infrared's notebook)\n\n**Suggestions by Gaju Ahmed** [here](https://www.kaggle.com/code/gazu468/tps-oct-22-simple-eda-and-xgboost#Class-Imbalance)\n- Random under-sampling or over_sampling\n- SMOTE (Synthetic Minority Oversampling Technique)\n- Change performance metric (Confusion Matrix, Precision, Recall, F1 Score, AUROC)\n\nTherefore, I have created a notebook to test out following method \n\n**Example code** [here](https://www.kaggle.com/code/alvinleenh/tpsoct22-imbalance-dataset-with-4-techniques)\n- Stratified cross validation (must have)\n- Class weight in ML model\n- Bagging with subsample\n- Undersampling (tested on event_time, may not be useful)\n\nPlease upvote 👍 or leave a comment if you wish to share other techniques/notebooks to handle class imbalance. \nI will include your suggestion in this topic as reference. Thank you!\n\n\nGood luck!"
  }
}