{
  "id": 72947,
  "title": "How to deal with class imbalance in this case ?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/72947",
  "author_name": "",
  "post_date": "2018-11-28T14:29:32.305953800Z",
  "votes": 15,
  "comment_count": 8,
  "views": 0,
  "content": "<p>As we can see highly skewed data is there , bringing more data is not allowed neither pretrained model on toxic comments i have is also not allowed so now what to do in this case , it's really confusing </p>",
  "messages": [
    {
      "id": "429211",
      "postDate": "11/28/2018 14:29:32",
      "content": "<p>As we can see highly skewed data is there , bringing more data is not allowed neither pretrained model on toxic comments i have is also not allowed so now what to do in this case , it's really confusing </p>",
      "rawMarkdown": "As we can see highly skewed data is there , bringing more data is not allowed neither pretrained model on toxic comments i have is also not allowed so now what to do in this case , it's really confusing",
      "votes": null
    },
    {
      "id": "429259",
      "postDate": "11/28/2018 15:43:07",
      "content": "<p>Here are three methods to balance your data, feel free to give them a look :</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/theoviel/using-word-embeddings-for-data-augmentation\">https://www.kaggle.com/theoviel/using-word-embeddings-for-data-augmentation</a></li>\n<li><a href=\"https://www.kaggle.com/shujian/fake-some-positive-data-data-augmentation\">https://www.kaggle.com/shujian/fake-some-positive-data-data-augmentation</a></li>\n<li><a href=\"https://www.kaggle.com/theoviel/dealing-with-class-imbalance-with-smote\">https://www.kaggle.com/theoviel/dealing-with-class-imbalance-with-smote</a></li>\n</ul>\n\n<p>Keep me updated if you get anykind of results with them!</p>",
      "rawMarkdown": "Here are three methods to balance your data, feel free to give them a look :\n\n* https://www.kaggle.com/theoviel/using-word-embeddings-for-data-augmentation\n* https://www.kaggle.com/shujian/fake-some-positive-data-data-augmentation\n* https://www.kaggle.com/theoviel/dealing-with-class-imbalance-with-smote\n\nKeep me updated if you get anykind of results with them!",
      "votes": null
    },
    {
      "id": "429260",
      "postDate": "11/28/2018 15:44:34",
      "content": "<p>Thanks theol , actually exams are running , and i have to pass in all , will surely coming with new versions of stack models</p>",
      "rawMarkdown": "Thanks theol , actually exams are running , and i have to pass in all , will surely coming with new versions of stack models",
      "votes": null
    },
    {
      "id": "429619",
      "postDate": "11/29/2018 04:36:04",
      "content": "<p>Hi Rohan, I have listed down 2 of the best approaches to deal with class imbalance.</p>\n\n<ol>\n<li><p>Sampling Solutions:\na. Oversampling:\ni. Random Oversampling: Create duplicates of minority class rows [Might lead to overfitting]\nii. Informed Oversampling: e.g. SMOTE\nb. Undersampling:\ni. Random Undersampling: Randomly remove rows of majority class [Might lose out on important data points]\nii. Informed Undersampling: e.g. Tomek Undersampling</p></li>\n<li><p>Cost-Sensitive Learning:\nAssociates a different cost to false positives and false negatives while training the model.\nDifferent cost combinations will have to be tried using grid search to arrive at the most optimal solution.\nrpart function in R has a feature that allows for cost-sensitive learning</p></li>\n</ol>\n\n<p>Choice of solution:\nI would personally suggest you to use cost-sensitive learning wherever possible but not every algorithm allows the user to do so. In such cases go with oversampling/undersampling.</p>",
      "rawMarkdown": "Hi Rohan, I have listed down 2 of the best approaches to deal with class imbalance.\n\n1. Sampling Solutions:\na. Oversampling:\ni. Random Oversampling: Create duplicates of minority class rows [Might lead to overfitting]\nii. Informed Oversampling: e.g. SMOTE\nb. Undersampling:\ni. Random Undersampling: Randomly remove rows of majority class [Might lose out on important data points]\nii. Informed Undersampling: e.g. Tomek Undersampling\n\n2. Cost-Sensitive Learning:\nAssociates a different cost to false positives and false negatives while training the model.\nDifferent cost combinations will have to be tried using grid search to arrive at the most optimal solution.\nrpart function in R has a feature that allows for cost-sensitive learning\n\nChoice of solution:\nI would personally suggest you to use cost-sensitive learning wherever possible but not every algorithm allows the user to do so. In such cases go with oversampling/undersampling.",
      "votes": null
    },
    {
      "id": "429672",
      "postDate": "11/29/2018 06:56:51",
      "content": "<p>Nice outline! thanks.</p>",
      "rawMarkdown": "Nice outline! thanks.",
      "votes": null
    },
    {
      "id": "431829",
      "postDate": "12/03/2018 01:18:41",
      "content": "<p>You might try anomaly detection with autoencoders as well:</p>\n\n<p><a href=\"https://medium.com/&lt;a href=\">@curiousily</a>/credit-card-fraud-detection-using-autoencoders-in-keras-tensorflow-for-hackers-part-vii-20e0c85301bd\"&gt;https://medium.com/<a href=\"/curiousily\">@curiousily</a>/credit-card-fraud-detection-using-autoencoders-in-keras-tensorflow-for-hackers-part-vii-20e0c85301bd</p>",
      "rawMarkdown": "You might try anomaly detection with autoencoders as well:\n\nhttps://medium.com/@curiousily/credit-card-fraud-detection-using-autoencoders-in-keras-tensorflow-for-hackers-part-vii-20e0c85301bd",
      "votes": null
    },
    {
      "id": "431997",
      "postDate": "12/03/2018 08:33:37",
      "content": "<p>well i have intuition of classical English grammar, to do data augmentation but thanks i will also look onto this , nice article mark </p>",
      "rawMarkdown": "well i have intuition of classical English grammar, to do data augmentation but thanks i will also look onto this , nice article mark",
      "votes": null
    },
    {
      "id": "432014",
      "postDate": "12/03/2018 08:56:39",
      "content": "<p>Another approach for credit card fraud detection using Self Organizing Maps:\n<a href=\"https://github.com/Jestin-278/fraud_detection_using_selforganizingmaps_and_anns\">https://github.com/Jestin-278/fraud_detection_using_selforganizingmaps_and_anns</a></p>",
      "rawMarkdown": "Another approach for credit card fraud detection using Self Organizing Maps:\nhttps://github.com/Jestin-278/fraud_detection_using_selforganizingmaps_and_anns",
      "votes": null
    },
    {
      "id": "433805",
      "postDate": "12/05/2018 13:52:48",
      "content": "<p>Maybe you can try Focal Loss, but it cant have anything improvement except the ascent of running speed.</p>",
      "rawMarkdown": "Maybe you can try Focal Loss, but it cant have anything improvement except the ascent of running speed.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 429259,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "11/28/2018 15:43:07",
      "content": "<p>Here are three methods to balance your data, feel free to give them a look :</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/theoviel/using-word-embeddings-for-data-augmentation\">https://www.kaggle.com/theoviel/using-word-embeddings-for-data-augmentation</a></li>\n<li><a href=\"https://www.kaggle.com/shujian/fake-some-positive-data-data-augmentation\">https://www.kaggle.com/shujian/fake-some-positive-data-data-augmentation</a></li>\n<li><a href=\"https://www.kaggle.com/theoviel/dealing-with-class-imbalance-with-smote\">https://www.kaggle.com/theoviel/dealing-with-class-imbalance-with-smote</a></li>\n</ul>\n\n<p>Keep me updated if you get anykind of results with them!</p>",
      "votes": null,
      "replies": [
        {
          "id": 429260,
          "author_name": "rohandx1996",
          "author_url": "",
          "post_date": "11/28/2018 15:44:34",
          "content": "<p>Thanks theol , actually exams are running , and i have to pass in all , will surely coming with new versions of stack models</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 429619,
      "author_name": "jman278",
      "author_url": "",
      "post_date": "11/29/2018 04:36:04",
      "content": "<p>Hi Rohan, I have listed down 2 of the best approaches to deal with class imbalance.</p>\n\n<ol>\n<li><p>Sampling Solutions:\na. Oversampling:\ni. Random Oversampling: Create duplicates of minority class rows [Might lead to overfitting]\nii. Informed Oversampling: e.g. SMOTE\nb. Undersampling:\ni. Random Undersampling: Randomly remove rows of majority class [Might lose out on important data points]\nii. Informed Undersampling: e.g. Tomek Undersampling</p></li>\n<li><p>Cost-Sensitive Learning:\nAssociates a different cost to false positives and false negatives while training the model.\nDifferent cost combinations will have to be tried using grid search to arrive at the most optimal solution.\nrpart function in R has a feature that allows for cost-sensitive learning</p></li>\n</ol>\n\n<p>Choice of solution:\nI would personally suggest you to use cost-sensitive learning wherever possible but not every algorithm allows the user to do so. In such cases go with oversampling/undersampling.</p>",
      "votes": null,
      "replies": [
        {
          "id": 429672,
          "author_name": "iiiiiiillllllli",
          "author_url": "",
          "post_date": "11/29/2018 06:56:51",
          "content": "<p>Nice outline! thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 431829,
      "author_name": "mconway",
      "author_url": "",
      "post_date": "12/03/2018 01:18:41",
      "content": "<p>You might try anomaly detection with autoencoders as well:</p>\n\n<p><a href=\"https://medium.com/&lt;a href=\">@curiousily</a>/credit-card-fraud-detection-using-autoencoders-in-keras-tensorflow-for-hackers-part-vii-20e0c85301bd\"&gt;https://medium.com/<a href=\"/curiousily\">@curiousily</a>/credit-card-fraud-detection-using-autoencoders-in-keras-tensorflow-for-hackers-part-vii-20e0c85301bd</p>",
      "votes": null,
      "replies": [
        {
          "id": 431997,
          "author_name": "rohandx1996",
          "author_url": "",
          "post_date": "12/03/2018 08:33:37",
          "content": "<p>well i have intuition of classical English grammar, to do data augmentation but thanks i will also look onto this , nice article mark </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 432014,
          "author_name": "jman278",
          "author_url": "",
          "post_date": "12/03/2018 08:56:39",
          "content": "<p>Another approach for credit card fraud detection using Self Organizing Maps:\n<a href=\"https://github.com/Jestin-278/fraud_detection_using_selforganizingmaps_and_anns\">https://github.com/Jestin-278/fraud_detection_using_selforganizingmaps_and_anns</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 433805,
      "author_name": "xiaobai1123q",
      "author_url": "",
      "post_date": "12/05/2018 13:52:48",
      "content": "<p>Maybe you can try Focal Loss, but it cant have anything improvement except the ascent of running speed.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "429211": "As we can see highly skewed data is there , bringing more data is not allowed neither pretrained model on toxic comments i have is also not allowed so now what to do in this case , it's really confusing",
    "429259": "Here are three methods to balance your data, feel free to give them a look :\n\n* https://www.kaggle.com/theoviel/using-word-embeddings-for-data-augmentation\n* https://www.kaggle.com/shujian/fake-some-positive-data-data-augmentation\n* https://www.kaggle.com/theoviel/dealing-with-class-imbalance-with-smote\n\nKeep me updated if you get anykind of results with them!",
    "429260": "Thanks theol , actually exams are running , and i have to pass in all , will surely coming with new versions of stack models",
    "429619": "Hi Rohan, I have listed down 2 of the best approaches to deal with class imbalance.\n\n1. Sampling Solutions:\na. Oversampling:\ni. Random Oversampling: Create duplicates of minority class rows [Might lead to overfitting]\nii. Informed Oversampling: e.g. SMOTE\nb. Undersampling:\ni. Random Undersampling: Randomly remove rows of majority class [Might lose out on important data points]\nii. Informed Undersampling: e.g. Tomek Undersampling\n\n2. Cost-Sensitive Learning:\nAssociates a different cost to false positives and false negatives while training the model.\nDifferent cost combinations will have to be tried using grid search to arrive at the most optimal solution.\nrpart function in R has a feature that allows for cost-sensitive learning\n\nChoice of solution:\nI would personally suggest you to use cost-sensitive learning wherever possible but not every algorithm allows the user to do so. In such cases go with oversampling/undersampling.",
    "429672": "Nice outline! thanks.",
    "431829": "You might try anomaly detection with autoencoders as well:\n\nhttps://medium.com/@curiousily/credit-card-fraud-detection-using-autoencoders-in-keras-tensorflow-for-hackers-part-vii-20e0c85301bd",
    "431997": "well i have intuition of classical English grammar, to do data augmentation but thanks i will also look onto this , nice article mark",
    "432014": "Another approach for credit card fraud detection using Self Organizing Maps:\nhttps://github.com/Jestin-278/fraud_detection_using_selforganizingmaps_and_anns",
    "433805": "Maybe you can try Focal Loss, but it cant have anything improvement except the ascent of running speed."
  },
  "source": "meta"
}