{
  "id": 51544,
  "title": "How to resolve the imbalanced data",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/51544",
  "author_name": "",
  "post_date": "2018-03-10T05:38:45.458723100Z",
  "votes": 3,
  "comment_count": 12,
  "views": 0,
  "content": "<p>90% of is_attributed is 0  </p>",
  "messages": [
    {
      "id": "293558",
      "postDate": "03/10/2018 05:38:45",
      "content": "<p>90% of is_attributed is 0  </p>",
      "rawMarkdown": "90% of is_attributed is 0",
      "votes": null
    },
    {
      "id": "294083",
      "postDate": "03/11/2018 08:31:02",
      "content": "<p>I believe it is even more than 90%. Anyway, there are at least 3 conventional ways : undersampling those that have is_attributed == 0 , oversampling the is_attributed == 1 or SMORE (you can google it). I personally think a \"smart\" undersampling will be the way to go. We have a lot of data, and some of it is just for IPs that clicked once and never ever again returned... just an initial thought, for sure one can do even beter.</p>",
      "rawMarkdown": "I believe it is even more than 90%. Anyway, there are at least 3 conventional ways : undersampling those that have is_attributed == 0 , oversampling the is_attributed == 1 or SMORE (you can google it). I personally think a \"smart\" undersampling will be the way to go. We have a lot of data, and some of it is just for IPs that clicked once and never ever again returned... just an initial thought, for sure one can do even beter.",
      "votes": null
    },
    {
      "id": "300174",
      "postDate": "03/21/2018 11:26:11",
      "content": "<p>As asparuh mentioned there are methods available like over and under sampling, but these have drawbacks as oversampling leads to overfit the data as we just duplicate rows, undersampling can always lead to loss of information for the \"Fraud\" class in our concerned data, there are various new modern techniques like SMOTE(Synthetic minority oversampling technique) this is a kind of oversample but here the rows arent dupicated but new minority class rows are produced using K nearest neighbour technique.\nThere are various oter techniques that combines boosting and data sampling techniques , like RUSboost, Smoteboost etc(Google it various good research papers are available)</p>",
      "rawMarkdown": "As asparuh mentioned there are methods available like over and under sampling, but these have drawbacks as oversampling leads to overfit the data as we just duplicate rows, undersampling can always lead to loss of information for the \"Fraud\" class in our concerned data, there are various new modern techniques like SMOTE(Synthetic minority oversampling technique) this is a kind of oversample but here the rows arent dupicated but new minority class rows are produced using K nearest neighbour technique.\nThere are various oter techniques that combines boosting and data sampling techniques , like RUSboost, Smoteboost etc(Google it various good research papers are available)",
      "votes": null
    },
    {
      "id": "303105",
      "postDate": "03/25/2018 14:24:09",
      "content": "<p>I used SMOTE + tsNE in a fraud detection algorithm . Does this mean its more or less acceptable to use SMOTE data to train a model ? </p>\n\n<p><a href=\"https://www.kaggle.com/mlg-ulb/creditcardfraud/discussion/52948\">https://www.kaggle.com/mlg-ulb/creditcardfraud/discussion/52948</a></p>",
      "rawMarkdown": "I used SMOTE + tsNE in a fraud detection algorithm . Does this mean its more or less acceptable to use SMOTE data to train a model ? \n\nhttps://www.kaggle.com/mlg-ulb/creditcardfraud/discussion/52948",
      "votes": null
    },
    {
      "id": "303116",
      "postDate": "03/25/2018 14:56:44",
      "content": "<p>If you're using LightGBM or XGBoost then I would recommend tuning <code>scale_pos_weight</code> parameter. </p>\n\n<ol>\n<li>Both algorithms (in R and Python) has <code>scale_pos_weight</code> argument which controls the weights of the positive observations which is very useful for unbalanced classes.</li>\n<li>According to LightGBM documentation, default value for scale_pos_weight is 1.0 and it represents weight of positive class in binary classification task. There are different ways to calculate weight for positive class. I have taken reference from LightGBM’s Github issues page for calculation.</li>\n</ol>\n\n<p>Below is the example from one of my LightGBM <a href=\"https://www.kaggle.com/pranav84/talkingdata-eda-to-model-evaluation-lb-0-9683/notebook\">model</a> which provides me following information based on training and validation data:</p>\n\n<pre><code>Number of positive: 143540, number of negative: 56856460 \nTotal Bins 1704 \nNumber of data: 57000000, number of used features: 11\n</code></pre>\n\n<p>With this information, I have calculated exact value for <code>scale_pos_weight</code> with following formula :</p>\n\n<pre><code>scale_pos_weight = 100 - ( [number of positive samples / total samples ] * 100 )\nscale_pos_weight = 100 - ( [ 143540 / 57000000 ] * 100 ) \nscale_pos_weight = 99.74\n</code></pre>\n\n<p>As Asparuh Hristov has correctly pointed out, in full dataset, number of negative classes are more than 90% i.e \n <strong>around 99.8%</strong>   </p>",
      "rawMarkdown": "If you're using LightGBM or XGBoost then I would recommend tuning `scale_pos_weight` parameter. \n\n 1. Both algorithms (in R and Python) has `scale_pos_weight` argument which controls the weights of the positive observations which is very useful for unbalanced classes.\n 2. According to LightGBM documentation, default value for scale_pos_weight is 1.0 and it represents weight of positive class in binary classification task. There are different ways to calculate weight for positive class. I have taken reference from LightGBM’s Github issues page for calculation.\n\nBelow is the example from one of my LightGBM [model][1] which provides me following information based on training and validation data:\n\n    Number of positive: 143540, number of negative: 56856460 \n    Total Bins 1704 \n    Number of data: 57000000, number of used features: 11\n\n\nWith this information, I have calculated exact value for `scale_pos_weight` with following formula :\n\n    scale_pos_weight = 100 - ( [number of positive samples / total samples ] * 100 )\n    scale_pos_weight = 100 - ( [ 143540 / 57000000 ] * 100 ) \n    scale_pos_weight = 99.74\n\nAs Asparuh Hristov has correctly pointed out, in full dataset, number of negative classes are more than 90% i.e \n **around 99.8%**   \n\n  [1]: https://www.kaggle.com/pranav84/talkingdata-eda-to-model-evaluation-lb-0-9683/notebook",
      "votes": null
    },
    {
      "id": "308203",
      "postDate": "04/03/2018 05:50:34",
      "content": "<p>Smote data is used for training but as per my experience you should always keep a validation set aside before smote , then apply smote on the rest of the data and train it and first check on the validation set weather it is of some help or not.</p>",
      "rawMarkdown": "Smote data is used for training but as per my experience you should always keep a validation set aside before smote , then apply smote on the rest of the data and train it and first check on the validation set weather it is of some help or not.",
      "votes": null
    },
    {
      "id": "308362",
      "postDate": "04/03/2018 12:12:22",
      "content": "<p>Your formula looks wrong to me.  You are setting scale_pos_weigh to be the percentage of negative samples.</p>\n\n<p>Your formula gives a scale_pos_weight of 50 instead of 1 for a balanced dataset.</p>",
      "rawMarkdown": "Your formula looks wrong to me.  You are setting scale_pos_weigh to be the percentage of negative samples.\n\nYour formula gives a scale_pos_weight of 50 instead of 1 for a balanced dataset.",
      "votes": null
    },
    {
      "id": "308418",
      "postDate": "04/03/2018 13:26:00",
      "content": "<p>Thanks a lot @CPMP. I also have doubt about this formula and unfortunately, there is no explanation in documentation. I appreciate your thoughts on this. I have now <a href=\"https://github.com/Microsoft/LightGBM/issues/1299\">created</a>  new issue on LightGBM repo to see if creators/contributors can provide correct procedure to calculate it. </p>",
      "rawMarkdown": "Thanks a lot @CPMP. I also have doubt about this formula and unfortunately, there is no explanation in documentation. I appreciate your thoughts on this. I have now [created][1]  new issue on LightGBM repo to see if creators/contributors can provide correct procedure to calculate it. \n\n\n  [1]: https://github.com/Microsoft/LightGBM/issues/1299",
      "votes": null
    },
    {
      "id": "308621",
      "postDate": "04/03/2018 19:59:09",
      "content": "<p>I've came up with the formula that I believe is giving proper values:</p>\n\n<pre><code># T - no. of total samples\n# P - no. of positive samples\n# scale_pos_weight = percent of negative / percent of positive\n# which translates to:\n# scale_pos_weight = (100*(T-P)/T) / (100*P/T)\n# which further simplifies to beautiful:\nscale_pos_weight = T/P - 1\n</code></pre>\n\n<p>I'm bit tired atm and I hope I haven't made an error - but I am sure you'll double-check it.\nHope it helps!</p>",
      "rawMarkdown": "I've came up with the formula that I believe is giving proper values:\n\n    # T - no. of total samples\n    # P - no. of positive samples\n    # scale_pos_weight = percent of negative / percent of positive\n    # which translates to:\n    # scale_pos_weight = (100*(T-P)/T) / (100*P/T)\n    # which further simplifies to beautiful:\n    scale_pos_weight = T/P - 1\n\nI'm bit tired atm and I hope I haven't made an error - but I am sure you'll double-check it.\nHope it helps!",
      "votes": null
    },
    {
      "id": "308626",
      "postDate": "04/03/2018 20:10:53",
      "content": "<blockquote>\n  <p>I've came up with the formula that I believe is giving proper values:</p>\n</blockquote>\n\n<p>This seems to be correct. You can also let python <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53696\"><strong>do the counting and formatting</strong></a>.</p>",
      "rawMarkdown": "&gt; I've came up with the formula that I believe is giving proper values:\n\nThis seems to be correct. You can also let python [__do the counting and formatting__](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53696).",
      "votes": null
    },
    {
      "id": "308631",
      "postDate": "04/03/2018 20:21:51",
      "content": "<p>Thanks! Hahah, I've commented on your thread a minute ago :)</p>",
      "rawMarkdown": "Thanks! Hahah, I've commented on your thread a minute ago :)",
      "votes": null
    },
    {
      "id": "477895",
      "postDate": "02/25/2019 12:45:43",
      "content": "<p>It would be simply </p>\n\n<p>scale_pos_weight =sum(negative instances) / sum(positive instances)</p>",
      "rawMarkdown": "It would be simply \n\nscale_pos_weight =sum(negative instances) / sum(positive instances)",
      "votes": null
    },
    {
      "id": "761236",
      "postDate": "03/02/2020 09:53:28",
      "content": "<p>&gt; <strong>Asparuh Hristov wrote:</strong>\n&gt; \n&gt; SMORE (you can google it)</p>\n\n<p>You mean SMOTE (Synthetic Minority Oversampling Technique)</p>",
      "rawMarkdown": "&gt; **Asparuh Hristov wrote:**\n&gt; \n&gt; SMORE (you can google it)\n\nYou mean SMOTE (Synthetic Minority Oversampling Technique)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 294083,
      "author_name": "asparuhhristov",
      "author_url": "",
      "post_date": "03/11/2018 08:31:02",
      "content": "<p>I believe it is even more than 90%. Anyway, there are at least 3 conventional ways : undersampling those that have is_attributed == 0 , oversampling the is_attributed == 1 or SMORE (you can google it). I personally think a \"smart\" undersampling will be the way to go. We have a lot of data, and some of it is just for IPs that clicked once and never ever again returned... just an initial thought, for sure one can do even beter.</p>",
      "votes": null,
      "replies": [
        {
          "id": 761236,
          "author_name": "smcinerney",
          "author_url": "",
          "post_date": "03/02/2020 09:53:28",
          "content": "<p>&gt; <strong>Asparuh Hristov wrote:</strong>\n&gt; \n&gt; SMORE (you can google it)</p>\n\n<p>You mean SMOTE (Synthetic Minority Oversampling Technique)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 300174,
      "author_name": "shrivastavarpit",
      "author_url": "",
      "post_date": "03/21/2018 11:26:11",
      "content": "<p>As asparuh mentioned there are methods available like over and under sampling, but these have drawbacks as oversampling leads to overfit the data as we just duplicate rows, undersampling can always lead to loss of information for the \"Fraud\" class in our concerned data, there are various new modern techniques like SMOTE(Synthetic minority oversampling technique) this is a kind of oversample but here the rows arent dupicated but new minority class rows are produced using K nearest neighbour technique.\nThere are various oter techniques that combines boosting and data sampling techniques , like RUSboost, Smoteboost etc(Google it various good research papers are available)</p>",
      "votes": null,
      "replies": [
        {
          "id": 303105,
          "author_name": "arbardhan",
          "author_url": "",
          "post_date": "03/25/2018 14:24:09",
          "content": "<p>I used SMOTE + tsNE in a fraud detection algorithm . Does this mean its more or less acceptable to use SMOTE data to train a model ? </p>\n\n<p><a href=\"https://www.kaggle.com/mlg-ulb/creditcardfraud/discussion/52948\">https://www.kaggle.com/mlg-ulb/creditcardfraud/discussion/52948</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 308203,
          "author_name": "shrivastavarpit",
          "author_url": "",
          "post_date": "04/03/2018 05:50:34",
          "content": "<p>Smote data is used for training but as per my experience you should always keep a validation set aside before smote , then apply smote on the rest of the data and train it and first check on the validation set weather it is of some help or not.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 303116,
      "author_name": "pranav84",
      "author_url": "",
      "post_date": "03/25/2018 14:56:44",
      "content": "<p>If you're using LightGBM or XGBoost then I would recommend tuning <code>scale_pos_weight</code> parameter. </p>\n\n<ol>\n<li>Both algorithms (in R and Python) has <code>scale_pos_weight</code> argument which controls the weights of the positive observations which is very useful for unbalanced classes.</li>\n<li>According to LightGBM documentation, default value for scale_pos_weight is 1.0 and it represents weight of positive class in binary classification task. There are different ways to calculate weight for positive class. I have taken reference from LightGBM’s Github issues page for calculation.</li>\n</ol>\n\n<p>Below is the example from one of my LightGBM <a href=\"https://www.kaggle.com/pranav84/talkingdata-eda-to-model-evaluation-lb-0-9683/notebook\">model</a> which provides me following information based on training and validation data:</p>\n\n<pre><code>Number of positive: 143540, number of negative: 56856460 \nTotal Bins 1704 \nNumber of data: 57000000, number of used features: 11\n</code></pre>\n\n<p>With this information, I have calculated exact value for <code>scale_pos_weight</code> with following formula :</p>\n\n<pre><code>scale_pos_weight = 100 - ( [number of positive samples / total samples ] * 100 )\nscale_pos_weight = 100 - ( [ 143540 / 57000000 ] * 100 ) \nscale_pos_weight = 99.74\n</code></pre>\n\n<p>As Asparuh Hristov has correctly pointed out, in full dataset, number of negative classes are more than 90% i.e \n <strong>around 99.8%</strong>   </p>",
      "votes": null,
      "replies": [
        {
          "id": 308362,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/03/2018 12:12:22",
          "content": "<p>Your formula looks wrong to me.  You are setting scale_pos_weigh to be the percentage of negative samples.</p>\n\n<p>Your formula gives a scale_pos_weight of 50 instead of 1 for a balanced dataset.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 308418,
          "author_name": "pranav84",
          "author_url": "",
          "post_date": "04/03/2018 13:26:00",
          "content": "<p>Thanks a lot @CPMP. I also have doubt about this formula and unfortunately, there is no explanation in documentation. I appreciate your thoughts on this. I have now <a href=\"https://github.com/Microsoft/LightGBM/issues/1299\">created</a>  new issue on LightGBM repo to see if creators/contributors can provide correct procedure to calculate it. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 308621,
          "author_name": "konchar",
          "author_url": "",
          "post_date": "04/03/2018 19:59:09",
          "content": "<p>I've came up with the formula that I believe is giving proper values:</p>\n\n<pre><code># T - no. of total samples\n# P - no. of positive samples\n# scale_pos_weight = percent of negative / percent of positive\n# which translates to:\n# scale_pos_weight = (100*(T-P)/T) / (100*P/T)\n# which further simplifies to beautiful:\nscale_pos_weight = T/P - 1\n</code></pre>\n\n<p>I'm bit tired atm and I hope I haven't made an error - but I am sure you'll double-check it.\nHope it helps!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 308626,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "04/03/2018 20:10:53",
          "content": "<blockquote>\n  <p>I've came up with the formula that I believe is giving proper values:</p>\n</blockquote>\n\n<p>This seems to be correct. You can also let python <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53696\"><strong>do the counting and formatting</strong></a>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 308631,
          "author_name": "konchar",
          "author_url": "",
          "post_date": "04/03/2018 20:21:51",
          "content": "<p>Thanks! Hahah, I've commented on your thread a minute ago :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 477895,
          "author_name": "buntyshah",
          "author_url": "",
          "post_date": "02/25/2019 12:45:43",
          "content": "<p>It would be simply </p>\n\n<p>scale_pos_weight =sum(negative instances) / sum(positive instances)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "293558": "90% of is_attributed is 0",
    "294083": "I believe it is even more than 90%. Anyway, there are at least 3 conventional ways : undersampling those that have is_attributed == 0 , oversampling the is_attributed == 1 or SMORE (you can google it). I personally think a \"smart\" undersampling will be the way to go. We have a lot of data, and some of it is just for IPs that clicked once and never ever again returned... just an initial thought, for sure one can do even beter.",
    "300174": "As asparuh mentioned there are methods available like over and under sampling, but these have drawbacks as oversampling leads to overfit the data as we just duplicate rows, undersampling can always lead to loss of information for the \"Fraud\" class in our concerned data, there are various new modern techniques like SMOTE(Synthetic minority oversampling technique) this is a kind of oversample but here the rows arent dupicated but new minority class rows are produced using K nearest neighbour technique.\nThere are various oter techniques that combines boosting and data sampling techniques , like RUSboost, Smoteboost etc(Google it various good research papers are available)",
    "303105": "I used SMOTE + tsNE in a fraud detection algorithm . Does this mean its more or less acceptable to use SMOTE data to train a model ? \n\nhttps://www.kaggle.com/mlg-ulb/creditcardfraud/discussion/52948",
    "303116": "If you're using LightGBM or XGBoost then I would recommend tuning `scale_pos_weight` parameter. \n\n 1. Both algorithms (in R and Python) has `scale_pos_weight` argument which controls the weights of the positive observations which is very useful for unbalanced classes.\n 2. According to LightGBM documentation, default value for scale_pos_weight is 1.0 and it represents weight of positive class in binary classification task. There are different ways to calculate weight for positive class. I have taken reference from LightGBM’s Github issues page for calculation.\n\nBelow is the example from one of my LightGBM [model][1] which provides me following information based on training and validation data:\n\n    Number of positive: 143540, number of negative: 56856460 \n    Total Bins 1704 \n    Number of data: 57000000, number of used features: 11\n\n\nWith this information, I have calculated exact value for `scale_pos_weight` with following formula :\n\n    scale_pos_weight = 100 - ( [number of positive samples / total samples ] * 100 )\n    scale_pos_weight = 100 - ( [ 143540 / 57000000 ] * 100 ) \n    scale_pos_weight = 99.74\n\nAs Asparuh Hristov has correctly pointed out, in full dataset, number of negative classes are more than 90% i.e \n **around 99.8%**   \n\n  [1]: https://www.kaggle.com/pranav84/talkingdata-eda-to-model-evaluation-lb-0-9683/notebook",
    "308203": "Smote data is used for training but as per my experience you should always keep a validation set aside before smote , then apply smote on the rest of the data and train it and first check on the validation set weather it is of some help or not.",
    "308362": "Your formula looks wrong to me.  You are setting scale_pos_weigh to be the percentage of negative samples.\n\nYour formula gives a scale_pos_weight of 50 instead of 1 for a balanced dataset.",
    "308418": "Thanks a lot @CPMP. I also have doubt about this formula and unfortunately, there is no explanation in documentation. I appreciate your thoughts on this. I have now [created][1]  new issue on LightGBM repo to see if creators/contributors can provide correct procedure to calculate it. \n\n\n  [1]: https://github.com/Microsoft/LightGBM/issues/1299",
    "308621": "I've came up with the formula that I believe is giving proper values:\n\n    # T - no. of total samples\n    # P - no. of positive samples\n    # scale_pos_weight = percent of negative / percent of positive\n    # which translates to:\n    # scale_pos_weight = (100*(T-P)/T) / (100*P/T)\n    # which further simplifies to beautiful:\n    scale_pos_weight = T/P - 1\n\nI'm bit tired atm and I hope I haven't made an error - but I am sure you'll double-check it.\nHope it helps!",
    "308626": "&gt; I've came up with the formula that I believe is giving proper values:\n\nThis seems to be correct. You can also let python [__do the counting and formatting__](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53696).",
    "308631": "Thanks! Hahah, I've commented on your thread a minute ago :)",
    "477895": "It would be simply \n\nscale_pos_weight =sum(negative instances) / sum(positive instances)",
    "761236": "&gt; **Asparuh Hristov wrote:**\n&gt; \n&gt; SMORE (you can google it)\n\nYou mean SMOTE (Synthetic Minority Oversampling Technique)"
  },
  "source": "meta"
}