{
  "id": 53696,
  "title": "Calculating class weights",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53696",
  "author_name": "",
  "post_date": "2018-04-03T19:49:01.369336600Z",
  "votes": 35,
  "comment_count": 26,
  "views": 0,
  "content": "<p>There seems to be some confusion about calculating class weights, as this dataset is very imbalanced. This dataset has ~0.247% 1s, and the rest are 0s (~99.753%). <code>scale_pos_weight</code> of LightGBM is \"weight of positive class in binary classification task\" according to LightGBM documentation. I think that translates into a multiplication factor that has to be applied to number of 1s in order to get the same sample number as in 0s. So:</p>\n\n<p>99.753 / 0.247 = ~ 403.8</p>\n\n<p>If you want to do it programatically rather than using a calculator:</p>\n\n<pre><code>import pandas as pd\nfrom collections import Counter\n\ndef get_class_weights(y):\n    counter = Counter(y)\n    majority = max(counter.values())\n    return  {cls: round(float(majority)/float(count), 2) for cls, count in counter.items()}\n\ntrain = pd.read('train.csv')\nclass_weights = get_class_weights(train.is_attributed.values)\nprint(class_weights)\n</code></pre>\n\n<p>It will print <code>{0: 1.0, 1: 403.74}</code> and this can be used directly in Keras .fit function (<code>class_weight=class_weights</code>) or with LightGBM (<code>scale_pos_weight=class_weights[1]</code>).</p>\n\n<p>By the way, I do not recommend that you set class 1 weight to 400. This only shows how class weights are calculated.</p>",
  "messages": [
    {
      "id": "308618",
      "postDate": "04/03/2018 19:49:01",
      "content": "<p>There seems to be some confusion about calculating class weights, as this dataset is very imbalanced. This dataset has ~0.247% 1s, and the rest are 0s (~99.753%). <code>scale_pos_weight</code> of LightGBM is \"weight of positive class in binary classification task\" according to LightGBM documentation. I think that translates into a multiplication factor that has to be applied to number of 1s in order to get the same sample number as in 0s. So:</p>\n\n<p>99.753 / 0.247 = ~ 403.8</p>\n\n<p>If you want to do it programatically rather than using a calculator:</p>\n\n<pre><code>import pandas as pd\nfrom collections import Counter\n\ndef get_class_weights(y):\n    counter = Counter(y)\n    majority = max(counter.values())\n    return  {cls: round(float(majority)/float(count), 2) for cls, count in counter.items()}\n\ntrain = pd.read('train.csv')\nclass_weights = get_class_weights(train.is_attributed.values)\nprint(class_weights)\n</code></pre>\n\n<p>It will print <code>{0: 1.0, 1: 403.74}</code> and this can be used directly in Keras .fit function (<code>class_weight=class_weights</code>) or with LightGBM (<code>scale_pos_weight=class_weights[1]</code>).</p>\n\n<p>By the way, I do not recommend that you set class 1 weight to 400. This only shows how class weights are calculated.</p>",
      "rawMarkdown": "There seems to be some confusion about calculating class weights, as this dataset is very imbalanced. This dataset has ~0.247% 1s, and the rest are 0s (~99.753%). `scale_pos_weight` of LightGBM is \"weight of positive class in binary classification task\" according to LightGBM documentation. I think that translates into a multiplication factor that has to be applied to number of 1s in order to get the same sample number as in 0s. So:\n\n99.753 / 0.247 = ~ 403.8\n\nIf you want to do it programatically rather than using a calculator:\n\n    import pandas as pd\n    from collections import Counter\n    \n    def get_class_weights(y):\n        counter = Counter(y)\n        majority = max(counter.values())\n        return  {cls: round(float(majority)/float(count), 2) for cls, count in counter.items()}\n    \n    train = pd.read('train.csv')\n    class_weights = get_class_weights(train.is_attributed.values)\n    print(class_weights)\n\nIt will print `{0: 1.0, 1: 403.74}` and this can be used directly in Keras .fit function (`class_weight=class_weights`) or with LightGBM (`scale_pos_weight=class_weights[1]`).\n\nBy the way, I do not recommend that you set class 1 weight to 400. This only shows how class weights are calculated.",
      "votes": null
    },
    {
      "id": "308630",
      "postDate": "04/03/2018 20:19:52",
      "content": "<p>For fans of simple and clean formulas:<br>\nscale_pos_weight = T/P - 1 <br>\nwhere T is total no. of samples and P is no. of positive samples</p>",
      "rawMarkdown": "For fans of simple and clean formulas:<br>\nscale_pos_weight = T/P - 1 <br>\nwhere T is total no. of samples and P is no. of positive samples",
      "votes": null
    },
    {
      "id": "308633",
      "postDate": "04/03/2018 20:25:31",
      "content": "<p>Thanks @Tilii for knowledge sharing.</p>",
      "rawMarkdown": "Thanks @Tilii for knowledge sharing.",
      "votes": null
    },
    {
      "id": "308635",
      "postDate": "04/03/2018 20:35:09",
      "content": "<p>Thanks alot @Tilii. </p>",
      "rawMarkdown": "Thanks alot @Tilii.",
      "votes": null
    },
    {
      "id": "308637",
      "postDate": "04/03/2018 20:37:49",
      "content": "<p>Do you know how sklearn's <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.utils.class_weight.compute_class_weight.html\">compute_class_weight</a> is different from weights compute function provided by you, and when to use which one?<br> I ran following to get class weights </p>\n\n<pre><code>from sklearn.utils import class_weight\n\nclass_weight.compute_class_weight(class_weight='balanced', classes=[0,1]\\, y=train.is_attributed.values)\n</code></pre>\n\n<p>and got following weights for classes</p>\n\n<pre><code>[   0.50123842,  202.37004373]\n</code></pre>\n\n<p>Thanks</p>",
      "rawMarkdown": "Do you know how sklearn's [compute_class_weight][1] is different from weights compute function provided by you, and when to use which one?<br> I ran following to get class weights \n\n    from sklearn.utils import class_weight\n\n    class_weight.compute_class_weight(class_weight='balanced', classes=[0,1]\\, y=train.is_attributed.values)\nand got following weights for classes\n\n    [   0.50123842,  202.37004373]\n\nThanks\n\n  [1]: http://scikit-learn.org/stable/modules/generated/sklearn.utils.class_weight.compute_class_weight.html",
      "votes": null
    },
    {
      "id": "308643",
      "postDate": "04/03/2018 20:51:10",
      "content": "<p>If I'm not mistaken, only the relative weights matter, so the result sklearn calculated is equivalent to what Tilli's function calculated above.  I'm not sure why the scale is different, but it shouldn't matter for the results.</p>",
      "rawMarkdown": "If I'm not mistaken, only the relative weights matter, so the result sklearn calculated is equivalent to what Tilli's function calculated above.  I'm not sure why the scale is different, but it shouldn't matter for the results.",
      "votes": null
    },
    {
      "id": "308644",
      "postDate": "04/03/2018 20:56:27",
      "content": "<blockquote>\n  <p>Do you know how sklearn's compute_class_weight is different from weights compute function provided by you, and when to use which one?</p>\n</blockquote>\n\n<p>Yup, @Andy Harless is correct that only relative weights matter. The result above is equivalent to what I got for the purposes of modules that need class weight as a tuple.</p>\n\n<p>However, note that <code>scale_pos_weight</code> has to be specified as a single number, and this is where sklearn's calculation would fail.</p>",
      "rawMarkdown": "&gt; Do you know how sklearn's compute_class_weight is different from weights compute function provided by you, and when to use which one?\n\nYup, @Andy Harless is correct that only relative weights matter. The result above is equivalent to what I got for the purposes of modules that need class weight as a tuple.\n\nHowever, note that `scale_pos_weight` has to be specified as a single number, and this is where sklearn's calculation would fail.",
      "votes": null
    },
    {
      "id": "308750",
      "postDate": "04/04/2018 02:47:11",
      "content": "<p>However, I found something interesting. In LGB, if you set scale_pos_weight smaller than 403.74(for example, 300), you'll get a better score. Maybe it's better to let the model know that the data is somewhat unbalanced. :)</p>",
      "rawMarkdown": "However, I found something interesting. In LGB, if you set scale_pos_weight smaller than 403.74(for example, 300), you'll get a better score. Maybe it's better to let the model know that the data is somewhat unbalanced. :)",
      "votes": null
    },
    {
      "id": "308751",
      "postDate": "04/04/2018 02:52:57",
      "content": "<p>Doesn't scale_pos_weight = 403.74 tells LGB the data is imbalanced? Emmm...Why a smaller scale_pos_weight like 300 would outperform 403? And could we ignore scale_pos_weight and just set is_unbalanced = True?</p>",
      "rawMarkdown": "Doesn't scale_pos_weight = 403.74 tells LGB the data is imbalanced? Emmm...Why a smaller scale_pos_weight like 300 would outperform 403? And could we ignore scale_pos_weight and just set is_unbalanced = True?",
      "votes": null
    },
    {
      "id": "308756",
      "postDate": "04/04/2018 03:06:31",
      "content": "<p>I'd second this. In my experience it's typically been the case that an intermediate class weight multiplier (between 1 and the full multiplier to balance) works better.   </p>",
      "rawMarkdown": "I'd second this. In my experience it's typically been the case that an intermediate class weight multiplier (between 1 and the full multiplier to balance) works better.",
      "votes": null
    },
    {
      "id": "308757",
      "postDate": "04/04/2018 03:08:10",
      "content": "<p>Based on my limited experience with lgb , I have noticed <code>scale_pos_weight</code> mostly does a better job(optimizing metrics ) than <code>is_unbalanced = True</code> . </p>",
      "rawMarkdown": "Based on my limited experience with lgb , I have noticed `scale_pos_weight` mostly does a better job(optimizing metrics ) than `is_unbalanced = True` .",
      "votes": null
    },
    {
      "id": "308792",
      "postDate": "04/04/2018 05:36:57",
      "content": "<p>exactly</p>",
      "rawMarkdown": "exactly",
      "votes": null
    },
    {
      "id": "308794",
      "postDate": "04/04/2018 05:38:03",
      "content": "<p>scale_pos_weight is better from my experience, just the same as @shivraj said.</p>",
      "rawMarkdown": "scale_pos_weight is better from my experience, just the same as @shivraj said.",
      "votes": null
    },
    {
      "id": "308941",
      "postDate": "04/04/2018 12:04:05",
      "content": "<p>@Snorlax @Joe Eddy @shivraj @Laevatein</p>\n\n<p>All of you are saying the same thing, but I will try to formalize it. I think that using <code>is_unbalanced = True</code> is equivalent to setting <code>scale_pos_weight</code> to the true 1s/0s ratio, meaning that the parameter is fixed. On the other hand, setting <code>scale_pos_weight</code> allows us to tune that number as it frequently should be lower than its \"true\" value.</p>\n\n<p>I think <code>scale_pos_weight</code> should be tunable like any other hyperparameter for datasets that are more than 10x imbalanced.</p>",
      "rawMarkdown": "Snorlax @Joe Eddy @shivraj @Laevatein\n\nAll of you are saying the same thing, but I will try to formalize it. I think that using `is_unbalanced = True` is equivalent to setting `scale_pos_weight` to the true 1s/0s ratio, meaning that the parameter is fixed. On the other hand, setting `scale_pos_weight` allows us to tune that number as it frequently should be lower than its \"true\" value.\n\nI think `scale_pos_weight` should be tunable like any other hyperparameter for datasets that are more than 10x imbalanced.",
      "votes": null
    },
    {
      "id": "309034",
      "postDate": "04/04/2018 14:32:53",
      "content": "<p>I prefer to think of it as \n$$\ns \\cdot P = N = T - P\n$$ \nand then solve for <em>s</em>.</p>\n\n<p>I kid, of course. :P</p>",
      "rawMarkdown": "I prefer to think of it as \n$$\ns \\cdot P = N = T - P\n$$ \nand then solve for *s*.\n\nI kid, of course. :P",
      "votes": null
    },
    {
      "id": "309075",
      "postDate": "04/04/2018 15:44:01",
      "content": "<p>Thank you very much for your help guys! upvote upvote!</p>",
      "rawMarkdown": "Thank you very much for your help guys! upvote upvote!",
      "votes": null
    },
    {
      "id": "309136",
      "postDate": "04/04/2018 17:41:58",
      "content": "<p>Hahahha, that was funny!</p>",
      "rawMarkdown": "Hahahha, that was funny!",
      "votes": null
    },
    {
      "id": "309334",
      "postDate": "04/05/2018 04:18:24",
      "content": "<p>@Sohaib Omar,\nIt's the almost same weight. That is, if you re-computed the ratio and re-base on the negative class (labeled as 0), you would get the same weight. </p>\n\n<blockquote>\n  <p>class_weights <br>\n  array([  0.50123842, 202.37004373]) <br>\n  class_weights[1] / class_weights[0] <br>\n  403.7400874693004   </p>\n</blockquote>\n\n<p>But based on the document link you provided, I'm not sure if the result would be the same when weighting on multiple classes since the equation used by LightGBM seems to only consider binary classes. <br>\nAlso I'm not sure if using the class weight computed by scikit-learn would introduce more numerical error than using 1.0 to 403.xxx.  I could imagine weights computed by scikit-learn would introduce more unstable numerical computation if applying to the online training setting (batch mode) where the negative examples is much smaller than the total number. However, I didn't go through the implementation details. There might be no any difference if class weights are eventually normalized to one in the probability sense. <br>\nI also wonder if it's more reasonable to use sample_weight when fitting batch of training examples.  Using sample_weight seems to have advantage of dynamically weighting. <br>\nI wonder if any expert is so kind as to answering these questions or maybe give some insights how to analyze these questions more effectively.  Thanks! </p>",
      "rawMarkdown": "Sohaib Omar,\nIt's the almost same weight. That is, if you re-computed the ratio and re-base on the negative class (labeled as 0), you would get the same weight. \n\n &gt; class_weights   \narray([  0.50123842, 202.37004373])   \n&gt; class_weights[1] / class_weights[0]   \n403.7400874693004   \n\nBut based on the document link you provided, I'm not sure if the result would be the same when weighting on multiple classes since the equation used by LightGBM seems to only consider binary classes.    \nAlso I'm not sure if using the class weight computed by scikit-learn would introduce more numerical error than using 1.0 to 403.xxx.  I could imagine weights computed by scikit-learn would introduce more unstable numerical computation if applying to the online training setting (batch mode) where the negative examples is much smaller than the total number. However, I didn't go through the implementation details. There might be no any difference if class weights are eventually normalized to one in the probability sense.   \nI also wonder if it's more reasonable to use sample_weight when fitting batch of training examples.  Using sample_weight seems to have advantage of dynamically weighting.    \nI wonder if any expert is so kind as to answering these questions or maybe give some insights how to analyze these questions more effectively.  Thanks!",
      "votes": null
    },
    {
      "id": "309469",
      "postDate": "04/05/2018 12:16:34",
      "content": "<p>Thanks @Tilii for sharing this knowledge. One question: Why do you not recommend class weight 400 though it's just slightly smaller than 403? Is it a huge deal?</p>",
      "rawMarkdown": "Thanks @Tilii for sharing this knowledge. One question: Why do you not recommend class weight 400 though it's just slightly smaller than 403? Is it a huge deal?",
      "votes": null
    },
    {
      "id": "309474",
      "postDate": "04/05/2018 12:23:44",
      "content": "<p>It's not about the difference between 400 and 403. Experience shows that very large but formally correct values often show worse performance than smaller values (up to ~100). <code>scale_pos_weight</code> should probably be treated as a tunable hyperparameter.</p>",
      "rawMarkdown": "It's not about the difference between 400 and 403. Experience shows that very large but formally correct values often show worse performance than smaller values (up to ~100). `scale_pos_weight` should probably be treated as a tunable hyperparameter.",
      "votes": null
    },
    {
      "id": "309475",
      "postDate": "04/05/2018 12:27:48",
      "content": "<p>I think the reason may be that we have min_sample_leaf parameter is set. If we assume each positive sample as 400 samples, then this parameter is as effective as x/400. Then the model can overfit easier.</p>",
      "rawMarkdown": "I think the reason may be that we have min_sample_leaf parameter is set. If we assume each positive sample as 400 samples, then this parameter is as effective as x/400. Then the model can overfit easier.",
      "votes": null
    },
    {
      "id": "309480",
      "postDate": "04/05/2018 12:35:59",
      "content": "<p>Hi @AhmetErdem,  to my understanding. Do you mean, if we set min_sample_leaf = 400 and scale_pos_weight = 400, then some leaf may just have 1 positive sample because this positive sample has weight 400 by itself. And in this case, overfitting happens.</p>",
      "rawMarkdown": "Hi @AhmetErdem,  to my understanding. Do you mean, if we set min_sample_leaf = 400 and scale_pos_weight = 400, then some leaf may just have 1 positive sample because this positive sample has weight 400 by itself. And in this case, overfitting happens.",
      "votes": null
    },
    {
      "id": "309489",
      "postDate": "04/05/2018 13:06:27",
      "content": "<blockquote>\n  <p>I think the reason may be that we have min_sample_leaf parameter is set. If we assume each positive sample as 400 samples, then this parameter is as effective as x/400. Then the model can overfit easier.</p>\n</blockquote>\n\n<p>You're right.  One positive is worth 400 samples.  But negative samples are unweighted, hence the effect of scaling positive class has a mixed effect with min_sample_leaf.</p>",
      "rawMarkdown": "&gt; I think the reason may be that we have min_sample_leaf parameter is set. If we assume each positive sample as 400 samples, then this parameter is as effective as x/400. Then the model can overfit easier.\n\nYou're right.  One positive is worth 400 samples.  But negative samples are unweighted, hence the effect of scaling positive class has a mixed effect with min_sample_leaf.",
      "votes": null
    },
    {
      "id": "309608",
      "postDate": "04/05/2018 17:12:57",
      "content": "<blockquote>\n  <p>I think the reason may be that we have min_sample_leaf parameter is set. If we assume each positive sample as 400 samples, then this parameter is as effective as x/400. Then the model can overfit easier.</p>\n</blockquote>\n\n<p>@AhmetErdem What you describe is probably not the only problem when artificially augmenting the data by weighting. Though I don't have anything but my intuition to vouch for it, I deal with hyperparameter optimization of heavily weighted data by allowing large <code>min_child_weight</code> values.</p>",
      "rawMarkdown": "&gt; I think the reason may be that we have min_sample_leaf parameter is set. If we assume each positive sample as 400 samples, then this parameter is as effective as x/400. Then the model can overfit easier.\n\n@AhmetErdem What you describe is probably not the only problem when artificially augmenting the data by weighting. Though I don't have anything but my intuition to vouch for it, I deal with hyperparameter optimization of heavily weighted data by allowing large `min_child_weight` values.",
      "votes": null
    },
    {
      "id": "309700",
      "postDate": "04/05/2018 20:38:17",
      "content": "<p>Thanks for sharing this!</p>",
      "rawMarkdown": "Thanks for sharing this!",
      "votes": null
    },
    {
      "id": "310126",
      "postDate": "04/06/2018 16:56:28",
      "content": "<p>An <a href=\"https://github.com/Microsoft/LightGBM/issues/1299\"> <strong>explanation</strong> </a>  received from one of the LightGBM creators/contributors <a href=\"https://www.kaggle.com/laurae2\">laurae</a>  (who is also a Kaggler :)) regarding <code>scale_pos_weight</code> parameter. </p>\n\n<p>More simple explanation: <a href=\"https://sites.google.com/view/lauraepp/parameters\">https://sites.google.com/view/lauraepp/parameters</a> and type \"scale\" in the search box, then click on \"Positive Binary Scaling\".</p>\n\n<p><img src=\"https://user-images.githubusercontent.com/9083669/38409445-c9513d94-3981-11e8-8f11-1baebf26ef24.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "An [ **explanation** ][1]  received from one of the LightGBM creators/contributors [laurae][2]  (who is also a Kaggler :)) regarding `scale_pos_weight` parameter. \n\nMore simple explanation: https://sites.google.com/view/lauraepp/parameters and type \"scale\" in the search box, then click on \"Positive Binary Scaling\".\n\n![enter image description here][3]\n\n\n  [1]: https://github.com/Microsoft/LightGBM/issues/1299\n  [2]: https://www.kaggle.com/laurae2\n  [3]: https://user-images.githubusercontent.com/9083669/38409445-c9513d94-3981-11e8-8f11-1baebf26ef24.png",
      "votes": null
    },
    {
      "id": "761234",
      "postDate": "03/02/2020 09:50:01",
      "content": "<p>&gt; <strong>Tilli wrote:</strong>\n&gt; \n&gt; ...I deal with hyperparameter optimization of heavily weighted data by allowing large <code>min_child_weight</code> values</p>\n\n<p>Roughly what <code>min_child_weight</code> values do you use, to prevent overfitting to one positive exemplar? Do you use (say) <code>2 * scale_pos_weight</code> = <code>2*400</code>?</p>",
      "rawMarkdown": "&gt; **Tilli wrote:**\n&gt; \n&gt; ...I deal with hyperparameter optimization of heavily weighted data by allowing large `min_child_weight` values\n\nRoughly what `min_child_weight` values do you use, to prevent overfitting to one positive exemplar? Do you use (say) `2 * scale_pos_weight` = `2*400`?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 308630,
      "author_name": "konchar",
      "author_url": "",
      "post_date": "04/03/2018 20:19:52",
      "content": "<p>For fans of simple and clean formulas:<br>\nscale_pos_weight = T/P - 1 <br>\nwhere T is total no. of samples and P is no. of positive samples</p>",
      "votes": null,
      "replies": [
        {
          "id": 309034,
          "author_name": "puremath86",
          "author_url": "",
          "post_date": "04/04/2018 14:32:53",
          "content": "<p>I prefer to think of it as \n$$\ns \\cdot P = N = T - P\n$$ \nand then solve for <em>s</em>.</p>\n\n<p>I kid, of course. :P</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 309136,
          "author_name": "konchar",
          "author_url": "",
          "post_date": "04/04/2018 17:41:58",
          "content": "<p>Hahahha, that was funny!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 308633,
      "author_name": "sohaibomar",
      "author_url": "",
      "post_date": "04/03/2018 20:25:31",
      "content": "<p>Thanks @Tilii for knowledge sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 308637,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "04/03/2018 20:37:49",
          "content": "<p>Do you know how sklearn's <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.utils.class_weight.compute_class_weight.html\">compute_class_weight</a> is different from weights compute function provided by you, and when to use which one?<br> I ran following to get class weights </p>\n\n<pre><code>from sklearn.utils import class_weight\n\nclass_weight.compute_class_weight(class_weight='balanced', classes=[0,1]\\, y=train.is_attributed.values)\n</code></pre>\n\n<p>and got following weights for classes</p>\n\n<pre><code>[   0.50123842,  202.37004373]\n</code></pre>\n\n<p>Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 308643,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "04/03/2018 20:51:10",
          "content": "<p>If I'm not mistaken, only the relative weights matter, so the result sklearn calculated is equivalent to what Tilli's function calculated above.  I'm not sure why the scale is different, but it shouldn't matter for the results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 308644,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "04/03/2018 20:56:27",
          "content": "<blockquote>\n  <p>Do you know how sklearn's compute_class_weight is different from weights compute function provided by you, and when to use which one?</p>\n</blockquote>\n\n<p>Yup, @Andy Harless is correct that only relative weights matter. The result above is equivalent to what I got for the purposes of modules that need class weight as a tuple.</p>\n\n<p>However, note that <code>scale_pos_weight</code> has to be specified as a single number, and this is where sklearn's calculation would fail.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 309334,
          "author_name": "renewang",
          "author_url": "",
          "post_date": "04/05/2018 04:18:24",
          "content": "<p>@Sohaib Omar,\nIt's the almost same weight. That is, if you re-computed the ratio and re-base on the negative class (labeled as 0), you would get the same weight. </p>\n\n<blockquote>\n  <p>class_weights <br>\n  array([  0.50123842, 202.37004373]) <br>\n  class_weights[1] / class_weights[0] <br>\n  403.7400874693004   </p>\n</blockquote>\n\n<p>But based on the document link you provided, I'm not sure if the result would be the same when weighting on multiple classes since the equation used by LightGBM seems to only consider binary classes. <br>\nAlso I'm not sure if using the class weight computed by scikit-learn would introduce more numerical error than using 1.0 to 403.xxx.  I could imagine weights computed by scikit-learn would introduce more unstable numerical computation if applying to the online training setting (batch mode) where the negative examples is much smaller than the total number. However, I didn't go through the implementation details. There might be no any difference if class weights are eventually normalized to one in the probability sense. <br>\nI also wonder if it's more reasonable to use sample_weight when fitting batch of training examples.  Using sample_weight seems to have advantage of dynamically weighting. <br>\nI wonder if any expert is so kind as to answering these questions or maybe give some insights how to analyze these questions more effectively.  Thanks! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 308635,
      "author_name": "pranav84",
      "author_url": "",
      "post_date": "04/03/2018 20:35:09",
      "content": "<p>Thanks alot @Tilii. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 308750,
      "author_name": "laevatein",
      "author_url": "",
      "post_date": "04/04/2018 02:47:11",
      "content": "<p>However, I found something interesting. In LGB, if you set scale_pos_weight smaller than 403.74(for example, 300), you'll get a better score. Maybe it's better to let the model know that the data is somewhat unbalanced. :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 308751,
          "author_name": "wythhh",
          "author_url": "",
          "post_date": "04/04/2018 02:52:57",
          "content": "<p>Doesn't scale_pos_weight = 403.74 tells LGB the data is imbalanced? Emmm...Why a smaller scale_pos_weight like 300 would outperform 403? And could we ignore scale_pos_weight and just set is_unbalanced = True?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 308756,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "04/04/2018 03:06:31",
          "content": "<p>I'd second this. In my experience it's typically been the case that an intermediate class weight multiplier (between 1 and the full multiplier to balance) works better.   </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 308757,
          "author_name": "shivrajp",
          "author_url": "",
          "post_date": "04/04/2018 03:08:10",
          "content": "<p>Based on my limited experience with lgb , I have noticed <code>scale_pos_weight</code> mostly does a better job(optimizing metrics ) than <code>is_unbalanced = True</code> . </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 308792,
          "author_name": "laevatein",
          "author_url": "",
          "post_date": "04/04/2018 05:36:57",
          "content": "<p>exactly</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 308794,
          "author_name": "laevatein",
          "author_url": "",
          "post_date": "04/04/2018 05:38:03",
          "content": "<p>scale_pos_weight is better from my experience, just the same as @shivraj said.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 308941,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "04/04/2018 12:04:05",
          "content": "<p>@Snorlax @Joe Eddy @shivraj @Laevatein</p>\n\n<p>All of you are saying the same thing, but I will try to formalize it. I think that using <code>is_unbalanced = True</code> is equivalent to setting <code>scale_pos_weight</code> to the true 1s/0s ratio, meaning that the parameter is fixed. On the other hand, setting <code>scale_pos_weight</code> allows us to tune that number as it frequently should be lower than its \"true\" value.</p>\n\n<p>I think <code>scale_pos_weight</code> should be tunable like any other hyperparameter for datasets that are more than 10x imbalanced.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 309075,
          "author_name": "wythhh",
          "author_url": "",
          "post_date": "04/04/2018 15:44:01",
          "content": "<p>Thank you very much for your help guys! upvote upvote!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 309475,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "04/05/2018 12:27:48",
          "content": "<p>I think the reason may be that we have min_sample_leaf parameter is set. If we assume each positive sample as 400 samples, then this parameter is as effective as x/400. Then the model can overfit easier.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 309480,
          "author_name": "wythhh",
          "author_url": "",
          "post_date": "04/05/2018 12:35:59",
          "content": "<p>Hi @AhmetErdem,  to my understanding. Do you mean, if we set min_sample_leaf = 400 and scale_pos_weight = 400, then some leaf may just have 1 positive sample because this positive sample has weight 400 by itself. And in this case, overfitting happens.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 309489,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/05/2018 13:06:27",
          "content": "<blockquote>\n  <p>I think the reason may be that we have min_sample_leaf parameter is set. If we assume each positive sample as 400 samples, then this parameter is as effective as x/400. Then the model can overfit easier.</p>\n</blockquote>\n\n<p>You're right.  One positive is worth 400 samples.  But negative samples are unweighted, hence the effect of scaling positive class has a mixed effect with min_sample_leaf.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 309608,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "04/05/2018 17:12:57",
          "content": "<blockquote>\n  <p>I think the reason may be that we have min_sample_leaf parameter is set. If we assume each positive sample as 400 samples, then this parameter is as effective as x/400. Then the model can overfit easier.</p>\n</blockquote>\n\n<p>@AhmetErdem What you describe is probably not the only problem when artificially augmenting the data by weighting. Though I don't have anything but my intuition to vouch for it, I deal with hyperparameter optimization of heavily weighted data by allowing large <code>min_child_weight</code> values.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761234,
          "author_name": "smcinerney",
          "author_url": "",
          "post_date": "03/02/2020 09:50:01",
          "content": "<p>&gt; <strong>Tilli wrote:</strong>\n&gt; \n&gt; ...I deal with hyperparameter optimization of heavily weighted data by allowing large <code>min_child_weight</code> values</p>\n\n<p>Roughly what <code>min_child_weight</code> values do you use, to prevent overfitting to one positive exemplar? Do you use (say) <code>2 * scale_pos_weight</code> = <code>2*400</code>?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 309469,
      "author_name": "enfeizhan",
      "author_url": "",
      "post_date": "04/05/2018 12:16:34",
      "content": "<p>Thanks @Tilii for sharing this knowledge. One question: Why do you not recommend class weight 400 though it's just slightly smaller than 403? Is it a huge deal?</p>",
      "votes": null,
      "replies": [
        {
          "id": 309474,
          "author_name": "joergdietrich",
          "author_url": "",
          "post_date": "04/05/2018 12:23:44",
          "content": "<p>It's not about the difference between 400 and 403. Experience shows that very large but formally correct values often show worse performance than smaller values (up to ~100). <code>scale_pos_weight</code> should probably be treated as a tunable hyperparameter.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 309700,
      "author_name": "kmsbmadhan",
      "author_url": "",
      "post_date": "04/05/2018 20:38:17",
      "content": "<p>Thanks for sharing this!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 310126,
      "author_name": "pranav84",
      "author_url": "",
      "post_date": "04/06/2018 16:56:28",
      "content": "<p>An <a href=\"https://github.com/Microsoft/LightGBM/issues/1299\"> <strong>explanation</strong> </a>  received from one of the LightGBM creators/contributors <a href=\"https://www.kaggle.com/laurae2\">laurae</a>  (who is also a Kaggler :)) regarding <code>scale_pos_weight</code> parameter. </p>\n\n<p>More simple explanation: <a href=\"https://sites.google.com/view/lauraepp/parameters\">https://sites.google.com/view/lauraepp/parameters</a> and type \"scale\" in the search box, then click on \"Positive Binary Scaling\".</p>\n\n<p><img src=\"https://user-images.githubusercontent.com/9083669/38409445-c9513d94-3981-11e8-8f11-1baebf26ef24.png\" alt=\"enter image description here\"></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "308618": "There seems to be some confusion about calculating class weights, as this dataset is very imbalanced. This dataset has ~0.247% 1s, and the rest are 0s (~99.753%). `scale_pos_weight` of LightGBM is \"weight of positive class in binary classification task\" according to LightGBM documentation. I think that translates into a multiplication factor that has to be applied to number of 1s in order to get the same sample number as in 0s. So:\n\n99.753 / 0.247 = ~ 403.8\n\nIf you want to do it programatically rather than using a calculator:\n\n    import pandas as pd\n    from collections import Counter\n    \n    def get_class_weights(y):\n        counter = Counter(y)\n        majority = max(counter.values())\n        return  {cls: round(float(majority)/float(count), 2) for cls, count in counter.items()}\n    \n    train = pd.read('train.csv')\n    class_weights = get_class_weights(train.is_attributed.values)\n    print(class_weights)\n\nIt will print `{0: 1.0, 1: 403.74}` and this can be used directly in Keras .fit function (`class_weight=class_weights`) or with LightGBM (`scale_pos_weight=class_weights[1]`).\n\nBy the way, I do not recommend that you set class 1 weight to 400. This only shows how class weights are calculated.",
    "308630": "For fans of simple and clean formulas:<br>\nscale_pos_weight = T/P - 1 <br>\nwhere T is total no. of samples and P is no. of positive samples",
    "308633": "Thanks @Tilii for knowledge sharing.",
    "308635": "Thanks alot @Tilii.",
    "308637": "Do you know how sklearn's [compute_class_weight][1] is different from weights compute function provided by you, and when to use which one?<br> I ran following to get class weights \n\n    from sklearn.utils import class_weight\n\n    class_weight.compute_class_weight(class_weight='balanced', classes=[0,1]\\, y=train.is_attributed.values)\nand got following weights for classes\n\n    [   0.50123842,  202.37004373]\n\nThanks\n\n  [1]: http://scikit-learn.org/stable/modules/generated/sklearn.utils.class_weight.compute_class_weight.html",
    "308643": "If I'm not mistaken, only the relative weights matter, so the result sklearn calculated is equivalent to what Tilli's function calculated above.  I'm not sure why the scale is different, but it shouldn't matter for the results.",
    "308644": "&gt; Do you know how sklearn's compute_class_weight is different from weights compute function provided by you, and when to use which one?\n\nYup, @Andy Harless is correct that only relative weights matter. The result above is equivalent to what I got for the purposes of modules that need class weight as a tuple.\n\nHowever, note that `scale_pos_weight` has to be specified as a single number, and this is where sklearn's calculation would fail.",
    "308750": "However, I found something interesting. In LGB, if you set scale_pos_weight smaller than 403.74(for example, 300), you'll get a better score. Maybe it's better to let the model know that the data is somewhat unbalanced. :)",
    "308751": "Doesn't scale_pos_weight = 403.74 tells LGB the data is imbalanced? Emmm...Why a smaller scale_pos_weight like 300 would outperform 403? And could we ignore scale_pos_weight and just set is_unbalanced = True?",
    "308756": "I'd second this. In my experience it's typically been the case that an intermediate class weight multiplier (between 1 and the full multiplier to balance) works better.",
    "308757": "Based on my limited experience with lgb , I have noticed `scale_pos_weight` mostly does a better job(optimizing metrics ) than `is_unbalanced = True` .",
    "308792": "exactly",
    "308794": "scale_pos_weight is better from my experience, just the same as @shivraj said.",
    "308941": "Snorlax @Joe Eddy @shivraj @Laevatein\n\nAll of you are saying the same thing, but I will try to formalize it. I think that using `is_unbalanced = True` is equivalent to setting `scale_pos_weight` to the true 1s/0s ratio, meaning that the parameter is fixed. On the other hand, setting `scale_pos_weight` allows us to tune that number as it frequently should be lower than its \"true\" value.\n\nI think `scale_pos_weight` should be tunable like any other hyperparameter for datasets that are more than 10x imbalanced.",
    "309034": "I prefer to think of it as \n$$\ns \\cdot P = N = T - P\n$$ \nand then solve for *s*.\n\nI kid, of course. :P",
    "309075": "Thank you very much for your help guys! upvote upvote!",
    "309136": "Hahahha, that was funny!",
    "309334": "Sohaib Omar,\nIt's the almost same weight. That is, if you re-computed the ratio and re-base on the negative class (labeled as 0), you would get the same weight. \n\n &gt; class_weights   \narray([  0.50123842, 202.37004373])   \n&gt; class_weights[1] / class_weights[0]   \n403.7400874693004   \n\nBut based on the document link you provided, I'm not sure if the result would be the same when weighting on multiple classes since the equation used by LightGBM seems to only consider binary classes.    \nAlso I'm not sure if using the class weight computed by scikit-learn would introduce more numerical error than using 1.0 to 403.xxx.  I could imagine weights computed by scikit-learn would introduce more unstable numerical computation if applying to the online training setting (batch mode) where the negative examples is much smaller than the total number. However, I didn't go through the implementation details. There might be no any difference if class weights are eventually normalized to one in the probability sense.   \nI also wonder if it's more reasonable to use sample_weight when fitting batch of training examples.  Using sample_weight seems to have advantage of dynamically weighting.    \nI wonder if any expert is so kind as to answering these questions or maybe give some insights how to analyze these questions more effectively.  Thanks!",
    "309469": "Thanks @Tilii for sharing this knowledge. One question: Why do you not recommend class weight 400 though it's just slightly smaller than 403? Is it a huge deal?",
    "309474": "It's not about the difference between 400 and 403. Experience shows that very large but formally correct values often show worse performance than smaller values (up to ~100). `scale_pos_weight` should probably be treated as a tunable hyperparameter.",
    "309475": "I think the reason may be that we have min_sample_leaf parameter is set. If we assume each positive sample as 400 samples, then this parameter is as effective as x/400. Then the model can overfit easier.",
    "309480": "Hi @AhmetErdem,  to my understanding. Do you mean, if we set min_sample_leaf = 400 and scale_pos_weight = 400, then some leaf may just have 1 positive sample because this positive sample has weight 400 by itself. And in this case, overfitting happens.",
    "309489": "&gt; I think the reason may be that we have min_sample_leaf parameter is set. If we assume each positive sample as 400 samples, then this parameter is as effective as x/400. Then the model can overfit easier.\n\nYou're right.  One positive is worth 400 samples.  But negative samples are unweighted, hence the effect of scaling positive class has a mixed effect with min_sample_leaf.",
    "309608": "&gt; I think the reason may be that we have min_sample_leaf parameter is set. If we assume each positive sample as 400 samples, then this parameter is as effective as x/400. Then the model can overfit easier.\n\n@AhmetErdem What you describe is probably not the only problem when artificially augmenting the data by weighting. Though I don't have anything but my intuition to vouch for it, I deal with hyperparameter optimization of heavily weighted data by allowing large `min_child_weight` values.",
    "309700": "Thanks for sharing this!",
    "310126": "An [ **explanation** ][1]  received from one of the LightGBM creators/contributors [laurae][2]  (who is also a Kaggler :)) regarding `scale_pos_weight` parameter. \n\nMore simple explanation: https://sites.google.com/view/lauraepp/parameters and type \"scale\" in the search box, then click on \"Positive Binary Scaling\".\n\n![enter image description here][3]\n\n\n  [1]: https://github.com/Microsoft/LightGBM/issues/1299\n  [2]: https://www.kaggle.com/laurae2\n  [3]: https://user-images.githubusercontent.com/9083669/38409445-c9513d94-3981-11e8-8f11-1baebf26ef24.png",
    "761234": "&gt; **Tilli wrote:**\n&gt; \n&gt; ...I deal with hyperparameter optimization of heavily weighted data by allowing large `min_child_weight` values\n\nRoughly what `min_child_weight` values do you use, to prevent overfitting to one positive exemplar? Do you use (say) `2 * scale_pos_weight` = `2*400`?"
  },
  "source": "meta"
}