{
  "id": 74065,
  "title": "If you want to play with class weights ...",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/74065",
  "author_name": "Tilii",
  "post_date": "2018-12-08T04:50:05.937000",
  "votes": 27,
  "comment_count": 19,
  "views": 0,
  "content": "<pre>    #mu in \"create_class_weight\" is a dampening parameter that could be tuned\n\n<code>import numpy as np\nimport math\n\ndef create_class_weight(labels_dict, mu=0.5):\n    total = np.sum(labels_dict.values())\n    keys = labels_dict.keys()\n    class_weight = dict()\n    class_weight_log = dict()\n\n    for key in keys:\n        score = total / float(labels_dict[key])\n        score_log = math.log(mu * total / float(labels_dict[key]))\n        class_weight[key] = round(score, 2) if score &gt; 1.0 else round(1.0, 2)\n        class_weight_log[key] = round(score_log, 2) if score_log &gt; 1.0 else round(1.0, 2)\n\n    return class_weight, class_weight_log\n\n# Class abundance for protein dataset\nlabels_dict = {\n    0: 12885,\n    1: 1254,\n    2: 3621,\n    3: 1561,\n    4: 1858,\n    5: 2513,\n    6: 1008,\n    7: 2822,\n    8: 53,\n    9: 45,\n    10: 28,\n    11: 1093,\n    12: 688,\n    13: 537,\n    14: 1066,\n    15: 21,\n    16: 530,\n    17: 210,\n    18: 902,\n    19: 1482,\n    20: 172,\n    21: 3777,\n    22: 802,\n    23: 2965,\n    24: 322,\n    25: 8228,\n    26: 328,\n    27: 11\n}\n\nprint('\\nTrue class weights:')\nprint(create_class_weight(labels_dict)[0])\nprint('\\nLog-dampened class weights:')\nprint(create_class_weight(labels_dict)[1])&lt;code&gt;\n</code></pre>",
  "messages": [
    {
      "id": 435480,
      "postDate": "2018-12-08T04:50:05.937Z",
      "content": "<pre>    #mu in \"create_class_weight\" is a dampening parameter that could be tuned\n\n<code>import numpy as np\nimport math\n\ndef create_class_weight(labels_dict, mu=0.5):\n    total = np.sum(labels_dict.values())\n    keys = labels_dict.keys()\n    class_weight = dict()\n    class_weight_log = dict()\n\n    for key in keys:\n        score = total / float(labels_dict[key])\n        score_log = math.log(mu * total / float(labels_dict[key]))\n        class_weight[key] = round(score, 2) if score &gt; 1.0 else round(1.0, 2)\n        class_weight_log[key] = round(score_log, 2) if score_log &gt; 1.0 else round(1.0, 2)\n\n    return class_weight, class_weight_log\n\n# Class abundance for protein dataset\nlabels_dict = {\n    0: 12885,\n    1: 1254,\n    2: 3621,\n    3: 1561,\n    4: 1858,\n    5: 2513,\n    6: 1008,\n    7: 2822,\n    8: 53,\n    9: 45,\n    10: 28,\n    11: 1093,\n    12: 688,\n    13: 537,\n    14: 1066,\n    15: 21,\n    16: 530,\n    17: 210,\n    18: 902,\n    19: 1482,\n    20: 172,\n    21: 3777,\n    22: 802,\n    23: 2965,\n    24: 322,\n    25: 8228,\n    26: 328,\n    27: 11\n}\n\nprint('\\nTrue class weights:')\nprint(create_class_weight(labels_dict)[0])\nprint('\\nLog-dampened class weights:')\nprint(create_class_weight(labels_dict)[1])&lt;code&gt;\n</code></pre>",
      "rawMarkdown": "<pre>    #mu in \"create_class_weight\" is a dampening parameter that could be tuned\n\n    import numpy as np\n    import math\n\n    def create_class_weight(labels_dict, mu=0.5):\n        total = np.sum(labels_dict.values())\n        keys = labels_dict.keys()\n        class_weight = dict()\n        class_weight_log = dict()\n\n        for key in keys:\n            score = total / float(labels_dict[key])\n            score_log = math.log(mu * total / float(labels_dict[key]))\n            class_weight[key] = round(score, 2) if score &gt; 1.0 else round(1.0, 2)\n            class_weight_log[key] = round(score_log, 2) if score_log &gt; 1.0 else round(1.0, 2)\n\n        return class_weight, class_weight_log\n\n    # Class abundance for protein dataset\n    labels_dict = {\n        0: 12885,\n        1: 1254,\n        2: 3621,\n        3: 1561,\n        4: 1858,\n        5: 2513,\n        6: 1008,\n        7: 2822,\n        8: 53,\n        9: 45,\n        10: 28,\n        11: 1093,\n        12: 688,\n        13: 537,\n        14: 1066,\n        15: 21,\n        16: 530,\n        17: 210,\n        18: 902,\n        19: 1482,\n        20: 172,\n        21: 3777,\n        22: 802,\n        23: 2965,\n        24: 322,\n        25: 8228,\n        26: 328,\n        27: 11\n    }\n\n    print('\\nTrue class weights:')\n    print(create_class_weight(labels_dict)[0])\n    print('\\nLog-dampened class weights:')\n    print(create_class_weight(labels_dict)[1])<code></code></pre>",
      "votes": 27
    },
    {
      "id": 436063,
      "postDate": "2018-12-09T13:45:41.757Z",
      "content": "<p>Another (simple) approach is just to use a linear interpolation of the prior probability of a label being in a class towards equal weights. </p>\n\n<p>In the below, setting mu = 0 is the original probability distribution and as mu increases we (linearly) tend to the distribution with all class weights equal (the maximum entropy distribution).</p>\n\n<pre><code>import numpy as np\n\nname_label_dict = {\n    0:   ('Nucleoplasm', 12885),\n    1:   ('Nuclear membrane', 1254),\n    2:   ('Nucleoli', 3621),\n    3:   ('Nucleoli fibrillar center', 1561),\n    4:   ('Nuclear speckles', 1858),\n    5:   ('Nuclear bodies', 2513),\n    6:   ('Endoplasmic reticulum', 1008),   \n    7:   ('Golgi apparatus', 2822),\n    8:   ('Peroxisomes', 53), \n    9:   ('Endosomes', 45),\n    10:  ('Lysosomes', 28),\n    11:  ('Intermediate filaments', 1093), \n    12:  ('Actin filaments', 688),\n    13:  ('Focal adhesion sites', 537),  \n    14:  ('Microtubules', 1066), \n    15:  ('Microtubule ends', 21),\n    16:  ('Cytokinetic bridge', 530),\n    17:  ('Mitotic spindle', 210),\n    18:  ('Microtubule organizing center', 902),\n    19:  ('Centrosome', 1482),\n    20:  ('Lipid droplets', 172),\n    21:  ('Plasma membrane', 3777),\n    22:  ('Cell junctions', 802),\n    23:  ('Mitochondria', 2965),\n    24:  ('Aggresome', 322),\n    25:  ('Cytosol', 8228),\n    26:  ('Cytoplasmic bodies', 328),   \n    27:  ('Rods &amp;amp; rings', 11)\n    }\n\nn_labels = 50782\n\ndef cls_wts(label_dict, mu=0.5):\n    prob_dict, prob_dict_bal = {}, {}\n    max_ent_wt = 1/28\n    for i in range(28):\n        prob_dict[i] = label_dict[i][1]/n_labels\n        if prob_dict[i] &amp;gt; max_ent_wt:\n            prob_dict_bal[i] = prob_dict[i]-mu*(prob_dict[i] - max_ent_wt)\n        else:\n            prob_dict_bal[i] = prob_dict[i]+mu*(max_ent_wt - prob_dict[i])            \n    return prob_dict, prob_dict_bal\n</code></pre>",
      "rawMarkdown": "Another (simple) approach is just to use a linear interpolation of the prior probability of a label being in a class towards equal weights. \n\nIn the below, setting mu = 0 is the original probability distribution and as mu increases we (linearly) tend to the distribution with all class weights equal (the maximum entropy distribution).\n\n\n    import numpy as np\n\n    name_label_dict = {\n        0:   ('Nucleoplasm', 12885),\n        1:   ('Nuclear membrane', 1254),\n        2:   ('Nucleoli', 3621),\n        3:   ('Nucleoli fibrillar center', 1561),\n        4:   ('Nuclear speckles', 1858),\n        5:   ('Nuclear bodies', 2513),\n        6:   ('Endoplasmic reticulum', 1008),   \n        7:   ('Golgi apparatus', 2822),\n        8:   ('Peroxisomes', 53), \n        9:   ('Endosomes', 45),\n        10:  ('Lysosomes', 28),\n        11:  ('Intermediate filaments', 1093), \n        12:  ('Actin filaments', 688),\n        13:  ('Focal adhesion sites', 537),  \n        14:  ('Microtubules', 1066), \n        15:  ('Microtubule ends', 21),\n        16:  ('Cytokinetic bridge', 530),\n        17:  ('Mitotic spindle', 210),\n        18:  ('Microtubule organizing center', 902),\n        19:  ('Centrosome', 1482),\n        20:  ('Lipid droplets', 172),\n        21:  ('Plasma membrane', 3777),\n        22:  ('Cell junctions', 802),\n        23:  ('Mitochondria', 2965),\n        24:  ('Aggresome', 322),\n        25:  ('Cytosol', 8228),\n        26:  ('Cytoplasmic bodies', 328),   \n        27:  ('Rods &amp; rings', 11)\n        }\n\n    n_labels = 50782\n\n    def cls_wts(label_dict, mu=0.5):\n        prob_dict, prob_dict_bal = {}, {}\n        max_ent_wt = 1/28\n        for i in range(28):\n            prob_dict[i] = label_dict[i][1]/n_labels\n            if prob_dict[i] &gt; max_ent_wt:\n                prob_dict_bal[i] = prob_dict[i]-mu*(prob_dict[i] - max_ent_wt)\n            else:\n                prob_dict_bal[i] = prob_dict[i]+mu*(max_ent_wt - prob_dict[i])            \n        return prob_dict, prob_dict_bal\n\n\n",
      "votes": 6,
      "replies": [
        {
          "id": 436147,
          "postDate": "2018-12-09T17:21:30.107Z",
          "content": "<blockquote>\n  <p>Another (simple) approach is just to use a linear interpolation of the prior probability of a label being in a class towards equal weights.</p>\n</blockquote>\n\n<p>My choice of scale was arbitrary, so linear dampening works fine. In fact, for most of us it may be easier to understand its effect than log adjustment.</p>",
          "rawMarkdown": "&gt; Another (simple) approach is just to use a linear interpolation of the prior probability of a label being in a class towards equal weights.\n\nMy choice of scale was arbitrary, so linear dampening works fine. In fact, for most of us it may be easier to understand its effect than log adjustment.",
          "votes": 1
        }
      ]
    },
    {
      "id": 439001,
      "postDate": "2018-12-14T14:56:09.873Z",
      "content": "<p>Wanted to report that my initial results are showing that the log-dampened weights do better than just the direct weights I was using before. Thank you for this!</p>",
      "rawMarkdown": "Wanted to report that my initial results are showing that the log-dampened weights do better than just the direct weights I was using before. Thank you for this!",
      "votes": 3,
      "replies": [
        {
          "id": 439145,
          "postDate": "2018-12-14T19:42:15.977Z",
          "content": "<p><a href=\"/hortonhearsafoo\">@hortonhearsafoo</a> Thank you for reporting it. I have dealt with enough imbalanced datasets to know that applying class weights based on true class proportions is almost never the right thing to do. It becomes a matter of finding a ratio that corrects the imbalance. I am sure that linear interpolation suggested by <a href=\"/maw501\">@maw501</a> works as well, though most likely with a different value of mu.</p>",
          "rawMarkdown": "@hortonhearsafoo Thank you for reporting it. I have dealt with enough imbalanced datasets to know that applying class weights based on true class proportions is almost never the right thing to do. It becomes a matter of finding a ratio that corrects the imbalance. I am sure that linear interpolation suggested by @maw501 works as well, though most likely with a different value of mu.",
          "votes": 1
        }
      ]
    },
    {
      "id": 440274,
      "postDate": "2018-12-17T10:14:08.177Z",
      "content": "<p>hi ,thanks for input...\ncan you please guide where to fit this one... imeant where to pass this one ,the weights  to the loss</p>",
      "rawMarkdown": "hi ,thanks for input...\ncan you please guide where to fit this one... imeant where to pass this one ,the weights  to the loss",
      "votes": 1,
      "replies": [
        {
          "id": 440664,
          "postDate": "2018-12-17T21:34:05.453Z",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> Use one of the script's outputs to define class weights:</p>\n\n<p><code>cw = {0: 1.0, 1: 3.01, 2: 1.95, 3: 2.79, 4: 2.61, ....  26: 4.35, 27: 7.74}</code></p>\n\n<p>This dictionary is used as an argument in keras <code>.fit</code> function:</p>\n\n<p><code>class_weight=cw</code></p>\n\n<p>Other classifiers use <code>class_weight</code> in a similar way.</p>",
          "rawMarkdown": "@jaideepvalani Use one of the script's outputs to define class weights:\n\n`cw = {0: 1.0, 1: 3.01, 2: 1.95, 3: 2.79, 4: 2.61, ....  26: 4.35, 27: 7.74}`\n\nThis dictionary is used as an argument in keras `.fit` function:\n\n`class_weight=cw`\n\nOther classifiers use `class_weight` in a similar way.",
          "votes": 1
        },
        {
          "id": 447824,
          "postDate": "2018-12-30T16:48:57.790Z",
          "content": "<p>hi,i tried using it with BCE cross entropy ,but val loss is very very fluctuating ,it reaches as high as 200,low less than 10 ,every alternate epochs..\nare the log scale values are not in scale of loss ??\ni get values similar as above.. </p>",
          "rawMarkdown": "hi,i tried using it with BCE cross entropy ,but val loss is very very fluctuating ,it reaches as high as 200,low less than 10 ,every alternate epochs..\nare the log scale values are not in scale of loss ??\ni get values similar as above.. \n"
        }
      ]
    },
    {
      "id": 437035,
      "postDate": "2018-12-11T09:25:25.953Z",
      "content": "<p>This is very nicely done, and easy to understand. Thank you! </p>",
      "rawMarkdown": "This is very nicely done, and easy to understand. Thank you! ",
      "votes": 1
    },
    {
      "id": 539014,
      "postDate": "2019-05-29T12:00:54.010Z",
      "content": "<p>Excuse me, should the parameter mu  be bigger than 1？Or mu must be in the range of (0,1)</p>",
      "rawMarkdown": "Excuse me, should the parameter mu  be bigger than 1？Or mu must be in the range of (0,1)"
    },
    {
      "id": 531070,
      "postDate": "2019-05-14T08:34:04.160Z",
      "content": "<p>Excuse me, is there any general rule of thumb when choosing the parameter <code>mu</code>？</p>",
      "rawMarkdown": "Excuse me, is there any general rule of thumb when choosing the parameter `mu`？",
      "replies": [
        {
          "id": 531344,
          "postDate": "2019-05-14T17:29:18.637Z",
          "content": "<blockquote>\n  <p>Excuse me, is there any general rule of thumb when choosing the parameter mu</p>\n</blockquote>\n\n<p><a href=\"/qunyang\">@qunyang</a> It is a parameter that needs to be tuned for each dataset. Larger values of <code>mu</code> mean the weights are more similar to true class ratios. Small values of <code>mu</code> dampen the difference.</p>",
          "rawMarkdown": "&gt; Excuse me, is there any general rule of thumb when choosing the parameter mu\n\n@qunyang It is a parameter that needs to be tuned for each dataset. Larger values of `mu` mean the weights are more similar to true class ratios. Small values of `mu` dampen the difference."
        },
        {
          "id": 531471,
          "postDate": "2019-05-15T01:26:04.227Z",
          "content": "<p>Thanks, I understand</p>",
          "rawMarkdown": "Thanks, I understand"
        }
      ]
    },
    {
      "id": 442278,
      "postDate": "2018-12-19T18:14:59.427Z",
      "content": "<p>hi thanks for explaination</p>\n\n<p>i get this error\nunsupported operand type(s) for /: 'dict_values' and 'float'\nwhile calculating score\nwhen i print total variable i get\ndict_values([12885, 1254, 3621,......\nprobably because of that it fails</p>\n\n<p>should this be a single value ??  sum of all the counts </p>",
      "rawMarkdown": "hi thanks for explaination\n\ni get this error\nunsupported operand type(s) for /: 'dict_values' and 'float'\nwhile calculating score\nwhen i print total variable i get\ndict_values([12885, 1254, 3621,......\nprobably because of that it fails\n\nshould this be a single value ??  sum of all the counts \n",
      "replies": [
        {
          "id": 442285,
          "postDate": "2018-12-19T18:25:01.767Z",
          "content": "<p>I'm not clear what error you are having but I think I recall in the code Tilii provided needing to wrap the <code>labels_dict.values()</code> in a <code>list</code> as I think it's based on python 2 - Tilii can correct me if I'm wrong.</p>",
          "rawMarkdown": "I'm not clear what error you are having but I think I recall in the code Tilii provided needing to wrap the `labels_dict.values()` in a `list` as I think it's based on python 2 - Tilii can correct me if I'm wrong.",
          "votes": 2
        },
        {
          "id": 442303,
          "postDate": "2018-12-19T19:03:37.313Z",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> As explained by <a href=\"/maw501\">@maw501</a> (Mark), this code works in python 2. If you are using python 3, change this line:</p>\n\n<p><code>total = np.sum(labels_dict.values())</code></p>\n\n<p>to:</p>\n\n<p><code>total = 50782</code></p>\n\n<p>It should work after that. The output looks like this:</p>\n\n<pre><code>True class weights:\n{0: 3.94, 1: 40.5, 2: 14.02, 3: 32.53, 4: 27.33, 5: 20.21, 6: 50.38, 7: 18.0, 8: 958.15, 9: 1128.49, 10: 1813.64, 11: 46.46, 12: 73.81, 13: 94.57, 14: 47.64, 15: 2418.19, 16: 95.82, 17: 241.82, 18: 56.3, 19: 34.27, 20: 295.24, 21: 13.45, 22: 63.32, 23: 17.13, 24: 157.71, 25: 6.17, 26: 154.82, 27: 4616.55}\n\nLog-dampened class weights:\n{0: 1.0, 1: 3.01, 2: 1.95, 3: 2.79, 4: 2.61, 5: 2.31, 6: 3.23, 7: 2.2, 8: 6.17, 9: 6.34, 10: 6.81, 11: 3.15, 12: 3.61, 13: 3.86, 14: 3.17, 15: 7.1, 16: 3.87, 17: 4.8, 18: 3.34, 19: 2.84, 20: 4.99, 21: 1.91, 22: 3.46, 23: 2.15, 24: 4.37, 25: 1.13, 26: 4.35, 27: 7.74}\n</code></pre>",
          "rawMarkdown": "@jaideepvalani As explained by @maw501 (Mark), this code works in python 2. If you are using python 3, change this line:\n\n`total = np.sum(labels_dict.values())`\n\nto:\n\n`total = 50782`\n\nIt should work after that. The output looks like this:\n\n    True class weights:\n    {0: 3.94, 1: 40.5, 2: 14.02, 3: 32.53, 4: 27.33, 5: 20.21, 6: 50.38, 7: 18.0, 8: 958.15, 9: 1128.49, 10: 1813.64, 11: 46.46, 12: 73.81, 13: 94.57, 14: 47.64, 15: 2418.19, 16: 95.82, 17: 241.82, 18: 56.3, 19: 34.27, 20: 295.24, 21: 13.45, 22: 63.32, 23: 17.13, 24: 157.71, 25: 6.17, 26: 154.82, 27: 4616.55}\n    \n    Log-dampened class weights:\n    {0: 1.0, 1: 3.01, 2: 1.95, 3: 2.79, 4: 2.61, 5: 2.31, 6: 3.23, 7: 2.2, 8: 6.17, 9: 6.34, 10: 6.81, 11: 3.15, 12: 3.61, 13: 3.86, 14: 3.17, 15: 7.1, 16: 3.87, 17: 4.8, 18: 3.34, 19: 2.84, 20: 4.99, 21: 1.91, 22: 3.46, 23: 2.15, 24: 4.37, 25: 1.13, 26: 4.35, 27: 7.74}\n    \n",
          "votes": 2
        },
        {
          "id": 442559,
          "postDate": "2018-12-20T05:58:05.120Z",
          "content": "<p>Thanks.. i corrected this by wraping dict into a list ...\nBtw . the library i use for fit ,hasnt got class wegith param and also the loss function i use also hasnt cant the class weight .\nCan you please guide if there is a way i can introduce this class weight in the loss function...</p>",
          "rawMarkdown": "Thanks.. i corrected this by wraping dict into a list ...\nBtw . the library i use for fit ,hasnt got class wegith param and also the loss function i use also hasnt cant the class weight .\nCan you please guide if there is a way i can introduce this class weight in the loss function...\n"
        },
        {
          "id": 442571,
          "postDate": "2018-12-20T06:24:29.117Z",
          "content": "<p>here is my loss fct \ndef forward(self, input, target,reduction='none'):\n        if not (target.size() == input.size()):\n            raise ValueError(\"Target size ({}) must be the same as input size ({})\"\n                             .format(target.size(), input.size()))</p>\n\n<pre><code>    max_val = (-input).clamp(min=0)\n    loss = input - input * target + max_val + \\\n        ((-max_val).exp() + (-input - max_val).exp()).log()\n\n    invprobs = F.logsigmoid(-input * (target * 2.0 - 1.0))\n    loss = (invprobs * self.gamma).exp() * loss\n\n    return loss.sum(dim=1).mean()\n</code></pre>",
          "rawMarkdown": "here is my loss fct \ndef forward(self, input, target,reduction='none'):\n        if not (target.size() == input.size()):\n            raise ValueError(\"Target size ({}) must be the same as input size ({})\"\n                             .format(target.size(), input.size()))\n\n        max_val = (-input).clamp(min=0)\n        loss = input - input * target + max_val + \\\n            ((-max_val).exp() + (-input - max_val).exp()).log()\n\n        invprobs = F.logsigmoid(-input * (target * 2.0 - 1.0))\n        loss = (invprobs * self.gamma).exp() * loss\n        \n        return loss.sum(dim=1).mean()"
        },
        {
          "id": 530943,
          "postDate": "2019-05-14T02:27:32.770Z",
          "content": "<p>Hi, I'm not clear about what the <code>total = np.sum(labels_dict.values())</code> means. I found that the <code>total</code> doesn't equal to the dataset's size since the data is multilabel. So, why do we calculate the weight like this? I think, maybe we could use <code>total = len(dataset)</code> to get the weight?</p>",
          "rawMarkdown": "Hi, I'm not clear about what the `total = np.sum(labels_dict.values())` means. I found that the `total` doesn't equal to the dataset's size since the data is multilabel. So, why do we calculate the weight like this? I think, maybe we could use `total = len(dataset)` to get the weight?"
        }
      ]
    },
    {
      "id": 442099,
      "postDate": "2018-12-19T13:40:40.467Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 436063,
      "author_name": "Mark Worrall",
      "author_url": "",
      "post_date": "2018-12-09T13:45:41.757000",
      "content": "<p>Another (simple) approach is just to use a linear interpolation of the prior probability of a label being in a class towards equal weights. </p>\n\n<p>In the below, setting mu = 0 is the original probability distribution and as mu increases we (linearly) tend to the distribution with all class weights equal (the maximum entropy distribution).</p>\n\n<pre><code>import numpy as np\n\nname_label_dict = {\n    0:   ('Nucleoplasm', 12885),\n    1:   ('Nuclear membrane', 1254),\n    2:   ('Nucleoli', 3621),\n    3:   ('Nucleoli fibrillar center', 1561),\n    4:   ('Nuclear speckles', 1858),\n    5:   ('Nuclear bodies', 2513),\n    6:   ('Endoplasmic reticulum', 1008),   \n    7:   ('Golgi apparatus', 2822),\n    8:   ('Peroxisomes', 53), \n    9:   ('Endosomes', 45),\n    10:  ('Lysosomes', 28),\n    11:  ('Intermediate filaments', 1093), \n    12:  ('Actin filaments', 688),\n    13:  ('Focal adhesion sites', 537),  \n    14:  ('Microtubules', 1066), \n    15:  ('Microtubule ends', 21),\n    16:  ('Cytokinetic bridge', 530),\n    17:  ('Mitotic spindle', 210),\n    18:  ('Microtubule organizing center', 902),\n    19:  ('Centrosome', 1482),\n    20:  ('Lipid droplets', 172),\n    21:  ('Plasma membrane', 3777),\n    22:  ('Cell junctions', 802),\n    23:  ('Mitochondria', 2965),\n    24:  ('Aggresome', 322),\n    25:  ('Cytosol', 8228),\n    26:  ('Cytoplasmic bodies', 328),   \n    27:  ('Rods &amp;amp; rings', 11)\n    }\n\nn_labels = 50782\n\ndef cls_wts(label_dict, mu=0.5):\n    prob_dict, prob_dict_bal = {}, {}\n    max_ent_wt = 1/28\n    for i in range(28):\n        prob_dict[i] = label_dict[i][1]/n_labels\n        if prob_dict[i] &amp;gt; max_ent_wt:\n            prob_dict_bal[i] = prob_dict[i]-mu*(prob_dict[i] - max_ent_wt)\n        else:\n            prob_dict_bal[i] = prob_dict[i]+mu*(max_ent_wt - prob_dict[i])            \n    return prob_dict, prob_dict_bal\n</code></pre>",
      "votes": 6,
      "replies": [
        {
          "id": 436147,
          "author_name": "Tilii",
          "author_url": "",
          "post_date": "2018-12-09T17:21:30.107000",
          "content": "<blockquote>\n  <p>Another (simple) approach is just to use a linear interpolation of the prior probability of a label being in a class towards equal weights.</p>\n</blockquote>\n\n<p>My choice of scale was arbitrary, so linear dampening works fine. In fact, for most of us it may be easier to understand its effect than log adjustment.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 439001,
      "author_name": "William Horton",
      "author_url": "",
      "post_date": "2018-12-14T14:56:09.873000",
      "content": "<p>Wanted to report that my initial results are showing that the log-dampened weights do better than just the direct weights I was using before. Thank you for this!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 439145,
          "author_name": "Tilii",
          "author_url": "",
          "post_date": "2018-12-14T19:42:15.977000",
          "content": "<p><a href=\"/hortonhearsafoo\">@hortonhearsafoo</a> Thank you for reporting it. I have dealt with enough imbalanced datasets to know that applying class weights based on true class proportions is almost never the right thing to do. It becomes a matter of finding a ratio that corrects the imbalance. I am sure that linear interpolation suggested by <a href=\"/maw501\">@maw501</a> works as well, though most likely with a different value of mu.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 440274,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2018-12-17T10:14:08.177000",
      "content": "<p>hi ,thanks for input...\ncan you please guide where to fit this one... imeant where to pass this one ,the weights  to the loss</p>",
      "votes": 1,
      "replies": [
        {
          "id": 440664,
          "author_name": "Tilii",
          "author_url": "",
          "post_date": "2018-12-17T21:34:05.453000",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> Use one of the script's outputs to define class weights:</p>\n\n<p><code>cw = {0: 1.0, 1: 3.01, 2: 1.95, 3: 2.79, 4: 2.61, ....  26: 4.35, 27: 7.74}</code></p>\n\n<p>This dictionary is used as an argument in keras <code>.fit</code> function:</p>\n\n<p><code>class_weight=cw</code></p>\n\n<p>Other classifiers use <code>class_weight</code> in a similar way.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 447824,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2018-12-30T16:48:57.790000",
          "content": "<p>hi,i tried using it with BCE cross entropy ,but val loss is very very fluctuating ,it reaches as high as 200,low less than 10 ,every alternate epochs..\nare the log scale values are not in scale of loss ??\ni get values similar as above.. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 437035,
      "author_name": "Panchajanya Banerjee (Pancham)",
      "author_url": "",
      "post_date": "2018-12-11T09:25:25.953000",
      "content": "<p>This is very nicely done, and easy to understand. Thank you! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 539014,
      "author_name": "ZY. Feng",
      "author_url": "",
      "post_date": "2019-05-29T12:00:54.010000",
      "content": "<p>Excuse me, should the parameter mu  be bigger than 1？Or mu must be in the range of (0,1)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 531070,
      "author_name": "QunYang",
      "author_url": "",
      "post_date": "2019-05-14T08:34:04.160000",
      "content": "<p>Excuse me, is there any general rule of thumb when choosing the parameter <code>mu</code>？</p>",
      "votes": 0,
      "replies": [
        {
          "id": 531344,
          "author_name": "Tilii",
          "author_url": "",
          "post_date": "2019-05-14T17:29:18.637000",
          "content": "<blockquote>\n  <p>Excuse me, is there any general rule of thumb when choosing the parameter mu</p>\n</blockquote>\n\n<p><a href=\"/qunyang\">@qunyang</a> It is a parameter that needs to be tuned for each dataset. Larger values of <code>mu</code> mean the weights are more similar to true class ratios. Small values of <code>mu</code> dampen the difference.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 531471,
          "author_name": "QunYang",
          "author_url": "",
          "post_date": "2019-05-15T01:26:04.227000",
          "content": "<p>Thanks, I understand</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 442278,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2018-12-19T18:14:59.427000",
      "content": "<p>hi thanks for explaination</p>\n\n<p>i get this error\nunsupported operand type(s) for /: 'dict_values' and 'float'\nwhile calculating score\nwhen i print total variable i get\ndict_values([12885, 1254, 3621,......\nprobably because of that it fails</p>\n\n<p>should this be a single value ??  sum of all the counts </p>",
      "votes": 0,
      "replies": [
        {
          "id": 442285,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-19T18:25:01.767000",
          "content": "<p>I'm not clear what error you are having but I think I recall in the code Tilii provided needing to wrap the <code>labels_dict.values()</code> in a <code>list</code> as I think it's based on python 2 - Tilii can correct me if I'm wrong.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 442303,
          "author_name": "Tilii",
          "author_url": "",
          "post_date": "2018-12-19T19:03:37.313000",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> As explained by <a href=\"/maw501\">@maw501</a> (Mark), this code works in python 2. If you are using python 3, change this line:</p>\n\n<p><code>total = np.sum(labels_dict.values())</code></p>\n\n<p>to:</p>\n\n<p><code>total = 50782</code></p>\n\n<p>It should work after that. The output looks like this:</p>\n\n<pre><code>True class weights:\n{0: 3.94, 1: 40.5, 2: 14.02, 3: 32.53, 4: 27.33, 5: 20.21, 6: 50.38, 7: 18.0, 8: 958.15, 9: 1128.49, 10: 1813.64, 11: 46.46, 12: 73.81, 13: 94.57, 14: 47.64, 15: 2418.19, 16: 95.82, 17: 241.82, 18: 56.3, 19: 34.27, 20: 295.24, 21: 13.45, 22: 63.32, 23: 17.13, 24: 157.71, 25: 6.17, 26: 154.82, 27: 4616.55}\n\nLog-dampened class weights:\n{0: 1.0, 1: 3.01, 2: 1.95, 3: 2.79, 4: 2.61, 5: 2.31, 6: 3.23, 7: 2.2, 8: 6.17, 9: 6.34, 10: 6.81, 11: 3.15, 12: 3.61, 13: 3.86, 14: 3.17, 15: 7.1, 16: 3.87, 17: 4.8, 18: 3.34, 19: 2.84, 20: 4.99, 21: 1.91, 22: 3.46, 23: 2.15, 24: 4.37, 25: 1.13, 26: 4.35, 27: 7.74}\n</code></pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 442559,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2018-12-20T05:58:05.120000",
          "content": "<p>Thanks.. i corrected this by wraping dict into a list ...\nBtw . the library i use for fit ,hasnt got class wegith param and also the loss function i use also hasnt cant the class weight .\nCan you please guide if there is a way i can introduce this class weight in the loss function...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442571,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2018-12-20T06:24:29.117000",
          "content": "<p>here is my loss fct \ndef forward(self, input, target,reduction='none'):\n        if not (target.size() == input.size()):\n            raise ValueError(\"Target size ({}) must be the same as input size ({})\"\n                             .format(target.size(), input.size()))</p>\n\n<pre><code>    max_val = (-input).clamp(min=0)\n    loss = input - input * target + max_val + \\\n        ((-max_val).exp() + (-input - max_val).exp()).log()\n\n    invprobs = F.logsigmoid(-input * (target * 2.0 - 1.0))\n    loss = (invprobs * self.gamma).exp() * loss\n\n    return loss.sum(dim=1).mean()\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 530943,
          "author_name": "QunYang",
          "author_url": "",
          "post_date": "2019-05-14T02:27:32.770000",
          "content": "<p>Hi, I'm not clear about what the <code>total = np.sum(labels_dict.values())</code> means. I found that the <code>total</code> doesn't equal to the dataset's size since the data is multilabel. So, why do we calculate the weight like this? I think, maybe we could use <code>total = len(dataset)</code> to get the weight?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 442099,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-19T13:40:40.467000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "435480": "<pre>    #mu in \"create_class_weight\" is a dampening parameter that could be tuned\n\n    import numpy as np\n    import math\n\n    def create_class_weight(labels_dict, mu=0.5):\n        total = np.sum(labels_dict.values())\n        keys = labels_dict.keys()\n        class_weight = dict()\n        class_weight_log = dict()\n\n        for key in keys:\n            score = total / float(labels_dict[key])\n            score_log = math.log(mu * total / float(labels_dict[key]))\n            class_weight[key] = round(score, 2) if score &gt; 1.0 else round(1.0, 2)\n            class_weight_log[key] = round(score_log, 2) if score_log &gt; 1.0 else round(1.0, 2)\n\n        return class_weight, class_weight_log\n\n    # Class abundance for protein dataset\n    labels_dict = {\n        0: 12885,\n        1: 1254,\n        2: 3621,\n        3: 1561,\n        4: 1858,\n        5: 2513,\n        6: 1008,\n        7: 2822,\n        8: 53,\n        9: 45,\n        10: 28,\n        11: 1093,\n        12: 688,\n        13: 537,\n        14: 1066,\n        15: 21,\n        16: 530,\n        17: 210,\n        18: 902,\n        19: 1482,\n        20: 172,\n        21: 3777,\n        22: 802,\n        23: 2965,\n        24: 322,\n        25: 8228,\n        26: 328,\n        27: 11\n    }\n\n    print('\\nTrue class weights:')\n    print(create_class_weight(labels_dict)[0])\n    print('\\nLog-dampened class weights:')\n    print(create_class_weight(labels_dict)[1])<code></code></pre>",
    "436063": "Another (simple) approach is just to use a linear interpolation of the prior probability of a label being in a class towards equal weights. \n\nIn the below, setting mu = 0 is the original probability distribution and as mu increases we (linearly) tend to the distribution with all class weights equal (the maximum entropy distribution).\n\n\n    import numpy as np\n\n    name_label_dict = {\n        0:   ('Nucleoplasm', 12885),\n        1:   ('Nuclear membrane', 1254),\n        2:   ('Nucleoli', 3621),\n        3:   ('Nucleoli fibrillar center', 1561),\n        4:   ('Nuclear speckles', 1858),\n        5:   ('Nuclear bodies', 2513),\n        6:   ('Endoplasmic reticulum', 1008),   \n        7:   ('Golgi apparatus', 2822),\n        8:   ('Peroxisomes', 53), \n        9:   ('Endosomes', 45),\n        10:  ('Lysosomes', 28),\n        11:  ('Intermediate filaments', 1093), \n        12:  ('Actin filaments', 688),\n        13:  ('Focal adhesion sites', 537),  \n        14:  ('Microtubules', 1066), \n        15:  ('Microtubule ends', 21),\n        16:  ('Cytokinetic bridge', 530),\n        17:  ('Mitotic spindle', 210),\n        18:  ('Microtubule organizing center', 902),\n        19:  ('Centrosome', 1482),\n        20:  ('Lipid droplets', 172),\n        21:  ('Plasma membrane', 3777),\n        22:  ('Cell junctions', 802),\n        23:  ('Mitochondria', 2965),\n        24:  ('Aggresome', 322),\n        25:  ('Cytosol', 8228),\n        26:  ('Cytoplasmic bodies', 328),   \n        27:  ('Rods &amp; rings', 11)\n        }\n\n    n_labels = 50782\n\n    def cls_wts(label_dict, mu=0.5):\n        prob_dict, prob_dict_bal = {}, {}\n        max_ent_wt = 1/28\n        for i in range(28):\n            prob_dict[i] = label_dict[i][1]/n_labels\n            if prob_dict[i] &gt; max_ent_wt:\n                prob_dict_bal[i] = prob_dict[i]-mu*(prob_dict[i] - max_ent_wt)\n            else:\n                prob_dict_bal[i] = prob_dict[i]+mu*(max_ent_wt - prob_dict[i])            \n        return prob_dict, prob_dict_bal\n\n\n",
    "439001": "Wanted to report that my initial results are showing that the log-dampened weights do better than just the direct weights I was using before. Thank you for this!",
    "440274": "hi ,thanks for input...\ncan you please guide where to fit this one... imeant where to pass this one ,the weights  to the loss",
    "437035": "This is very nicely done, and easy to understand. Thank you! ",
    "539014": "Excuse me, should the parameter mu  be bigger than 1？Or mu must be in the range of (0,1)",
    "531070": "Excuse me, is there any general rule of thumb when choosing the parameter `mu`？",
    "442278": "hi thanks for explaination\n\ni get this error\nunsupported operand type(s) for /: 'dict_values' and 'float'\nwhile calculating score\nwhen i print total variable i get\ndict_values([12885, 1254, 3621,......\nprobably because of that it fails\n\nshould this be a single value ??  sum of all the counts \n",
    "442099": ""
  }
}