{
  "id": 344578,
  "title": "Custom LGBM Obj Version 2: Weighted LogLoss Function",
  "url": "/competitions/amex-default-prediction/discussion/344578",
  "author_name": "",
  "post_date": "2022-08-15T18:33:51.859667800Z",
  "votes": 11,
  "comment_count": 4,
  "views": 0,
  "content": "<p>In this <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/344038\" target=\"_blank\">topic</a> and <a href=\"https://www.kaggle.com/code/jpison/custom-lgbm-obj-weighted-logloss-function\" target=\"_blank\">notebook</a> I proposed a tentative loss function based on ranking <strong>to penalize both the Gradient and the Hessian of the positive cases that are further away from the top positions</strong>. </p>\n<p>However, as <a href=\"https://www.kaggle.com/davidirudel\" target=\"_blank\">@davidirudel</a> commented <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/344038#1898761\" target=\"_blank\">here</a> <strong>I was not correctly multiplying the weights to the hessian</strong>. </p>\n<p>If the loss for one case is L and p(x) the sigmoid function. I want to multiply the weight alpha &gt;1.0 which is proportional to the rank position for each true positive. </p>\n<p>Then:</p>\n<p>$$L = -\\alpha y \\ln p - (1-y) \\ln(1-p)$$</p>\n<p>$$\\frac{\\partial p}{\\partial x}  = p(1-p)$$</p>\n<p>$$\\text{grad} = \\frac{\\partial L}{\\partial x} = \\frac{\\partial L}{\\partial p} \\frac{\\partial p}{\\partial x} = p(1 + \\alpha y - y) - \\alpha y$$</p>\n<p>$$\\text{hess} = \\frac{\\partial^2 L}{\\partial x^2} = (1 + \\alpha y - y)p(1-p) $$</p>\n<p>I think, the tentative version 2 of the <em>Weighted LogLoss Function</em> based on rank could be:</p>\n<pre><code>def weighted_logloss(preds, dtrain):\n   global MULT_NO4PERC, MAX_WEIGHTS\n   eps = 1e-16\n   labels = dtrain.get_label()\n   preds = 1.0 / (1.0 + np.exp(-preds))\n\n   # top 4 perc\n   labels_mat = np.transpose(np.array([np.arange(len(labels)), labels, preds]))\n   pos_ord = labels_mat[:, 2].argsort()[::-1]\n   labels_mat = labels_mat[pos_ord]\n   weights_4perc = np.where(labels_mat[:,1]==0, 20, 1)\n   top4 = np.cumsum(weights_4perc) &lt;= int(0.04 * np.sum(weights_4perc))\n   top4 = top4[labels_mat[:, 0].argsort()]\n\n   weights = 1+np.exp(-MULT_NO4PERC*np.linspace(MAX_WEIGHTS-1,0,len(top4)))[labels_mat[:, 0].argsort()]\n   # Set to one weights of positive labels in top 4perc\n   weights[top4 &amp; (labels==1.0)] = 1.0 \n   # Set to one weights of negative labels\n   weights[(labels==0.0)] = 1.0 \n\n  grad = preds * (1 + weights*labels - labels) - (weights * labels)\n  hess = np.maximum(preds * (1.0 - preds) * (1 + weights*labels - labels) , eps)\n\n   return grad, hess\n</code></pre>",
  "messages": [
    {
      "id": "1900149",
      "postDate": "08/15/2022 18:33:51",
      "content": "<p>In this <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/344038\" target=\"_blank\">topic</a> and <a href=\"https://www.kaggle.com/code/jpison/custom-lgbm-obj-weighted-logloss-function\" target=\"_blank\">notebook</a> I proposed a tentative loss function based on ranking <strong>to penalize both the Gradient and the Hessian of the positive cases that are further away from the top positions</strong>. </p>\n<p>However, as <a href=\"https://www.kaggle.com/davidirudel\" target=\"_blank\">@davidirudel</a> commented <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/344038#1898761\" target=\"_blank\">here</a> <strong>I was not correctly multiplying the weights to the hessian</strong>. </p>\n<p>If the loss for one case is L and p(x) the sigmoid function. I want to multiply the weight alpha &gt;1.0 which is proportional to the rank position for each true positive. </p>\n<p>Then:</p>\n<p>$$L = -\\alpha y \\ln p - (1-y) \\ln(1-p)$$</p>\n<p>$$\\frac{\\partial p}{\\partial x}  = p(1-p)$$</p>\n<p>$$\\text{grad} = \\frac{\\partial L}{\\partial x} = \\frac{\\partial L}{\\partial p} \\frac{\\partial p}{\\partial x} = p(1 + \\alpha y - y) - \\alpha y$$</p>\n<p>$$\\text{hess} = \\frac{\\partial^2 L}{\\partial x^2} = (1 + \\alpha y - y)p(1-p) $$</p>\n<p>I think, the tentative version 2 of the <em>Weighted LogLoss Function</em> based on rank could be:</p>\n<pre><code>def weighted_logloss(preds, dtrain):\n   global MULT_NO4PERC, MAX_WEIGHTS\n   eps = 1e-16\n   labels = dtrain.get_label()\n   preds = 1.0 / (1.0 + np.exp(-preds))\n\n   # top 4 perc\n   labels_mat = np.transpose(np.array([np.arange(len(labels)), labels, preds]))\n   pos_ord = labels_mat[:, 2].argsort()[::-1]\n   labels_mat = labels_mat[pos_ord]\n   weights_4perc = np.where(labels_mat[:,1]==0, 20, 1)\n   top4 = np.cumsum(weights_4perc) &lt;= int(0.04 * np.sum(weights_4perc))\n   top4 = top4[labels_mat[:, 0].argsort()]\n\n   weights = 1+np.exp(-MULT_NO4PERC*np.linspace(MAX_WEIGHTS-1,0,len(top4)))[labels_mat[:, 0].argsort()]\n   # Set to one weights of positive labels in top 4perc\n   weights[top4 &amp; (labels==1.0)] = 1.0 \n   # Set to one weights of negative labels\n   weights[(labels==0.0)] = 1.0 \n\n  grad = preds * (1 + weights*labels - labels) - (weights * labels)\n  hess = np.maximum(preds * (1.0 - preds) * (1 + weights*labels - labels) , eps)\n\n   return grad, hess\n</code></pre>",
      "rawMarkdown": "In this [topic](https://www.kaggle.com/competitions/amex-default-prediction/discussion/344038) and [notebook](https://www.kaggle.com/code/jpison/custom-lgbm-obj-weighted-logloss-function) I proposed a tentative loss function based on ranking **to penalize both the Gradient and the Hessian of the positive cases that are further away from the top positions**. \n\nHowever, as @davidirudel commented [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/344038#1898761) **I was not correctly multiplying the weights to the hessian**. \n\nIf the loss for one case is L and p(x) the sigmoid function. I want to multiply the weight alpha >1.0 which is proportional to the rank position for each true positive. \n\nThen:\n\n$$L = -\\alpha y \\ln p - (1-y) \\ln(1-p)$$\n\n$$\\frac{\\partial p}{\\partial x}  = p(1-p)$$\n\n$$\\text{grad} = \\frac{\\partial L}{\\partial x} = \\frac{\\partial L}{\\partial p} \\frac{\\partial p}{\\partial x} = p(1 + \\alpha y - y) - \\alpha y$$\n\n$$\\text{hess} = \\frac{\\partial^2 L}{\\partial x^2} = (1 + \\alpha y - y)p(1-p) $$\n\nI think, the tentative version 2 of the *Weighted LogLoss Function* based on rank could be:\n\n    def weighted_logloss(preds, dtrain):\n       global MULT_NO4PERC, MAX_WEIGHTS\n       eps = 1e-16\n       labels = dtrain.get_label()\n       preds = 1.0 / (1.0 + np.exp(-preds))\n    \n       # top 4 perc\n       labels_mat = np.transpose(np.array([np.arange(len(labels)), labels, preds]))\n       pos_ord = labels_mat[:, 2].argsort()[::-1]\n       labels_mat = labels_mat[pos_ord]\n       weights_4perc = np.where(labels_mat[:,1]==0, 20, 1)\n       top4 = np.cumsum(weights_4perc) <= int(0.04 * np.sum(weights_4perc))\n       top4 = top4[labels_mat[:, 0].argsort()]\n\n       weights = 1+np.exp(-MULT_NO4PERC*np.linspace(MAX_WEIGHTS-1,0,len(top4)))[labels_mat[:, 0].argsort()]\n       # Set to one weights of positive labels in top 4perc\n       weights[top4 & (labels==1.0)] = 1.0 \n       # Set to one weights of negative labels\n       weights[(labels==0.0)] = 1.0 \n\n      grad = preds * (1 + weights*labels - labels) - (weights * labels)\n      hess = np.maximum(preds * (1.0 - preds) * (1 + weights*labels - labels) , eps)\n\n       return grad, hess",
      "votes": null
    },
    {
      "id": "1901356",
      "postDate": "08/16/2022 15:44:35",
      "content": "<p>I think hessian is wrong. <br>\nCan you please confirm?</p>",
      "rawMarkdown": "I think hessian is wrong. \nCan you please confirm?",
      "votes": null
    },
    {
      "id": "1901808",
      "postDate": "08/16/2022 23:54:45",
      "content": "<p>$$\\text{grad} = p(1 + \\alpha y - y) - \\alpha y$$</p>\n<p>$$\\text{hess} = \\frac{\\partial grad}{\\partial x} = \\frac{\\partial grad}{\\partial p} \\frac{\\partial p}{\\partial x} = (1 + \\alpha y - y) p (1 - p)  $$</p>",
      "rawMarkdown": "$$\\text{grad} = p(1 + \\alpha y - y) - \\alpha y$$\n\n$$\\text{hess} = \\frac{\\partial grad}{\\partial x} = \\frac{\\partial grad}{\\partial p} \\frac{\\partial p}{\\partial x} = (1 + \\alpha y - y) p (1 - p)  $$",
      "votes": null
    },
    {
      "id": "1902986",
      "postDate": "08/17/2022 03:28:03",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "1903256",
      "postDate": "08/17/2022 08:32:26",
      "content": "<p>Why do you use <code>maximum()</code> operation for hessian in the code?</p>",
      "rawMarkdown": "Why do you use `maximum()` operation for hessian in the code?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1901356,
      "author_name": "leewook",
      "author_url": "",
      "post_date": "08/16/2022 15:44:35",
      "content": "<p>I think hessian is wrong. <br>\nCan you please confirm?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1901808,
          "author_name": "jpison",
          "author_url": "",
          "post_date": "08/16/2022 23:54:45",
          "content": "<p>$$\\text{grad} = p(1 + \\alpha y - y) - \\alpha y$$</p>\n<p>$$\\text{hess} = \\frac{\\partial grad}{\\partial x} = \\frac{\\partial grad}{\\partial p} \\frac{\\partial p}{\\partial x} = (1 + \\alpha y - y) p (1 - p)  $$</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1902986,
          "author_name": "leewook",
          "author_url": "",
          "post_date": "08/17/2022 03:28:03",
          "content": "<p>Thanks for sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1903256,
      "author_name": "rasoulmojtahedzadeh",
      "author_url": "",
      "post_date": "08/17/2022 08:32:26",
      "content": "<p>Why do you use <code>maximum()</code> operation for hessian in the code?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1900149": "In this [topic](https://www.kaggle.com/competitions/amex-default-prediction/discussion/344038) and [notebook](https://www.kaggle.com/code/jpison/custom-lgbm-obj-weighted-logloss-function) I proposed a tentative loss function based on ranking **to penalize both the Gradient and the Hessian of the positive cases that are further away from the top positions**. \n\nHowever, as @davidirudel commented [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/344038#1898761) **I was not correctly multiplying the weights to the hessian**. \n\nIf the loss for one case is L and p(x) the sigmoid function. I want to multiply the weight alpha >1.0 which is proportional to the rank position for each true positive. \n\nThen:\n\n$$L = -\\alpha y \\ln p - (1-y) \\ln(1-p)$$\n\n$$\\frac{\\partial p}{\\partial x}  = p(1-p)$$\n\n$$\\text{grad} = \\frac{\\partial L}{\\partial x} = \\frac{\\partial L}{\\partial p} \\frac{\\partial p}{\\partial x} = p(1 + \\alpha y - y) - \\alpha y$$\n\n$$\\text{hess} = \\frac{\\partial^2 L}{\\partial x^2} = (1 + \\alpha y - y)p(1-p) $$\n\nI think, the tentative version 2 of the *Weighted LogLoss Function* based on rank could be:\n\n    def weighted_logloss(preds, dtrain):\n       global MULT_NO4PERC, MAX_WEIGHTS\n       eps = 1e-16\n       labels = dtrain.get_label()\n       preds = 1.0 / (1.0 + np.exp(-preds))\n    \n       # top 4 perc\n       labels_mat = np.transpose(np.array([np.arange(len(labels)), labels, preds]))\n       pos_ord = labels_mat[:, 2].argsort()[::-1]\n       labels_mat = labels_mat[pos_ord]\n       weights_4perc = np.where(labels_mat[:,1]==0, 20, 1)\n       top4 = np.cumsum(weights_4perc) <= int(0.04 * np.sum(weights_4perc))\n       top4 = top4[labels_mat[:, 0].argsort()]\n\n       weights = 1+np.exp(-MULT_NO4PERC*np.linspace(MAX_WEIGHTS-1,0,len(top4)))[labels_mat[:, 0].argsort()]\n       # Set to one weights of positive labels in top 4perc\n       weights[top4 & (labels==1.0)] = 1.0 \n       # Set to one weights of negative labels\n       weights[(labels==0.0)] = 1.0 \n\n      grad = preds * (1 + weights*labels - labels) - (weights * labels)\n      hess = np.maximum(preds * (1.0 - preds) * (1 + weights*labels - labels) , eps)\n\n       return grad, hess",
    "1901356": "I think hessian is wrong. \nCan you please confirm?",
    "1901808": "$$\\text{grad} = p(1 + \\alpha y - y) - \\alpha y$$\n\n$$\\text{hess} = \\frac{\\partial grad}{\\partial x} = \\frac{\\partial grad}{\\partial p} \\frac{\\partial p}{\\partial x} = (1 + \\alpha y - y) p (1 - p)  $$",
    "1902986": "Thanks for sharing!",
    "1903256": "Why do you use `maximum()` operation for hessian in the code?"
  },
  "source": "meta"
}