{
  "id": 535691,
  "title": "custom objective",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/535691",
  "author_name": "",
  "post_date": "2024-09-23T16:43:41.177215800Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi kagglers! In many notebooks here, I saw that If we want to make custom objective for optimizing kappa we can use something like this. Can somebody explain what are \"a\" and \"b\" constants and how they was calculated? Is it something to scale the target distribution or what ? </p>\n<p>a = 2.998<br>\nb = 1.092<br>\ndef MakeObj(y_true, y_pred):<br>\n    labels = y_true + a<br>\n    preds  = y_pred + a<br>\n    preds  = preds.clip(0, np.inf)<br>\n    f      = 1/2<em>np.sum((preds-labels)<strong>2)\n    g      = 1/2</strong></em>np.sum((preds-a)2+b)<br>\n    df     = preds - labels<br>\n    dg     = preds - a<br>\n    grad   = (df/g - f<em>dg/g</em><em>2)</em>len(labels)<br>\n    hess   = np.ones(len(labels))<br>\n    return grad, hess</p>\n<p>P.S. for this example thanks to <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>",
  "messages": [
    {
      "id": "2996578",
      "postDate": "09/23/2024 16:43:41",
      "content": "<p>Hi kagglers! In many notebooks here, I saw that If we want to make custom objective for optimizing kappa we can use something like this. Can somebody explain what are \"a\" and \"b\" constants and how they was calculated? Is it something to scale the target distribution or what ? </p>\n<p>a = 2.998<br>\nb = 1.092<br>\ndef MakeObj(y_true, y_pred):<br>\n    labels = y_true + a<br>\n    preds  = y_pred + a<br>\n    preds  = preds.clip(0, np.inf)<br>\n    f      = 1/2<em>np.sum((preds-labels)<strong>2)\n    g      = 1/2</strong></em>np.sum((preds-a)2+b)<br>\n    df     = preds - labels<br>\n    dg     = preds - a<br>\n    grad   = (df/g - f<em>dg/g</em><em>2)</em>len(labels)<br>\n    hess   = np.ones(len(labels))<br>\n    return grad, hess</p>\n<p>P.S. for this example thanks to <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>",
      "rawMarkdown": "Hi kagglers! In many notebooks here, I saw that If we want to make custom objective for optimizing kappa we can use something like this. Can somebody explain what are \"a\" and \"b\" constants and how they was calculated? Is it something to scale the target distribution or what ? \n\na = 2.998\nb = 1.092\ndef MakeObj(y_true, y_pred):\n    labels = y_true + a\n    preds  = y_pred + a\n    preds  = preds.clip(0, np.inf)\n    f      = 1/2*np.sum((preds-labels)**2)\n    g      = 1/2*np.sum((preds-a)**2+b)\n    df     = preds - labels\n    dg     = preds - a\n    grad   = (df/g - f*dg/g**2)*len(labels)\n    hess   = np.ones(len(labels))\n    return grad, hess\n\nP.S. for this example thanks to @ravi20076",
      "votes": null
    },
    {
      "id": "2996703",
      "postDate": "09/23/2024 20:38:21",
      "content": "<p><a href=\"https://www.kaggle.com/alexshumilov\" target=\"_blank\">@alexshumilov</a> I simply took them from another competition artefacts. They were used there to good effect </p>",
      "rawMarkdown": "alexshumilov I simply took them from another competition artefacts. They were used there to good effect",
      "votes": null
    },
    {
      "id": "2996758",
      "postDate": "09/23/2024 21:45:36",
      "content": "<p>According to comment <a href=\"https://www.kaggle.com/competitions/learning-agency-lab-automated-essay-scoring-2/discussion/511130#2872396\" target=\"_blank\">here</a> in Essay Scoring competition, <code>a</code> is mean of train target and <code>b</code> is std of train target. And the following blog <a href=\"https://zenn.dev/jackthekaggler/articles/cf988ca341e34ed83034\" target=\"_blank\">here</a> explains why</p>",
      "rawMarkdown": "According to comment [here][1] in Essay Scoring competition, `a` is mean of train target and `b` is std of train target. And the following blog [here][2] explains why\n\n[1]: https://www.kaggle.com/competitions/learning-agency-lab-automated-essay-scoring-2/discussion/511130#2872396\n[2]: https://zenn.dev/jackthekaggler/articles/cf988ca341e34ed83034",
      "votes": null
    },
    {
      "id": "2998762",
      "postDate": "09/26/2024 01:32:45",
      "content": "<p>The derivation is a straightforward calculation. In standard notations,<br>\n\\[<br>\nO_{ij}=\\sum_k\\delta_{iy_k}\\delta_{j\\hat{y_k}}<br>\n\\]<br>\n\\[<br>\nE_{ij}=p_i\\sum_k\\delta_{j\\hat{y_k}}<br>\n\\]<br>\nwhere \\(p_i\\) is the proportion of label \\(i\\) in the training set so that \\(\\sum_ip_i=1\\). Then<br>\n\\[<br>\n\\mathfrak{f}=\\sum_{i,j}(i-j)^2O_{ij}=\\sum_k\\sum_{i,j}(i-j)^2\\delta_{iy_k}\\delta_{j\\hat{y_k}}=\\sum_k(y_k-\\hat{y_k})^2<br>\n\\]<br>\n\\[<br>\n\\mathfrak{g}=\\sum_{i,j}(i-j)^2E_{ij}=\\sum_k\\sum_ip_i(i-\\hat{y_k})^2=\\sum_k\\mathbb{E}_y[(\\hat{y_k}-y)^2]=\\sum_k\\left((\\hat{y_k}-a)^2+b\\right)<br>\n\\]</p>\n<p>where \\(a\\) and \\(b\\) are respectively the mean and variance of \\(y_{train}\\). This follows from the identity \\(\\mathbb{E}_x[(t-x)^2]=(t-\\mathbb{E}[x])^2+\\mbox{Var}[x]\\). Then \\(1-\\kappa=\\mathfrak{f}/\\mathfrak{g}\\) which is a rational function of \\(\\hat{y_k}\\). Interestingly, it extends QWK to real-valued \\(\\hat{y_k}\\), and we can start taking the gradient and hessian with respect to \\(\\hat{y_k}\\).</p>\n<p>The gradient \\(g\\) should just be \\(\\frac{\\mathfrak{f}'}{\\mathfrak{g}}-\\frac{\\mathfrak{f}\\mathfrak{g}'}{\\mathfrak{g}^2}\\). I am not sure why the scaling factor <code>len(labels)</code> is needed; it would affect the learning rate. The hessian \\(h\\) can also be analytically computed, but there is no reason to believe that it would be bounded below by some positive number. In practice, we only need \\(const. + g\\cdot y_{pred}+\\frac{1}{2}h\\cdot y_{pred}^2\\) to be a majorant of the objective function (see the <a href=\"https://github.com/dmlc/xgboost/issues/1825\" target=\"_blank\">explanation</a> from the inventor of xgboost). Perhaps scaling the gradient by <code>len(labels)</code> and setting the hessian to \\(1\\) can make this happen.</p>",
      "rawMarkdown": "The derivation is a straightforward calculation. In standard notations,\n\\\\[\nO_{ij}=\\sum_k\\delta_{iy_k}\\delta_{j\\hat{y_k}}\n\\\\]\n\\\\[\nE_{ij}=p_i\\sum_k\\delta_{j\\hat{y_k}}\n\\\\]\nwhere \\\\(p_i\\\\) is the proportion of label \\\\(i\\\\) in the training set so that \\\\(\\sum_ip_i=1\\\\). Then\n\\\\[\n\\mathfrak{f}=\\sum_{i,j}(i-j)^2O_{ij}=\\sum_k\\sum_{i,j}(i-j)^2\\delta_{iy_k}\\delta_{j\\hat{y_k}}=\\sum_k(y_k-\\hat{y_k})^2\n\\\\]\n\\\\[\n\\mathfrak{g}=\\sum_{i,j}(i-j)^2E_{ij}=\\sum_k\\sum_ip_i(i-\\hat{y_k})^2=\\sum_k\\mathbb{E}_y[(\\hat{y_k}-y)^2]=\\sum_k\\left((\\hat{y_k}-a)^2+b\\right)\n\\\\]\n\nwhere \\\\(a\\\\) and \\\\(b\\\\) are respectively the mean and variance of \\\\(y_{train}\\\\). This follows from the identity \\\\(\\mathbb{E}_x[(t-x)^2]=(t-\\mathbb{E}[x])^2+\\mbox{Var}[x]\\\\). Then \\\\(1-\\kappa=\\mathfrak{f}/\\mathfrak{g}\\\\) which is a rational function of \\\\(\\hat{y_k}\\\\). Interestingly, it extends QWK to real-valued \\\\(\\hat{y_k}\\\\), and we can start taking the gradient and hessian with respect to \\\\(\\hat{y_k}\\\\).\n\nThe gradient \\\\(g\\\\) should just be \\\\(\\frac{\\mathfrak{f}'}{\\mathfrak{g}}-\\frac{\\mathfrak{f}\\mathfrak{g}'}{\\mathfrak{g}^2}\\\\). I am not sure why the scaling factor `len(labels)` is needed; it would affect the learning rate. The hessian \\\\(h\\\\) can also be analytically computed, but there is no reason to believe that it would be bounded below by some positive number. In practice, we only need \\\\(const. + g\\cdot y_{pred}+\\frac{1}{2}h\\cdot y_{pred}^2\\\\) to be a majorant of the objective function (see the [explanation](https://github.com/dmlc/xgboost/issues/1825) from the inventor of xgboost). Perhaps scaling the gradient by `len(labels)` and setting the hessian to \\\\(1\\\\) can make this happen.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2996703,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "09/23/2024 20:38:21",
      "content": "<p><a href=\"https://www.kaggle.com/alexshumilov\" target=\"_blank\">@alexshumilov</a> I simply took them from another competition artefacts. They were used there to good effect </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2996758,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "09/23/2024 21:45:36",
      "content": "<p>According to comment <a href=\"https://www.kaggle.com/competitions/learning-agency-lab-automated-essay-scoring-2/discussion/511130#2872396\" target=\"_blank\">here</a> in Essay Scoring competition, <code>a</code> is mean of train target and <code>b</code> is std of train target. And the following blog <a href=\"https://zenn.dev/jackthekaggler/articles/cf988ca341e34ed83034\" target=\"_blank\">here</a> explains why</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2998762,
      "author_name": "siukeitin",
      "author_url": "",
      "post_date": "09/26/2024 01:32:45",
      "content": "<p>The derivation is a straightforward calculation. In standard notations,<br>\n\\[<br>\nO_{ij}=\\sum_k\\delta_{iy_k}\\delta_{j\\hat{y_k}}<br>\n\\]<br>\n\\[<br>\nE_{ij}=p_i\\sum_k\\delta_{j\\hat{y_k}}<br>\n\\]<br>\nwhere \\(p_i\\) is the proportion of label \\(i\\) in the training set so that \\(\\sum_ip_i=1\\). Then<br>\n\\[<br>\n\\mathfrak{f}=\\sum_{i,j}(i-j)^2O_{ij}=\\sum_k\\sum_{i,j}(i-j)^2\\delta_{iy_k}\\delta_{j\\hat{y_k}}=\\sum_k(y_k-\\hat{y_k})^2<br>\n\\]<br>\n\\[<br>\n\\mathfrak{g}=\\sum_{i,j}(i-j)^2E_{ij}=\\sum_k\\sum_ip_i(i-\\hat{y_k})^2=\\sum_k\\mathbb{E}_y[(\\hat{y_k}-y)^2]=\\sum_k\\left((\\hat{y_k}-a)^2+b\\right)<br>\n\\]</p>\n<p>where \\(a\\) and \\(b\\) are respectively the mean and variance of \\(y_{train}\\). This follows from the identity \\(\\mathbb{E}_x[(t-x)^2]=(t-\\mathbb{E}[x])^2+\\mbox{Var}[x]\\). Then \\(1-\\kappa=\\mathfrak{f}/\\mathfrak{g}\\) which is a rational function of \\(\\hat{y_k}\\). Interestingly, it extends QWK to real-valued \\(\\hat{y_k}\\), and we can start taking the gradient and hessian with respect to \\(\\hat{y_k}\\).</p>\n<p>The gradient \\(g\\) should just be \\(\\frac{\\mathfrak{f}'}{\\mathfrak{g}}-\\frac{\\mathfrak{f}\\mathfrak{g}'}{\\mathfrak{g}^2}\\). I am not sure why the scaling factor <code>len(labels)</code> is needed; it would affect the learning rate. The hessian \\(h\\) can also be analytically computed, but there is no reason to believe that it would be bounded below by some positive number. In practice, we only need \\(const. + g\\cdot y_{pred}+\\frac{1}{2}h\\cdot y_{pred}^2\\) to be a majorant of the objective function (see the <a href=\"https://github.com/dmlc/xgboost/issues/1825\" target=\"_blank\">explanation</a> from the inventor of xgboost). Perhaps scaling the gradient by <code>len(labels)</code> and setting the hessian to \\(1\\) can make this happen.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2996578": "Hi kagglers! In many notebooks here, I saw that If we want to make custom objective for optimizing kappa we can use something like this. Can somebody explain what are \"a\" and \"b\" constants and how they was calculated? Is it something to scale the target distribution or what ? \n\na = 2.998\nb = 1.092\ndef MakeObj(y_true, y_pred):\n    labels = y_true + a\n    preds  = y_pred + a\n    preds  = preds.clip(0, np.inf)\n    f      = 1/2*np.sum((preds-labels)**2)\n    g      = 1/2*np.sum((preds-a)**2+b)\n    df     = preds - labels\n    dg     = preds - a\n    grad   = (df/g - f*dg/g**2)*len(labels)\n    hess   = np.ones(len(labels))\n    return grad, hess\n\nP.S. for this example thanks to @ravi20076",
    "2996703": "alexshumilov I simply took them from another competition artefacts. They were used there to good effect",
    "2996758": "According to comment [here][1] in Essay Scoring competition, `a` is mean of train target and `b` is std of train target. And the following blog [here][2] explains why\n\n[1]: https://www.kaggle.com/competitions/learning-agency-lab-automated-essay-scoring-2/discussion/511130#2872396\n[2]: https://zenn.dev/jackthekaggler/articles/cf988ca341e34ed83034",
    "2998762": "The derivation is a straightforward calculation. In standard notations,\n\\\\[\nO_{ij}=\\sum_k\\delta_{iy_k}\\delta_{j\\hat{y_k}}\n\\\\]\n\\\\[\nE_{ij}=p_i\\sum_k\\delta_{j\\hat{y_k}}\n\\\\]\nwhere \\\\(p_i\\\\) is the proportion of label \\\\(i\\\\) in the training set so that \\\\(\\sum_ip_i=1\\\\). Then\n\\\\[\n\\mathfrak{f}=\\sum_{i,j}(i-j)^2O_{ij}=\\sum_k\\sum_{i,j}(i-j)^2\\delta_{iy_k}\\delta_{j\\hat{y_k}}=\\sum_k(y_k-\\hat{y_k})^2\n\\\\]\n\\\\[\n\\mathfrak{g}=\\sum_{i,j}(i-j)^2E_{ij}=\\sum_k\\sum_ip_i(i-\\hat{y_k})^2=\\sum_k\\mathbb{E}_y[(\\hat{y_k}-y)^2]=\\sum_k\\left((\\hat{y_k}-a)^2+b\\right)\n\\\\]\n\nwhere \\\\(a\\\\) and \\\\(b\\\\) are respectively the mean and variance of \\\\(y_{train}\\\\). This follows from the identity \\\\(\\mathbb{E}_x[(t-x)^2]=(t-\\mathbb{E}[x])^2+\\mbox{Var}[x]\\\\). Then \\\\(1-\\kappa=\\mathfrak{f}/\\mathfrak{g}\\\\) which is a rational function of \\\\(\\hat{y_k}\\\\). Interestingly, it extends QWK to real-valued \\\\(\\hat{y_k}\\\\), and we can start taking the gradient and hessian with respect to \\\\(\\hat{y_k}\\\\).\n\nThe gradient \\\\(g\\\\) should just be \\\\(\\frac{\\mathfrak{f}'}{\\mathfrak{g}}-\\frac{\\mathfrak{f}\\mathfrak{g}'}{\\mathfrak{g}^2}\\\\). I am not sure why the scaling factor `len(labels)` is needed; it would affect the learning rate. The hessian \\\\(h\\\\) can also be analytically computed, but there is no reason to believe that it would be bounded below by some positive number. In practice, we only need \\\\(const. + g\\cdot y_{pred}+\\frac{1}{2}h\\cdot y_{pred}^2\\\\) to be a majorant of the objective function (see the [explanation](https://github.com/dmlc/xgboost/issues/1825) from the inventor of xgboost). Perhaps scaling the gradient by `len(labels)` and setting the hessian to \\\\(1\\\\) can make this happen."
  },
  "source": "meta"
}