{
  "id": 548619,
  "title": "Weighted R2 metric as custom loss function",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/548619",
  "author_name": "",
  "post_date": "2024-11-27T20:27:28.050921900Z",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>We can train our model by directly optimizing R2 score as in <a href=\"https://www.kaggle.com/code/eu1234/weighted-r2-as-custom-loss-function\" target=\"_blank\">this example</a>, using custom objective function:</p>\n<pre><code> ():\n    \n    residuals = y_pred - y_true\n    weighted_residual_sum = np.(sample_weight * (residuals ** ))\n\n    \n     weighted_residual_sum == :\n        weighted_residual_sum = \n\n    grad =  * sample_weight * residuals / weighted_residual_sum\n    hess =  * sample_weight / weighted_residual_sum\n\n     grad, hess\n</code></pre>",
  "messages": [
    {
      "id": "3057188",
      "postDate": "11/27/2024 20:27:28",
      "content": "<p>We can train our model by directly optimizing R2 score as in <a href=\"https://www.kaggle.com/code/eu1234/weighted-r2-as-custom-loss-function\" target=\"_blank\">this example</a>, using custom objective function:</p>\n<pre><code> ():\n    \n    residuals = y_pred - y_true\n    weighted_residual_sum = np.(sample_weight * (residuals ** ))\n\n    \n     weighted_residual_sum == :\n        weighted_residual_sum = \n\n    grad =  * sample_weight * residuals / weighted_residual_sum\n    hess =  * sample_weight / weighted_residual_sum\n\n     grad, hess\n</code></pre>",
      "rawMarkdown": "We can train our model by directly optimizing R2 score as in [this example](https://www.kaggle.com/code/eu1234/weighted-r2-as-custom-loss-function), using custom objective function:\n\n```python\ndef weighted_r2_loss(y_true, y_pred, sample_weight):\n    \"\"\"\n    Loss function for sample-weighted R2 score\n    \"\"\"\n    residuals = y_pred - y_true\n    weighted_residual_sum = np.sum(sample_weight * (residuals ** 2))\n\n    # Avoid division by zero\n    if weighted_residual_sum == 0:\n        weighted_residual_sum = 1e-10\n\n    grad = 2 * sample_weight * residuals / weighted_residual_sum\n    hess = 2 * sample_weight / weighted_residual_sum\n\n    return grad, hess\n```",
      "votes": null
    },
    {
      "id": "3057203",
      "postDate": "11/27/2024 20:59:47",
      "content": "<p>Have you tried it? How was the performance ?</p>",
      "rawMarkdown": "Have you tried it? How was the performance ?",
      "votes": null
    },
    {
      "id": "3058214",
      "postDate": "11/29/2024 08:21:58",
      "content": "<p>Except for the scaling parameter <code>weighted_residual_sum</code>, this should be the same as MSE with weights out of the box.</p>\n<p>Edit: I think it should be the weighted sum of zero-centred observations and not the weighted sum of squared residuals.</p>",
      "rawMarkdown": "Except for the scaling parameter `weighted_residual_sum`, this should be the same as MSE with weights out of the box.\n\nEdit: I think it should be the weighted sum of zero-centred observations and not the weighted sum of squared residuals.",
      "votes": null
    },
    {
      "id": "3058254",
      "postDate": "11/29/2024 09:06:58",
      "content": "<p>Can you expand the idea please?</p>",
      "rawMarkdown": "Can you expand the idea please?",
      "votes": null
    },
    {
      "id": "3058307",
      "postDate": "11/29/2024 10:06:43",
      "content": "<p>The sample weighted zero-mean R-squared score (R2) is <code>R2 = 1 - sum(sample_weight * (y_pred - y_true) ** 2) / sum(sample_weight * y_true ** 2)</code>. Weighted mean squared error (WMSE) is <code>WMSE = sum(sample_weight * (y_pred - y_true) ** 2) / sum(sample_weight)</code>.<br>\nWe can write <code>R2 = 1 - WMSE * sum(sample_weight) / sum(sample_weight * y_true ** 2)</code>.<br>\nSince <code>scale = sum(sample_weight) / sum(sample_weight * y_true ** 2)</code> is independent of the predictions <code>y_pred</code>, the first derivative of R2 will be a scaled version of the first derivative of WMSE: <code>R2' = -WMSE' * scale</code>. Naturally, the same is true for the second derviative: <code>R2'' = -WMSE'' * scale</code>.</p>\n<p>Given this, I think you have a mistake in your function, since your derivatives are divided by the weighted sum of squared residuals and not the weighted sum of zero-centred observations.</p>",
      "rawMarkdown": "The sample weighted zero-mean R-squared score (R2) is `R2 = 1 - sum(sample_weight * (y_pred - y_true) ** 2) / sum(sample_weight * y_true ** 2)`. Weighted mean squared error (WMSE) is `WMSE = sum(sample_weight * (y_pred - y_true) ** 2) / sum(sample_weight)`.\nWe can write `R2 = 1 - WMSE * sum(sample_weight) / sum(sample_weight * y_true ** 2)`.\nSince `scale = sum(sample_weight) / sum(sample_weight * y_true ** 2)` is independent of the predictions `y_pred`, the first derivative of R2 will be a scaled version of the first derivative of WMSE: `R2' = -WMSE' * scale`. Naturally, the same is true for the second derviative: `R2'' = -WMSE'' * scale`.\n\nGiven this, I think you have a mistake in your function, since your derivatives are divided by the weighted sum of squared residuals and not the weighted sum of zero-centred observations.",
      "votes": null
    },
    {
      "id": "3058332",
      "postDate": "11/29/2024 10:48:24",
      "content": "<p>I have tried this version too and the optimization was something better when normalizing by the sum of the weighted predicted values at least in CV, but we can test them both of course. This version should be something like this:</p>\n<pre><code> ():\n    \n    residuals = y_pred - y_true\n    weighted_true_sum = np.(sample_weight * (y_true ** ))\n\n    \n     weighted_true_sum == :\n        weighted_true_sum = \n\n    grad =  * sample_weight * residuals / weighted_true_sum\n    hess =  * sample_weight / weighted_true_sum\n\n     grad, hess\n</code></pre>",
      "rawMarkdown": "I have tried this version too and the optimization was something better when normalizing by the sum of the weighted predicted values at least in CV, but we can test them both of course. This version should be something like this:\n```python\ndef weighted_r2_loss(y_true, y_pred, sample_weight):\n    \"\"\" Loss function for sample-weighted R2 score \"\"\"\n    residuals = y_pred - y_true\n    weighted_true_sum = np.sum(sample_weight * (y_true ** 2))\n\n    # Avoid division by zero\n    if weighted_true_sum == 0:\n        weighted_true_sum = 1e-10\n   \n    grad = 2 * sample_weight * residuals / weighted_true_sum\n    hess = 2 * sample_weight / weighted_true_sum\n\n    return grad, hess\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3057203,
      "author_name": "aymanallawi",
      "author_url": "",
      "post_date": "11/27/2024 20:59:47",
      "content": "<p>Have you tried it? How was the performance ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3058214,
      "author_name": "jonassend",
      "author_url": "",
      "post_date": "11/29/2024 08:21:58",
      "content": "<p>Except for the scaling parameter <code>weighted_residual_sum</code>, this should be the same as MSE with weights out of the box.</p>\n<p>Edit: I think it should be the weighted sum of zero-centred observations and not the weighted sum of squared residuals.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3058254,
          "author_name": "eu1234",
          "author_url": "",
          "post_date": "11/29/2024 09:06:58",
          "content": "<p>Can you expand the idea please?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3058307,
              "author_name": "jonassend",
              "author_url": "",
              "post_date": "11/29/2024 10:06:43",
              "content": "<p>The sample weighted zero-mean R-squared score (R2) is <code>R2 = 1 - sum(sample_weight * (y_pred - y_true) ** 2) / sum(sample_weight * y_true ** 2)</code>. Weighted mean squared error (WMSE) is <code>WMSE = sum(sample_weight * (y_pred - y_true) ** 2) / sum(sample_weight)</code>.<br>\nWe can write <code>R2 = 1 - WMSE * sum(sample_weight) / sum(sample_weight * y_true ** 2)</code>.<br>\nSince <code>scale = sum(sample_weight) / sum(sample_weight * y_true ** 2)</code> is independent of the predictions <code>y_pred</code>, the first derivative of R2 will be a scaled version of the first derivative of WMSE: <code>R2' = -WMSE' * scale</code>. Naturally, the same is true for the second derviative: <code>R2'' = -WMSE'' * scale</code>.</p>\n<p>Given this, I think you have a mistake in your function, since your derivatives are divided by the weighted sum of squared residuals and not the weighted sum of zero-centred observations.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3058332,
                  "author_name": "eu1234",
                  "author_url": "",
                  "post_date": "11/29/2024 10:48:24",
                  "content": "<p>I have tried this version too and the optimization was something better when normalizing by the sum of the weighted predicted values at least in CV, but we can test them both of course. This version should be something like this:</p>\n<pre><code> ():\n    \n    residuals = y_pred - y_true\n    weighted_true_sum = np.(sample_weight * (y_true ** ))\n\n    \n     weighted_true_sum == :\n        weighted_true_sum = \n\n    grad =  * sample_weight * residuals / weighted_true_sum\n    hess =  * sample_weight / weighted_true_sum\n\n     grad, hess\n</code></pre>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3057188": "We can train our model by directly optimizing R2 score as in [this example](https://www.kaggle.com/code/eu1234/weighted-r2-as-custom-loss-function), using custom objective function:\n\n```python\ndef weighted_r2_loss(y_true, y_pred, sample_weight):\n    \"\"\"\n    Loss function for sample-weighted R2 score\n    \"\"\"\n    residuals = y_pred - y_true\n    weighted_residual_sum = np.sum(sample_weight * (residuals ** 2))\n\n    # Avoid division by zero\n    if weighted_residual_sum == 0:\n        weighted_residual_sum = 1e-10\n\n    grad = 2 * sample_weight * residuals / weighted_residual_sum\n    hess = 2 * sample_weight / weighted_residual_sum\n\n    return grad, hess\n```",
    "3057203": "Have you tried it? How was the performance ?",
    "3058214": "Except for the scaling parameter `weighted_residual_sum`, this should be the same as MSE with weights out of the box.\n\nEdit: I think it should be the weighted sum of zero-centred observations and not the weighted sum of squared residuals.",
    "3058254": "Can you expand the idea please?",
    "3058307": "The sample weighted zero-mean R-squared score (R2) is `R2 = 1 - sum(sample_weight * (y_pred - y_true) ** 2) / sum(sample_weight * y_true ** 2)`. Weighted mean squared error (WMSE) is `WMSE = sum(sample_weight * (y_pred - y_true) ** 2) / sum(sample_weight)`.\nWe can write `R2 = 1 - WMSE * sum(sample_weight) / sum(sample_weight * y_true ** 2)`.\nSince `scale = sum(sample_weight) / sum(sample_weight * y_true ** 2)` is independent of the predictions `y_pred`, the first derivative of R2 will be a scaled version of the first derivative of WMSE: `R2' = -WMSE' * scale`. Naturally, the same is true for the second derviative: `R2'' = -WMSE'' * scale`.\n\nGiven this, I think you have a mistake in your function, since your derivatives are divided by the weighted sum of squared residuals and not the weighted sum of zero-centred observations.",
    "3058332": "I have tried this version too and the optimization was something better when normalizing by the sum of the weighted predicted values at least in CV, but we can test them both of course. This version should be something like this:\n```python\ndef weighted_r2_loss(y_true, y_pred, sample_weight):\n    \"\"\" Loss function for sample-weighted R2 score \"\"\"\n    residuals = y_pred - y_true\n    weighted_true_sum = np.sum(sample_weight * (y_true ** 2))\n\n    # Avoid division by zero\n    if weighted_true_sum == 0:\n        weighted_true_sum = 1e-10\n   \n    grad = 2 * sample_weight * residuals / weighted_true_sum\n    hess = 2 * sample_weight / weighted_true_sum\n\n    return grad, hess\n```"
  },
  "source": "meta"
}