{
  "id": 544866,
  "title": "Implementing competition's metric (sample-weighted zero-mean R2 )",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/544866",
  "author_name": "",
  "post_date": "2024-11-07T10:21:52.606986100Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>As many are wondering why their CV score is greatly different from the LB score - we should first score our CV with the same metric as the competition uses and not just the sklearn's r2_score even with it's sample weight as it's not always zero mean. So this should help:</p>\n<pre><code> ():\n    \n    \n    residuals = sample_weight * (y_true - y_pred) ** \n    weighted_residual_sum = np.(residuals)\n\n    \n    weighted_true_sum = np.(sample_weight * (y_true) ** )\n\n    \n    w_r2 =  - weighted_residual_sum / weighted_true_sum\n\n     w_r2\n</code></pre>",
  "messages": [
    {
      "id": "3038765",
      "postDate": "11/07/2024 10:21:52",
      "content": "<p>As many are wondering why their CV score is greatly different from the LB score - we should first score our CV with the same metric as the competition uses and not just the sklearn's r2_score even with it's sample weight as it's not always zero mean. So this should help:</p>\n<pre><code> ():\n    \n    \n    residuals = sample_weight * (y_true - y_pred) ** \n    weighted_residual_sum = np.(residuals)\n\n    \n    weighted_true_sum = np.(sample_weight * (y_true) ** )\n\n    \n    w_r2 =  - weighted_residual_sum / weighted_true_sum\n\n     w_r2\n</code></pre>",
      "rawMarkdown": "As many are wondering why their CV score is greatly different from the LB score - we should first score our CV with the same metric as the competition uses and not just the sklearn's r2_score even with it's sample weight as it's not always zero mean. So this should help:\n\n```python\ndef weighted_r2(y_true, y_pred, sample_weight):\n    \"\"\"\n    Compute the sample-weighted zero-mean R2 score\n    \"\"\"\n    # Calculate weighted sum of squared residuals (numerator)\n    residuals = sample_weight * (y_true - y_pred) ** 2\n    weighted_residual_sum = np.sum(residuals)\n\n    # Calculate weighted sum of squared true values (denominator)\n    weighted_true_sum = np.sum(sample_weight * (y_true) ** 2)\n\n    # Calculate weighted R2\n    w_r2 = 1 - weighted_residual_sum / weighted_true_sum\n\n    return w_r2\n```",
      "votes": null
    },
    {
      "id": "3038886",
      "postDate": "11/07/2024 13:50:30",
      "content": "<p>I did this and still the CV score is not aligned with the leaderboard <a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a> </p>",
      "rawMarkdown": "I did this and still the CV score is not aligned with the leaderboard @eu1234",
      "votes": null
    },
    {
      "id": "3038902",
      "postDate": "11/07/2024 14:02:05",
      "content": "<p>Well, scoring alignement depends on the model and the CV approach as you perfectly know :) but at least with the correct metric we can avoid comparing apples with oranges</p>",
      "rawMarkdown": "Well, scoring alignement depends on the model and the CV approach as you perfectly know :) but at least with the correct metric we can avoid comparing apples with oranges",
      "votes": null
    },
    {
      "id": "3040814",
      "postDate": "11/09/2024 15:50:34",
      "content": "<p>I was using sklearn's r2 score with sample weights so far. I tried this but it seems like they are almost always same. They start to change after 5th decimal.</p>",
      "rawMarkdown": "I was using sklearn's r2 score with sample weights so far. I tried this but it seems like they are almost always same. They start to change after 5th decimal.",
      "votes": null
    },
    {
      "id": "3040836",
      "postDate": "11/09/2024 16:20:56",
      "content": "<p>It's just an implementation of the exact formula competition provided, but it's an individual choice what metric is good enough to use.</p>",
      "rawMarkdown": "It's just an implementation of the exact formula competition provided, but it's an individual choice what metric is good enough to use.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3038886,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "11/07/2024 13:50:30",
      "content": "<p>I did this and still the CV score is not aligned with the leaderboard <a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 3038902,
          "author_name": "eu1234",
          "author_url": "",
          "post_date": "11/07/2024 14:02:05",
          "content": "<p>Well, scoring alignement depends on the model and the CV approach as you perfectly know :) but at least with the correct metric we can avoid comparing apples with oranges</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3040814,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "11/09/2024 15:50:34",
      "content": "<p>I was using sklearn's r2 score with sample weights so far. I tried this but it seems like they are almost always same. They start to change after 5th decimal.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3040836,
          "author_name": "eu1234",
          "author_url": "",
          "post_date": "11/09/2024 16:20:56",
          "content": "<p>It's just an implementation of the exact formula competition provided, but it's an individual choice what metric is good enough to use.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3038765": "As many are wondering why their CV score is greatly different from the LB score - we should first score our CV with the same metric as the competition uses and not just the sklearn's r2_score even with it's sample weight as it's not always zero mean. So this should help:\n\n```python\ndef weighted_r2(y_true, y_pred, sample_weight):\n    \"\"\"\n    Compute the sample-weighted zero-mean R2 score\n    \"\"\"\n    # Calculate weighted sum of squared residuals (numerator)\n    residuals = sample_weight * (y_true - y_pred) ** 2\n    weighted_residual_sum = np.sum(residuals)\n\n    # Calculate weighted sum of squared true values (denominator)\n    weighted_true_sum = np.sum(sample_weight * (y_true) ** 2)\n\n    # Calculate weighted R2\n    w_r2 = 1 - weighted_residual_sum / weighted_true_sum\n\n    return w_r2\n```",
    "3038886": "I did this and still the CV score is not aligned with the leaderboard @eu1234",
    "3038902": "Well, scoring alignement depends on the model and the CV approach as you perfectly know :) but at least with the correct metric we can avoid comparing apples with oranges",
    "3040814": "I was using sklearn's r2 score with sample weights so far. I tried this but it seems like they are almost always same. They start to change after 5th decimal.",
    "3040836": "It's just an implementation of the exact formula competition provided, but it's an individual choice what metric is good enough to use."
  },
  "source": "meta"
}