{
  "id": 540554,
  "title": "Are metrics like MAE or MSE superior to R-squared for clipped targets ?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/540554",
  "author_name": "",
  "post_date": "2024-10-15T04:17:25.341224300Z",
  "votes": 21,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Looking at the data, the target is clipped within the range of ±5.</p>\n<p>I think that using a clipped target for the R-squared metric might distort the model's training and increase the likelihood that the model won't be able to fully explain the original variance.</p>\n<p>Wouldn't MAE or MSE be better than R-squared when clipping the target?</p>",
  "messages": [
    {
      "id": "3017625",
      "postDate": "10/15/2024 04:17:25",
      "content": "<p>Looking at the data, the target is clipped within the range of ±5.</p>\n<p>I think that using a clipped target for the R-squared metric might distort the model's training and increase the likelihood that the model won't be able to fully explain the original variance.</p>\n<p>Wouldn't MAE or MSE be better than R-squared when clipping the target?</p>",
      "rawMarkdown": "Looking at the data, the target is clipped within the range of ±5.\n\nI think that using a clipped target for the R-squared metric might distort the model's training and increase the likelihood that the model won't be able to fully explain the original variance.\n\nWouldn't MAE or MSE be better than R-squared when clipping the target?",
      "votes": null
    },
    {
      "id": "3017779",
      "postDate": "10/15/2024 07:39:52",
      "content": "<p>I do agree.</p>",
      "rawMarkdown": "I do agree.",
      "votes": null
    },
    {
      "id": "3017792",
      "postDate": "10/15/2024 07:51:46",
      "content": "<p>R-squared measures variance:  when the target is clipped, the variance of the true values is  reduced ,Since R-squared compares the model's predictions to the variance of the true (clipped) target,It may give an incorrect prediction of how well the model will perform, In addition to , R-squared may not reflect how well the model generalizes beyond the clipped range.</p>",
      "rawMarkdown": "R-squared measures variance:  when the target is clipped, the variance of the true values is  reduced ,Since R-squared compares the model's predictions to the variance of the true (clipped) target,It may give an incorrect prediction of how well the model will perform, In addition to , R-squared may not reflect how well the model generalizes beyond the clipped range.",
      "votes": null
    },
    {
      "id": "3018037",
      "postDate": "10/15/2024 13:28:31",
      "content": "<p><a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> Thank you for agreeing. I'm glad to confirm that my thinking wasn't wrong. </p>",
      "rawMarkdown": "ulrich07 Thank you for agreeing. I'm glad to confirm that my thinking wasn't wrong.",
      "votes": null
    },
    {
      "id": "3018039",
      "postDate": "10/15/2024 13:29:46",
      "content": "<p><a href=\"https://www.kaggle.com/adnanalaref\" target=\"_blank\">@adnanalaref</a> Thank you for your comment. Yes, exactly. It would be great if the host or Kaggle could address this topic.</p>",
      "rawMarkdown": "adnanalaref Thank you for your comment. Yes, exactly. It would be great if the host or Kaggle could address this topic.",
      "votes": null
    },
    {
      "id": "3020096",
      "postDate": "10/17/2024 06:44:25",
      "content": "<p>most models will perform worse in a more volatile market or on a more volatile instrument</p>\n<p>adjusting the the MSE(SSE) by variance helps evaluate the model's performance under different conditions</p>\n<p>without it it is possible that best performing model overfit market volatility of the testing set</p>\n<p>(if I have to guess why they chose R square)</p>",
      "rawMarkdown": "most models will perform worse in a more volatile market or on a more volatile instrument\n\nadjusting the the MSE(SSE) by variance helps evaluate the model's performance under different conditions\n\nwithout it it is possible that best performing model overfit market volatility of the testing set\n\n(if I have to guess why they chose R square)",
      "votes": null
    },
    {
      "id": "3020846",
      "postDate": "10/18/2024 00:26:22",
      "content": "<p><a href=\"https://www.kaggle.com/dc260123\" target=\"_blank\">@dc260123</a> Thank you. I’m starting to understand why the host wants to observe trends with R2. As a participant in the competition, however, using metrics like MSE or MAE might be better for building a more reliable model.</p>",
      "rawMarkdown": "dc260123 Thank you. I’m starting to understand why the host wants to observe trends with R2. As a participant in the competition, however, using metrics like MSE or MAE might be better for building a more reliable model.",
      "votes": null
    },
    {
      "id": "3021095",
      "postDate": "10/18/2024 07:28:05",
      "content": "<p>Yes, <a href=\"https://www.kaggle.com/dc260123\" target=\"_blank\">@dc260123</a> </p>",
      "rawMarkdown": "Yes, @dc260123",
      "votes": null
    },
    {
      "id": "3030034",
      "postDate": "10/28/2024 04:48:26",
      "content": "<p>Targets are normalized though. Don't you think mae or mse is applicable here?</p>",
      "rawMarkdown": "Targets are normalized though. Don't you think mae or mse is applicable here?",
      "votes": null
    },
    {
      "id": "3030072",
      "postDate": "10/28/2024 06:10:22",
      "content": "<p>That's also why clipping predictions makes your score worse.</p>",
      "rawMarkdown": "That's also why clipping predictions makes your score worse.",
      "votes": null
    },
    {
      "id": "3030174",
      "postDate": "10/28/2024 09:03:06",
      "content": "<p>Let’s look at the argument by itself: we don’t have the variance of the data because it is clipped so it is not a good idea to use R-squared because models trained using the “wrong” variance will be pushed to overfit the “wrong” variance. <br>\nBut the ultimate goal is not to explain the variance. Models’ outputs are aggregated as signals in trading system and you need a metric to measure model performance.<br>\nI highly doubt this is what they are using but something minimal that does not pick the wrong winner. We can guess what the intention of why did they pick this but it’s up to them to decide what’s valuable to them.</p>\n<p>Another way to look at it is that you want to award/punish predictions of large swing less(by dividing the loss by variance) because by nature these are most likely a result of new information being available and you can’t predict arbitrary event in the future with 90 features.(if the model made the right call — it’s most likely luck, if it did not, that’s fine too because that’s why they have traders) </p>\n<p>The goal here is to capture some sort of pattern from the features that can be captured not to predict the future.</p>",
      "rawMarkdown": "Let’s look at the argument by itself: we don’t have the variance of the data because it is clipped so it is not a good idea to use R-squared because models trained using the “wrong” variance will be pushed to overfit the “wrong” variance. \nBut the ultimate goal is not to explain the variance. Models’ outputs are aggregated as signals in trading system and you need a metric to measure model performance.\nI highly doubt this is what they are using but something minimal that does not pick the wrong winner. We can guess what the intention of why did they pick this but it’s up to them to decide what’s valuable to them.\n\nAnother way to look at it is that you want to award/punish predictions of large swing less(by dividing the loss by variance) because by nature these are most likely a result of new information being available and you can’t predict arbitrary event in the future with 90 features.(if the model made the right call — it’s most likely luck, if it did not, that’s fine too because that’s why they have traders) \n\nThe goal here is to capture some sort of pattern from the features that can be captured not to predict the future.",
      "votes": null
    },
    {
      "id": "3030459",
      "postDate": "10/28/2024 14:38:07",
      "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> <a href=\"https://www.kaggle.com/dc260123\" target=\"_blank\">@dc260123</a> Thank you for comment ! This is my linear regression sample result.<br>\nI made this <a href=\"https://www.kaggle.com/code/chumajin/metric-simulation-of-clipped-r2-score/notebook\" target=\"_blank\">public</a></p>\n<p>without clipping</p>\n<pre><code>=.\n=.\n=.\n</code></pre>\n<p>with clipping</p>\n<pre><code>=.\n=.\n=.\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F8a6cc44efee1c31d2073a6698f913344%2FClipboard07.jpg?generation=1730126128798795&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "gunesevitan @dc260123 Thank you for comment ! This is my linear regression sample result.\nI made this [public](https://www.kaggle.com/code/chumajin/metric-simulation-of-clipped-r2-score/notebook)\n\nwithout clipping\n~~~\nr2_true=0.9203016112934438\nmae_true=4.025541885409228\nmse_true=25.207771336715304\n~~~\n\nwith clipping\n~~~\nr2_clipped=0.7607555181735448\nmae_clipped=1.8509716653987454\nmse_clipped=5.156533423176572\n~~~\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F8a6cc44efee1c31d2073a6698f913344%2FClipboard07.jpg?generation=1730126128798795&alt=media)",
      "votes": null
    },
    {
      "id": "3030896",
      "postDate": "10/29/2024 03:17:39",
      "content": "<p>actually I just realized that they probably did make a mistake here if the loss is calculated using overall R^2 instead of averages like what they did before. <br>\nIf they did then using R^2 as a target for SGD based algorithm will hurt your lb score even though it will yield a 'better' model</p>",
      "rawMarkdown": "actually I just realized that they probably did make a mistake here if the loss is calculated using overall R^2 instead of averages like what they did before. \nIf they did then using R^2 as a target for SGD based algorithm will hurt your lb score even though it will yield a 'better' model",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3017779,
      "author_name": "ulrich07",
      "author_url": "",
      "post_date": "10/15/2024 07:39:52",
      "content": "<p>I do agree.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3018037,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "10/15/2024 13:28:31",
          "content": "<p><a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> Thank you for agreeing. I'm glad to confirm that my thinking wasn't wrong. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3017792,
      "author_name": "adnanalaref",
      "author_url": "",
      "post_date": "10/15/2024 07:51:46",
      "content": "<p>R-squared measures variance:  when the target is clipped, the variance of the true values is  reduced ,Since R-squared compares the model's predictions to the variance of the true (clipped) target,It may give an incorrect prediction of how well the model will perform, In addition to , R-squared may not reflect how well the model generalizes beyond the clipped range.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3018039,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "10/15/2024 13:29:46",
          "content": "<p><a href=\"https://www.kaggle.com/adnanalaref\" target=\"_blank\">@adnanalaref</a> Thank you for your comment. Yes, exactly. It would be great if the host or Kaggle could address this topic.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3020096,
      "author_name": "dc260123",
      "author_url": "",
      "post_date": "10/17/2024 06:44:25",
      "content": "<p>most models will perform worse in a more volatile market or on a more volatile instrument</p>\n<p>adjusting the the MSE(SSE) by variance helps evaluate the model's performance under different conditions</p>\n<p>without it it is possible that best performing model overfit market volatility of the testing set</p>\n<p>(if I have to guess why they chose R square)</p>",
      "votes": null,
      "replies": [
        {
          "id": 3020846,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "10/18/2024 00:26:22",
          "content": "<p><a href=\"https://www.kaggle.com/dc260123\" target=\"_blank\">@dc260123</a> Thank you. I’m starting to understand why the host wants to observe trends with R2. As a participant in the competition, however, using metrics like MSE or MAE might be better for building a more reliable model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3021095,
          "author_name": "muhammedtausif",
          "author_url": "",
          "post_date": "10/18/2024 07:28:05",
          "content": "<p>Yes, <a href=\"https://www.kaggle.com/dc260123\" target=\"_blank\">@dc260123</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3030034,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "10/28/2024 04:48:26",
          "content": "<p>Targets are normalized though. Don't you think mae or mse is applicable here?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3030174,
              "author_name": "dc260123",
              "author_url": "",
              "post_date": "10/28/2024 09:03:06",
              "content": "<p>Let’s look at the argument by itself: we don’t have the variance of the data because it is clipped so it is not a good idea to use R-squared because models trained using the “wrong” variance will be pushed to overfit the “wrong” variance. <br>\nBut the ultimate goal is not to explain the variance. Models’ outputs are aggregated as signals in trading system and you need a metric to measure model performance.<br>\nI highly doubt this is what they are using but something minimal that does not pick the wrong winner. We can guess what the intention of why did they pick this but it’s up to them to decide what’s valuable to them.</p>\n<p>Another way to look at it is that you want to award/punish predictions of large swing less(by dividing the loss by variance) because by nature these are most likely a result of new information being available and you can’t predict arbitrary event in the future with 90 features.(if the model made the right call — it’s most likely luck, if it did not, that’s fine too because that’s why they have traders) </p>\n<p>The goal here is to capture some sort of pattern from the features that can be captured not to predict the future.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3030459,
                  "author_name": "chumajin",
                  "author_url": "",
                  "post_date": "10/28/2024 14:38:07",
                  "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> <a href=\"https://www.kaggle.com/dc260123\" target=\"_blank\">@dc260123</a> Thank you for comment ! This is my linear regression sample result.<br>\nI made this <a href=\"https://www.kaggle.com/code/chumajin/metric-simulation-of-clipped-r2-score/notebook\" target=\"_blank\">public</a></p>\n<p>without clipping</p>\n<pre><code>=.\n=.\n=.\n</code></pre>\n<p>with clipping</p>\n<pre><code>=.\n=.\n=.\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F8a6cc44efee1c31d2073a6698f913344%2FClipboard07.jpg?generation=1730126128798795&amp;alt=media\" alt=\"\"></p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3030896,
                      "author_name": "dc260123",
                      "author_url": "",
                      "post_date": "10/29/2024 03:17:39",
                      "content": "<p>actually I just realized that they probably did make a mistake here if the loss is calculated using overall R^2 instead of averages like what they did before. <br>\nIf they did then using R^2 as a target for SGD based algorithm will hurt your lb score even though it will yield a 'better' model</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3030072,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "10/28/2024 06:10:22",
      "content": "<p>That's also why clipping predictions makes your score worse.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3017625": "Looking at the data, the target is clipped within the range of ±5.\n\nI think that using a clipped target for the R-squared metric might distort the model's training and increase the likelihood that the model won't be able to fully explain the original variance.\n\nWouldn't MAE or MSE be better than R-squared when clipping the target?",
    "3017779": "I do agree.",
    "3017792": "R-squared measures variance:  when the target is clipped, the variance of the true values is  reduced ,Since R-squared compares the model's predictions to the variance of the true (clipped) target,It may give an incorrect prediction of how well the model will perform, In addition to , R-squared may not reflect how well the model generalizes beyond the clipped range.",
    "3018037": "ulrich07 Thank you for agreeing. I'm glad to confirm that my thinking wasn't wrong.",
    "3018039": "adnanalaref Thank you for your comment. Yes, exactly. It would be great if the host or Kaggle could address this topic.",
    "3020096": "most models will perform worse in a more volatile market or on a more volatile instrument\n\nadjusting the the MSE(SSE) by variance helps evaluate the model's performance under different conditions\n\nwithout it it is possible that best performing model overfit market volatility of the testing set\n\n(if I have to guess why they chose R square)",
    "3020846": "dc260123 Thank you. I’m starting to understand why the host wants to observe trends with R2. As a participant in the competition, however, using metrics like MSE or MAE might be better for building a more reliable model.",
    "3021095": "Yes, @dc260123",
    "3030034": "Targets are normalized though. Don't you think mae or mse is applicable here?",
    "3030072": "That's also why clipping predictions makes your score worse.",
    "3030174": "Let’s look at the argument by itself: we don’t have the variance of the data because it is clipped so it is not a good idea to use R-squared because models trained using the “wrong” variance will be pushed to overfit the “wrong” variance. \nBut the ultimate goal is not to explain the variance. Models’ outputs are aggregated as signals in trading system and you need a metric to measure model performance.\nI highly doubt this is what they are using but something minimal that does not pick the wrong winner. We can guess what the intention of why did they pick this but it’s up to them to decide what’s valuable to them.\n\nAnother way to look at it is that you want to award/punish predictions of large swing less(by dividing the loss by variance) because by nature these are most likely a result of new information being available and you can’t predict arbitrary event in the future with 90 features.(if the model made the right call — it’s most likely luck, if it did not, that’s fine too because that’s why they have traders) \n\nThe goal here is to capture some sort of pattern from the features that can be captured not to predict the future.",
    "3030459": "gunesevitan @dc260123 Thank you for comment ! This is my linear regression sample result.\nI made this [public](https://www.kaggle.com/code/chumajin/metric-simulation-of-clipped-r2-score/notebook)\n\nwithout clipping\n~~~\nr2_true=0.9203016112934438\nmae_true=4.025541885409228\nmse_true=25.207771336715304\n~~~\n\nwith clipping\n~~~\nr2_clipped=0.7607555181735448\nmae_clipped=1.8509716653987454\nmse_clipped=5.156533423176572\n~~~\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F8a6cc44efee1c31d2073a6698f913344%2FClipboard07.jpg?generation=1730126128798795&alt=media)",
    "3030896": "actually I just realized that they probably did make a mistake here if the loss is calculated using overall R^2 instead of averages like what they did before. \nIf they did then using R^2 as a target for SGD based algorithm will hurt your lb score even though it will yield a 'better' model"
  },
  "source": "meta"
}