{
  "id": 551741,
  "title": "Predicted vs. Actuals?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/551741",
  "author_name": "",
  "post_date": "2024-12-15T07:56:49.356926500Z",
  "votes": 6,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Does anyone have predicted vs. actuals plots they can share? Even for an average or basic model? I'm trying to figure out if my current approach is just poor or has a bug. My plot is basically a flatline.</p>",
  "messages": [
    {
      "id": "3072445",
      "postDate": "12/15/2024 07:56:49",
      "content": "<p>Does anyone have predicted vs. actuals plots they can share? Even for an average or basic model? I'm trying to figure out if my current approach is just poor or has a bug. My plot is basically a flatline.</p>",
      "rawMarkdown": "Does anyone have predicted vs. actuals plots they can share? Even for an average or basic model? I'm trying to figure out if my current approach is just poor or has a bug. My plot is basically a flatline.",
      "votes": null
    },
    {
      "id": "3072448",
      "postDate": "12/15/2024 08:05:45",
      "content": "<p>You may take this from any public kernel and make the plot <a href=\"https://www.kaggle.com/jimbeno\" target=\"_blank\">@jimbeno</a> </p>",
      "rawMarkdown": "You may take this from any public kernel and make the plot @jimbeno",
      "votes": null
    },
    {
      "id": "3072451",
      "postDate": "12/15/2024 08:14:28",
      "content": "<p>try to plot your predictions on a different axis :)</p>",
      "rawMarkdown": "try to plot your predictions on a different axis :)",
      "votes": null
    },
    {
      "id": "3072548",
      "postDate": "12/15/2024 11:35:39",
      "content": "<p>I made visual evaluations in my LGBM models. </p>\n<p>It fails miserably to predict the range -5 to +5.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2Fb6fe6814782903d2d40cdfd0ce8c1931%2Fbad%20model.png?generation=1734261949016183&amp;alt=media\" alt=\"\"></p>\n<p>The following chart should be a good model<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2F3dd52540975a6cf9c7e0cec55f601575%2Fgood%20model.png?generation=1734261986692256&amp;alt=media\" alt=\"\"></p>\n<p>I can see why it fails to predict the entire range. We can see that most values are in a narrow range between +1 and -1.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2F669b2ed66ca4e7bc25d3c3e656007f6d%2Fdensity.png?generation=1734262132420656&amp;alt=media\" alt=\"\"></p>\n<p>I think that even the top leaderboards solutions aren't doing a really good job predicting the responder. The maximum score is 1 and top leaderboards are 0.01x. I think their predicted x actual plot should be closer to my bad model than to the good one.</p>",
      "rawMarkdown": "I made visual evaluations in my LGBM models. \n\nIt fails miserably to predict the range -5 to +5.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2Fb6fe6814782903d2d40cdfd0ce8c1931%2Fbad%20model.png?generation=1734261949016183&alt=media)\n\nThe following chart should be a good model\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2F3dd52540975a6cf9c7e0cec55f601575%2Fgood%20model.png?generation=1734261986692256&alt=media)\n\nI can see why it fails to predict the entire range. We can see that most values are in a narrow range between +1 and -1.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2F669b2ed66ca4e7bc25d3c3e656007f6d%2Fdensity.png?generation=1734262132420656&alt=media)\n\nI think that even the top leaderboards solutions aren't doing a really good job predicting the responder. The maximum score is 1 and top leaderboards are 0.01x. I think their predicted x actual plot should be closer to my bad model than to the good one.",
      "votes": null
    },
    {
      "id": "3072668",
      "postDate": "12/15/2024 14:18:08",
      "content": "<p>The model doesn’t “dare” to predict large values. This is due to the extreme long tailed target distribution. You can design a weighted loss to encourage the model to predict larger values, similar to what is done in a classification task using the Focal Loss. But this implies that you have put some prior knowledge or assumptions on the hidden test data distributions. <br>\nNice plot and visuals!</p>",
      "rawMarkdown": "The model doesn’t “dare” to predict large values. This is due to the extreme long tailed target distribution. You can design a weighted loss to encourage the model to predict larger values, similar to what is done in a classification task using the Focal Loss. But this implies that you have put some prior knowledge or assumptions on the hidden test data distributions. \nNice plot and visuals!",
      "votes": null
    },
    {
      "id": "3072954",
      "postDate": "12/15/2024 20:37:38",
      "content": "<p>This is from a super simple mlp evaluated on symbol_id 1</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11063283%2Fbee7e11b073c7e442b32c4fd0416c48a%2Fscatter.png?generation=1734294957546393&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "This is from a super simple mlp evaluated on symbol_id 1\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11063283%2Fbee7e11b073c7e442b32c4fd0416c48a%2Fscatter.png?generation=1734294957546393&alt=media)",
      "votes": null
    },
    {
      "id": "3073014",
      "postDate": "12/15/2024 23:16:56",
      "content": "<p>Every time I have tried to update my loss function to stop it being \"safe\" and keeping close to 0 I ended up getting worse results. </p>",
      "rawMarkdown": "Every time I have tried to update my loss function to stop it being \"safe\" and keeping close to 0 I ended up getting worse results.",
      "votes": null
    },
    {
      "id": "3073025",
      "postDate": "12/16/2024 00:10:01",
      "content": "<p>Okay, that's basically what I'm seeing. I've never had a predicted vs. actuals plot look so flat 😆 So my takeaway is it may not be a bug in my code, but just a poor model that's not able to learn anything? Looks like it's always predicting between -1 and 1. Almost seems like it could be a bug with the target being normalized and then not de-normalized. But I'm getting this kind of plot without modifying the target responder_6 at all.</p>",
      "rawMarkdown": "Okay, that's basically what I'm seeing. I've never had a predicted vs. actuals plot look so flat 😆 So my takeaway is it may not be a bug in my code, but just a poor model that's not able to learn anything? Looks like it's always predicting between -1 and 1. Almost seems like it could be a bug with the target being normalized and then not de-normalized. But I'm getting this kind of plot without modifying the target responder_6 at all.",
      "votes": null
    },
    {
      "id": "3073026",
      "postDate": "12/16/2024 00:13:25",
      "content": "<p>Thank you, this is fascinating. I'm used to dealing with classification task imbalances. I've never really worked with a regression target imbalance. I wonder, do similar techniques like over-sampling, under-sampling, different weights, apply for regression models?</p>",
      "rawMarkdown": "Thank you, this is fascinating. I'm used to dealing with classification task imbalances. I've never really worked with a regression target imbalance. I wonder, do similar techniques like over-sampling, under-sampling, different weights, apply for regression models?",
      "votes": null
    },
    {
      "id": "3073275",
      "postDate": "12/16/2024 08:54:43",
      "content": "<p>I was asking me exaclty the same question. I did some research powered by AI in the weekend and found that sampling, despite unusual, could be applied. Different weights can be applied also. Lastly, it suggested me this <a href=\"https://towardsdatascience.com/strategies-and-tactics-for-regression-on-imbalanced-data-61eeb0921fca\" target=\"_blank\">article</a>. I am experimenting with it. For now, it has the same result as training without it.</p>",
      "rawMarkdown": "I was asking me exaclty the same question. I did some research powered by AI in the weekend and found that sampling, despite unusual, could be applied. Different weights can be applied also. Lastly, it suggested me this [article](https://towardsdatascience.com/strategies-and-tactics-for-regression-on-imbalanced-data-61eeb0921fca). I am experimenting with it. For now, it has the same result as training without it.",
      "votes": null
    },
    {
      "id": "3073285",
      "postDate": "12/16/2024 09:26:57",
      "content": "<p>Yes, same techniques dealing with imbalance in classifications can also be applied to regression here. For example, you can bin the targets into ordered classes and resample from each class to get a more balanced dataset. But the trap here is that you have changed the inherent distribution of the data when doing so. It is worth to think if we really want that. My experiment in the beginning of the competition did not give me good results :(</p>\n<p>BTW Thanks <a href=\"https://www.kaggle.com/serjhenrique\" target=\"_blank\">@serjhenrique</a> for the article. It is a great source. I found it too in my early explorations but totally forgot it xD. </p>",
      "rawMarkdown": "Yes, same techniques dealing with imbalance in classifications can also be applied to regression here. For example, you can bin the targets into ordered classes and resample from each class to get a more balanced dataset. But the trap here is that you have changed the inherent distribution of the data when doing so. It is worth to think if we really want that. My experiment in the beginning of the competition did not give me good results :(\n\nBTW Thanks @serjhenrique for the article. It is a great source. I found it too in my early explorations but totally forgot it xD.",
      "votes": null
    },
    {
      "id": "3073653",
      "postDate": "12/16/2024 17:41:04",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16529018%2F9e7b88bd0027b993c52f291177aa8aa4%2F2024-12-17%20013633.png?generation=1734370732783421&amp;alt=media\" alt=\"\"></p>\n<p>A xgb model trained with 800 date, cv 0.0080 (valid in the last 120 dates)<br>\nAs we can see, the prediction values are mostly in -1 ~ 1, and it is because most target data are between -1 and 1.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16529018%2F9e7b88bd0027b993c52f291177aa8aa4%2F2024-12-17%20013633.png?generation=1734370732783421&alt=media)\n\nA xgb model trained with 800 date, cv 0.0080 (valid in the last 120 dates)\nAs we can see, the prediction values are mostly in -1 ~ 1, and it is because most target data are between -1 and 1.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3072448,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "12/15/2024 08:05:45",
      "content": "<p>You may take this from any public kernel and make the plot <a href=\"https://www.kaggle.com/jimbeno\" target=\"_blank\">@jimbeno</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3072451,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "12/15/2024 08:14:28",
      "content": "<p>try to plot your predictions on a different axis :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3072548,
      "author_name": "serjhenrique",
      "author_url": "",
      "post_date": "12/15/2024 11:35:39",
      "content": "<p>I made visual evaluations in my LGBM models. </p>\n<p>It fails miserably to predict the range -5 to +5.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2Fb6fe6814782903d2d40cdfd0ce8c1931%2Fbad%20model.png?generation=1734261949016183&amp;alt=media\" alt=\"\"></p>\n<p>The following chart should be a good model<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2F3dd52540975a6cf9c7e0cec55f601575%2Fgood%20model.png?generation=1734261986692256&amp;alt=media\" alt=\"\"></p>\n<p>I can see why it fails to predict the entire range. We can see that most values are in a narrow range between +1 and -1.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2F669b2ed66ca4e7bc25d3c3e656007f6d%2Fdensity.png?generation=1734262132420656&amp;alt=media\" alt=\"\"></p>\n<p>I think that even the top leaderboards solutions aren't doing a really good job predicting the responder. The maximum score is 1 and top leaderboards are 0.01x. I think their predicted x actual plot should be closer to my bad model than to the good one.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3072668,
          "author_name": "shiyili",
          "author_url": "",
          "post_date": "12/15/2024 14:18:08",
          "content": "<p>The model doesn’t “dare” to predict large values. This is due to the extreme long tailed target distribution. You can design a weighted loss to encourage the model to predict larger values, similar to what is done in a classification task using the Focal Loss. But this implies that you have put some prior knowledge or assumptions on the hidden test data distributions. <br>\nNice plot and visuals!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3073026,
          "author_name": "jimbeno",
          "author_url": "",
          "post_date": "12/16/2024 00:13:25",
          "content": "<p>Thank you, this is fascinating. I'm used to dealing with classification task imbalances. I've never really worked with a regression target imbalance. I wonder, do similar techniques like over-sampling, under-sampling, different weights, apply for regression models?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3073275,
              "author_name": "serjhenrique",
              "author_url": "",
              "post_date": "12/16/2024 08:54:43",
              "content": "<p>I was asking me exaclty the same question. I did some research powered by AI in the weekend and found that sampling, despite unusual, could be applied. Different weights can be applied also. Lastly, it suggested me this <a href=\"https://towardsdatascience.com/strategies-and-tactics-for-regression-on-imbalanced-data-61eeb0921fca\" target=\"_blank\">article</a>. I am experimenting with it. For now, it has the same result as training without it.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3073285,
              "author_name": "shiyili",
              "author_url": "",
              "post_date": "12/16/2024 09:26:57",
              "content": "<p>Yes, same techniques dealing with imbalance in classifications can also be applied to regression here. For example, you can bin the targets into ordered classes and resample from each class to get a more balanced dataset. But the trap here is that you have changed the inherent distribution of the data when doing so. It is worth to think if we really want that. My experiment in the beginning of the competition did not give me good results :(</p>\n<p>BTW Thanks <a href=\"https://www.kaggle.com/serjhenrique\" target=\"_blank\">@serjhenrique</a> for the article. It is a great source. I found it too in my early explorations but totally forgot it xD. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3072954,
      "author_name": "bradywynn",
      "author_url": "",
      "post_date": "12/15/2024 20:37:38",
      "content": "<p>This is from a super simple mlp evaluated on symbol_id 1</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11063283%2Fbee7e11b073c7e442b32c4fd0416c48a%2Fscatter.png?generation=1734294957546393&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 3073025,
          "author_name": "jimbeno",
          "author_url": "",
          "post_date": "12/16/2024 00:10:01",
          "content": "<p>Okay, that's basically what I'm seeing. I've never had a predicted vs. actuals plot look so flat 😆 So my takeaway is it may not be a bug in my code, but just a poor model that's not able to learn anything? Looks like it's always predicting between -1 and 1. Almost seems like it could be a bug with the target being normalized and then not de-normalized. But I'm getting this kind of plot without modifying the target responder_6 at all.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3073014,
      "author_name": "michaeltimbs",
      "author_url": "",
      "post_date": "12/15/2024 23:16:56",
      "content": "<p>Every time I have tried to update my loss function to stop it being \"safe\" and keeping close to 0 I ended up getting worse results. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3073653,
      "author_name": "i2nfinit3y",
      "author_url": "",
      "post_date": "12/16/2024 17:41:04",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16529018%2F9e7b88bd0027b993c52f291177aa8aa4%2F2024-12-17%20013633.png?generation=1734370732783421&amp;alt=media\" alt=\"\"></p>\n<p>A xgb model trained with 800 date, cv 0.0080 (valid in the last 120 dates)<br>\nAs we can see, the prediction values are mostly in -1 ~ 1, and it is because most target data are between -1 and 1.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3072445": "Does anyone have predicted vs. actuals plots they can share? Even for an average or basic model? I'm trying to figure out if my current approach is just poor or has a bug. My plot is basically a flatline.",
    "3072448": "You may take this from any public kernel and make the plot @jimbeno",
    "3072451": "try to plot your predictions on a different axis :)",
    "3072548": "I made visual evaluations in my LGBM models. \n\nIt fails miserably to predict the range -5 to +5.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2Fb6fe6814782903d2d40cdfd0ce8c1931%2Fbad%20model.png?generation=1734261949016183&alt=media)\n\nThe following chart should be a good model\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2F3dd52540975a6cf9c7e0cec55f601575%2Fgood%20model.png?generation=1734261986692256&alt=media)\n\nI can see why it fails to predict the entire range. We can see that most values are in a narrow range between +1 and -1.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2F669b2ed66ca4e7bc25d3c3e656007f6d%2Fdensity.png?generation=1734262132420656&alt=media)\n\nI think that even the top leaderboards solutions aren't doing a really good job predicting the responder. The maximum score is 1 and top leaderboards are 0.01x. I think their predicted x actual plot should be closer to my bad model than to the good one.",
    "3072668": "The model doesn’t “dare” to predict large values. This is due to the extreme long tailed target distribution. You can design a weighted loss to encourage the model to predict larger values, similar to what is done in a classification task using the Focal Loss. But this implies that you have put some prior knowledge or assumptions on the hidden test data distributions. \nNice plot and visuals!",
    "3072954": "This is from a super simple mlp evaluated on symbol_id 1\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11063283%2Fbee7e11b073c7e442b32c4fd0416c48a%2Fscatter.png?generation=1734294957546393&alt=media)",
    "3073014": "Every time I have tried to update my loss function to stop it being \"safe\" and keeping close to 0 I ended up getting worse results.",
    "3073025": "Okay, that's basically what I'm seeing. I've never had a predicted vs. actuals plot look so flat 😆 So my takeaway is it may not be a bug in my code, but just a poor model that's not able to learn anything? Looks like it's always predicting between -1 and 1. Almost seems like it could be a bug with the target being normalized and then not de-normalized. But I'm getting this kind of plot without modifying the target responder_6 at all.",
    "3073026": "Thank you, this is fascinating. I'm used to dealing with classification task imbalances. I've never really worked with a regression target imbalance. I wonder, do similar techniques like over-sampling, under-sampling, different weights, apply for regression models?",
    "3073275": "I was asking me exaclty the same question. I did some research powered by AI in the weekend and found that sampling, despite unusual, could be applied. Different weights can be applied also. Lastly, it suggested me this [article](https://towardsdatascience.com/strategies-and-tactics-for-regression-on-imbalanced-data-61eeb0921fca). I am experimenting with it. For now, it has the same result as training without it.",
    "3073285": "Yes, same techniques dealing with imbalance in classifications can also be applied to regression here. For example, you can bin the targets into ordered classes and resample from each class to get a more balanced dataset. But the trap here is that you have changed the inherent distribution of the data when doing so. It is worth to think if we really want that. My experiment in the beginning of the competition did not give me good results :(\n\nBTW Thanks @serjhenrique for the article. It is a great source. I found it too in my early explorations but totally forgot it xD.",
    "3073653": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16529018%2F9e7b88bd0027b993c52f291177aa8aa4%2F2024-12-17%20013633.png?generation=1734370732783421&alt=media)\n\nA xgb model trained with 800 date, cv 0.0080 (valid in the last 120 dates)\nAs we can see, the prediction values are mostly in -1 ~ 1, and it is because most target data are between -1 and 1."
  },
  "source": "meta"
}