{
  "id": 556082,
  "title": "Responder 6 predict good after you know it is positive or negative",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/556082",
  "author_name": "",
  "post_date": "2025-01-11T05:25:42.256345200Z",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I found that if you divide data by responder 6 positive or negative, the trained model on each data set(positive, negative) can predict pretty well. A simple Xgboost can be predicted with a weighted R^2 to 0.5.  Anyone knows the reason?</p>",
  "messages": [
    {
      "id": "3093557",
      "postDate": "01/11/2025 05:25:42",
      "content": "<p>I found that if you divide data by responder 6 positive or negative, the trained model on each data set(positive, negative) can predict pretty well. A simple Xgboost can be predicted with a weighted R^2 to 0.5.  Anyone knows the reason?</p>",
      "rawMarkdown": "I found that if you divide data by responder 6 positive or negative, the trained model on each data set(positive, negative) can predict pretty well. A simple Xgboost can be predicted with a weighted R^2 to 0.5.  Anyone knows the reason?",
      "votes": null
    },
    {
      "id": "3093560",
      "postDate": "01/11/2025 05:30:34",
      "content": "<p>It's easy, just leakage. </p>",
      "rawMarkdown": "It's easy, just leakage.",
      "votes": null
    },
    {
      "id": "3093600",
      "postDate": "01/11/2025 07:33:04",
      "content": "<p>The train and test set are disjoint, why is this data leakage?</p>",
      "rawMarkdown": "The train and test set are disjoint, why is this data leakage?",
      "votes": null
    },
    {
      "id": "3093606",
      "postDate": "01/11/2025 07:44:19",
      "content": "<p>a classifier (based on target sign) will be as bad as the regressor - if you build a model that's capable of finding out responder_6' sign accurately then you've solved the problem …</p>",
      "rawMarkdown": "a classifier (based on target sign) will be as bad as the regressor - if you build a model that's capable of finding out responder_6' sign accurately then you've solved the problem ...",
      "votes": null
    },
    {
      "id": "3093609",
      "postDate": "01/11/2025 07:51:12",
      "content": "<p>Yeah. I tried to use the classifier and the accuracy is near 0.6… The result is not good enough to predict. </p>",
      "rawMarkdown": "Yeah. I tried to use the classifier and the accuracy is near 0.6... The result is not good enough to predict.",
      "votes": null
    },
    {
      "id": "3093679",
      "postDate": "01/11/2025 08:59:16",
      "content": "<p>R2 scores compare against all predict zero baseline. When you split the data, data’s mean changes. Then the zero baseline gets worse and your model gets relatively better. </p>",
      "rawMarkdown": "R2 scores compare against all predict zero baseline. When you split the data, data’s mean changes. Then the zero baseline gets worse and your model gets relatively better.",
      "votes": null
    },
    {
      "id": "3093699",
      "postDate": "01/11/2025 09:09:48",
      "content": "<p>Oh. I see. This makes sense. Thanks!</p>",
      "rawMarkdown": "Oh. I see. This makes sense. Thanks!",
      "votes": null
    },
    {
      "id": "3094334",
      "postDate": "01/12/2025 00:29:16",
      "content": "<p>if you predicted a constant .1 and -.1 for positive and negative splits, u would probably get a high r2. the model is probably just doing that.</p>",
      "rawMarkdown": "if you predicted a constant .1 and -.1 for positive and negative splits, u would probably get a high r2. the model is probably just doing that.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3093560,
      "author_name": "ray6767",
      "author_url": "",
      "post_date": "01/11/2025 05:30:34",
      "content": "<p>It's easy, just leakage. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3093600,
          "author_name": "zxmath",
          "author_url": "",
          "post_date": "01/11/2025 07:33:04",
          "content": "<p>The train and test set are disjoint, why is this data leakage?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3093606,
              "author_name": "sritichaimae",
              "author_url": "",
              "post_date": "01/11/2025 07:44:19",
              "content": "<p>a classifier (based on target sign) will be as bad as the regressor - if you build a model that's capable of finding out responder_6' sign accurately then you've solved the problem …</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3093609,
                  "author_name": "zxmath",
                  "author_url": "",
                  "post_date": "01/11/2025 07:51:12",
                  "content": "<p>Yeah. I tried to use the classifier and the accuracy is near 0.6… The result is not good enough to predict. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3093679,
      "author_name": "aerdem4",
      "author_url": "",
      "post_date": "01/11/2025 08:59:16",
      "content": "<p>R2 scores compare against all predict zero baseline. When you split the data, data’s mean changes. Then the zero baseline gets worse and your model gets relatively better. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3093699,
          "author_name": "zxmath",
          "author_url": "",
          "post_date": "01/11/2025 09:09:48",
          "content": "<p>Oh. I see. This makes sense. Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3094334,
      "author_name": "rulangu",
      "author_url": "",
      "post_date": "01/12/2025 00:29:16",
      "content": "<p>if you predicted a constant .1 and -.1 for positive and negative splits, u would probably get a high r2. the model is probably just doing that.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3093557": "I found that if you divide data by responder 6 positive or negative, the trained model on each data set(positive, negative) can predict pretty well. A simple Xgboost can be predicted with a weighted R^2 to 0.5.  Anyone knows the reason?",
    "3093560": "It's easy, just leakage.",
    "3093600": "The train and test set are disjoint, why is this data leakage?",
    "3093606": "a classifier (based on target sign) will be as bad as the regressor - if you build a model that's capable of finding out responder_6' sign accurately then you've solved the problem ...",
    "3093609": "Yeah. I tried to use the classifier and the accuracy is near 0.6... The result is not good enough to predict.",
    "3093679": "R2 scores compare against all predict zero baseline. When you split the data, data’s mean changes. Then the zero baseline gets worse and your model gets relatively better.",
    "3093699": "Oh. I see. This makes sense. Thanks!",
    "3094334": "if you predicted a constant .1 and -.1 for positive and negative splits, u would probably get a high r2. the model is probably just doing that."
  },
  "source": "meta"
}