{
  "id": 550678,
  "title": "Question about Using Other Responders as Features",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/550678",
  "author_name": "",
  "post_date": "2024-12-09T00:21:56.237427300Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello everyone,  </p>\n<p>I have a question regarding the use of the other responders (<code>responder_{0...5}</code> and <code>responder_{7...8}</code>) in this competition. Specifically, is it allowed to use these responders as input features for training my model to predict <code>responder_6</code>?  </p>\n<p>From the dataset description, I see that:  </p>\n<ol>\n<li>The training data includes all responders (<code>responder_{0...8}</code>), which are clipped between -5 and 5.  </li>\n<li>The file <code>lags.parquet</code> provides lagged values for all responders for a given <code>date_id</code>.  </li>\n</ol>\n<p>This suggests that these responders might be usable as additional features, particularly when creating lagged variables. However, since <code>responder_6</code> is the only target for scoring, I want to confirm if using the other responders during training aligns with the competition's guidelines. I’m concerned about the risk of overfitting if the other responders are highly correlated with <code>responder_6</code>.<br>\nThank you in advance!  </p>",
  "messages": [
    {
      "id": "3067133",
      "postDate": "12/09/2024 00:21:56",
      "content": "<p>Hello everyone,  </p>\n<p>I have a question regarding the use of the other responders (<code>responder_{0...5}</code> and <code>responder_{7...8}</code>) in this competition. Specifically, is it allowed to use these responders as input features for training my model to predict <code>responder_6</code>?  </p>\n<p>From the dataset description, I see that:  </p>\n<ol>\n<li>The training data includes all responders (<code>responder_{0...8}</code>), which are clipped between -5 and 5.  </li>\n<li>The file <code>lags.parquet</code> provides lagged values for all responders for a given <code>date_id</code>.  </li>\n</ol>\n<p>This suggests that these responders might be usable as additional features, particularly when creating lagged variables. However, since <code>responder_6</code> is the only target for scoring, I want to confirm if using the other responders during training aligns with the competition's guidelines. I’m concerned about the risk of overfitting if the other responders are highly correlated with <code>responder_6</code>.<br>\nThank you in advance!  </p>",
      "rawMarkdown": "Hello everyone,  \n\nI have a question regarding the use of the other responders (`responder_{0...5}` and `responder_{7...8}`) in this competition. Specifically, is it allowed to use these responders as input features for training my model to predict `responder_6`?  \n\nFrom the dataset description, I see that:  \n1. The training data includes all responders (`responder_{0...8}`), which are clipped between -5 and 5.  \n2. The file `lags.parquet` provides lagged values for all responders for a given `date_id`.  \n\nThis suggests that these responders might be usable as additional features, particularly when creating lagged variables. However, since `responder_6` is the only target for scoring, I want to confirm if using the other responders during training aligns with the competition's guidelines. I’m concerned about the risk of overfitting if the other responders are highly correlated with `responder_6`.\nThank you in advance!",
      "votes": null
    },
    {
      "id": "3067275",
      "postDate": "12/09/2024 06:00:11",
      "content": "<p>You can but obviously you have to use <em>predicted</em> responder values as they aren't passed in at test time.</p>\n<p>I had some success with offline training in predicting the other responders and feeding in those predictions to a main model to predict responder 6. The issue I faced is that on the holdout test set it made my models perform worse. I suspect the errors magnify due to non-stationarity of data. I'm not sure if there will be time to do online training for multiple auxiliary models but you can definitely try it (I might if I have time)</p>",
      "rawMarkdown": "You can but obviously you have to use *predicted* responder values as they aren't passed in at test time.\n\nI had some success with offline training in predicting the other responders and feeding in those predictions to a main model to predict responder 6. The issue I faced is that on the holdout test set it made my models perform worse. I suspect the errors magnify due to non-stationarity of data. I'm not sure if there will be time to do online training for multiple auxiliary models but you can definitely try it (I might if I have time)",
      "votes": null
    },
    {
      "id": "3094079",
      "postDate": "01/11/2025 16:23:23",
      "content": "<p>Thank you🙏</p>",
      "rawMarkdown": "Thank you🙏",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3067275,
      "author_name": "michaeltimbs",
      "author_url": "",
      "post_date": "12/09/2024 06:00:11",
      "content": "<p>You can but obviously you have to use <em>predicted</em> responder values as they aren't passed in at test time.</p>\n<p>I had some success with offline training in predicting the other responders and feeding in those predictions to a main model to predict responder 6. The issue I faced is that on the holdout test set it made my models perform worse. I suspect the errors magnify due to non-stationarity of data. I'm not sure if there will be time to do online training for multiple auxiliary models but you can definitely try it (I might if I have time)</p>",
      "votes": null,
      "replies": [
        {
          "id": 3094079,
          "author_name": "lordsantiman",
          "author_url": "",
          "post_date": "01/11/2025 16:23:23",
          "content": "<p>Thank you🙏</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3067133": "Hello everyone,  \n\nI have a question regarding the use of the other responders (`responder_{0...5}` and `responder_{7...8}`) in this competition. Specifically, is it allowed to use these responders as input features for training my model to predict `responder_6`?  \n\nFrom the dataset description, I see that:  \n1. The training data includes all responders (`responder_{0...8}`), which are clipped between -5 and 5.  \n2. The file `lags.parquet` provides lagged values for all responders for a given `date_id`.  \n\nThis suggests that these responders might be usable as additional features, particularly when creating lagged variables. However, since `responder_6` is the only target for scoring, I want to confirm if using the other responders during training aligns with the competition's guidelines. I’m concerned about the risk of overfitting if the other responders are highly correlated with `responder_6`.\nThank you in advance!",
    "3067275": "You can but obviously you have to use *predicted* responder values as they aren't passed in at test time.\n\nI had some success with offline training in predicting the other responders and feeding in those predictions to a main model to predict responder 6. The issue I faced is that on the holdout test set it made my models perform worse. I suspect the errors magnify due to non-stationarity of data. I'm not sure if there will be time to do online training for multiple auxiliary models but you can definitely try it (I might if I have time)",
    "3094079": "Thank you🙏"
  },
  "source": "meta"
}