{
  "id": 189855,
  "title": "Challenge of Feature Engineering on batchwise test data",
  "url": "/competitions/riiid-test-answer-prediction/discussion/189855",
  "author_name": "Abhishek Bhat",
  "post_date": "2020-10-09T03:39:54.418000",
  "votes": 13,
  "comment_count": 0,
  "views": 0,
  "content": "<p>In this competition feature engineering based on user's historical data will play a crucial role. But since the test data will be available only in batches it is important to understand whether we will have all the historical data for a particular user in the same batch?</p>\n<p>Lets assume this was true and 10 questions of a particular user are present in the same batch. Since all questions are displayed in order:</p>\n<ol>\n<li>For predicting the user response for 1st question we have the data for 9 questions from the future.</li>\n<li>For predicting the users response for 2nd question we will have one historical data point and 8 future data points.<br>\nand so on..</li>\n</ol>\n<p>The sole purpose of the test API format is to not give access to the future data points while predicting the response of a question. So for a particular user having more than one question per batch will not make sense. </p>\n<p>Is this understanding correct?</p>\n<p>This brings us to the real challenge of feature engineering as we might not have access to all historical datapoints for a particular user in the test data. Any thoughts?</p>",
  "messages": [
    {
      "id": 1043512,
      "postDate": "2020-10-09T03:39:54.417Z",
      "content": "<p>In this competition feature engineering based on user's historical data will play a crucial role. But since the test data will be available only in batches it is important to understand whether we will have all the historical data for a particular user in the same batch?</p>\n<p>Lets assume this was true and 10 questions of a particular user are present in the same batch. Since all questions are displayed in order:</p>\n<ol>\n<li>For predicting the user response for 1st question we have the data for 9 questions from the future.</li>\n<li>For predicting the users response for 2nd question we will have one historical data point and 8 future data points.<br>\nand so on..</li>\n</ol>\n<p>The sole purpose of the test API format is to not give access to the future data points while predicting the response of a question. So for a particular user having more than one question per batch will not make sense. </p>\n<p>Is this understanding correct?</p>\n<p>This brings us to the real challenge of feature engineering as we might not have access to all historical datapoints for a particular user in the test data. Any thoughts?</p>",
      "rawMarkdown": "In this competition feature engineering based on user's historical data will play a crucial role. But since the test data will be available only in batches it is important to understand whether we will have all the historical data for a particular user in the same batch?\n\nLets assume this was true and 10 questions of a particular user are present in the same batch. Since all questions are displayed in order:\n1. For predicting the user response for 1st question we have the data for 9 questions from the future.\n2. For predicting the users response for 2nd question we will have one historical data point and 8 future data points.\nand so on..\n\nThe sole purpose of the test API format is to not give access to the future data points while predicting the response of a question. So for a particular user having more than one question per batch will not make sense. \n\nIs this understanding correct?\n\nThis brings us to the real challenge of feature engineering as we might not have access to all historical datapoints for a particular user in the test data. Any thoughts?\n",
      "votes": 13
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1043512": "In this competition feature engineering based on user's historical data will play a crucial role. But since the test data will be available only in batches it is important to understand whether we will have all the historical data for a particular user in the same batch?\n\nLets assume this was true and 10 questions of a particular user are present in the same batch. Since all questions are displayed in order:\n1. For predicting the user response for 1st question we have the data for 9 questions from the future.\n2. For predicting the users response for 2nd question we will have one historical data point and 8 future data points.\nand so on..\n\nThe sole purpose of the test API format is to not give access to the future data points while predicting the response of a question. So for a particular user having more than one question per batch will not make sense. \n\nIs this understanding correct?\n\nThis brings us to the real challenge of feature engineering as we might not have access to all historical datapoints for a particular user in the test data. Any thoughts?\n"
  }
}