{
  "id": 386295,
  "title": "Catch 22 in test data?",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/386295",
  "author_name": "",
  "post_date": "2023-02-12T10:42:40.244224400Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>In order to come up with a solid CV strategy and understanding of submission I would like to understand the presence or absence of questions 19-22 in the test data.</p>\n<p>Nowhere in the data description are questions 19-22 ruled out from test data. The only limitation in terms of description is that the train data only includes questions 1-18.</p>\n<p>In train session data the levels 19-22 are present.</p>\n<p>Given that CV between public train results and LB are quite nice this leaves two hypothesis:<br>\n1) No questions 19-22 in test data<br>\nor …more likely?<br>\n2) questions 19-22 in hidden test data</p>\n<p>Scenario 2 leads to a distribution shift, unseen fqids and potentially <br>\nmissing submission code logic (when not including q19-22 processing)</p>\n<p>What do you think how this is set up? How do you understand the data description?</p>",
  "messages": [
    {
      "id": "2140980",
      "postDate": "02/12/2023 10:42:40",
      "content": "<p>In order to come up with a solid CV strategy and understanding of submission I would like to understand the presence or absence of questions 19-22 in the test data.</p>\n<p>Nowhere in the data description are questions 19-22 ruled out from test data. The only limitation in terms of description is that the train data only includes questions 1-18.</p>\n<p>In train session data the levels 19-22 are present.</p>\n<p>Given that CV between public train results and LB are quite nice this leaves two hypothesis:<br>\n1) No questions 19-22 in test data<br>\nor …more likely?<br>\n2) questions 19-22 in hidden test data</p>\n<p>Scenario 2 leads to a distribution shift, unseen fqids and potentially <br>\nmissing submission code logic (when not including q19-22 processing)</p>\n<p>What do you think how this is set up? How do you understand the data description?</p>",
      "rawMarkdown": "In order to come up with a solid CV strategy and understanding of submission I would like to understand the presence or absence of questions 19-22 in the test data.\n\nNowhere in the data description are questions 19-22 ruled out from test data. The only limitation in terms of description is that the train data only includes questions 1-18.\n\nIn train session data the levels 19-22 are present.\n\nGiven that CV between public train results and LB are quite nice this leaves two hypothesis:\n1) No questions 19-22 in test data\nor ...more likely?\n2) questions 19-22 in hidden test data\n\nScenario 2 leads to a distribution shift, unseen fqids and potentially \nmissing submission code logic (when not including q19-22 processing)\n\nWhat do you think how this is set up? How do you understand the data description?",
      "votes": null
    },
    {
      "id": "2141346",
      "postDate": "02/12/2023 17:14:47",
      "content": "<p>We only need to predict correctness for questions 1 thru 18 (for both public and private LB). A description of the game is <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384796\" target=\"_blank\">here</a>. I played the game but I forget if questions 19 thru 22 exist. But regardless if they exist, we do not need to predict them.</p>",
      "rawMarkdown": "We only need to predict correctness for questions 1 thru 18 (for both public and private LB). A description of the game is [here][1]. I played the game but I forget if questions 19 thru 22 exist. But regardless if they exist, we do not need to predict them.\n\n[1]: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384796",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2141346,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/12/2023 17:14:47",
      "content": "<p>We only need to predict correctness for questions 1 thru 18 (for both public and private LB). A description of the game is <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384796\" target=\"_blank\">here</a>. I played the game but I forget if questions 19 thru 22 exist. But regardless if they exist, we do not need to predict them.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2140980": "In order to come up with a solid CV strategy and understanding of submission I would like to understand the presence or absence of questions 19-22 in the test data.\n\nNowhere in the data description are questions 19-22 ruled out from test data. The only limitation in terms of description is that the train data only includes questions 1-18.\n\nIn train session data the levels 19-22 are present.\n\nGiven that CV between public train results and LB are quite nice this leaves two hypothesis:\n1) No questions 19-22 in test data\nor ...more likely?\n2) questions 19-22 in hidden test data\n\nScenario 2 leads to a distribution shift, unseen fqids and potentially \nmissing submission code logic (when not including q19-22 processing)\n\nWhat do you think how this is set up? How do you understand the data description?",
    "2141346": "We only need to predict correctness for questions 1 thru 18 (for both public and private LB). A description of the game is [here][1]. I played the game but I forget if questions 19 thru 22 exist. But regardless if they exist, we do not need to predict them.\n\n[1]: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384796"
  },
  "source": "meta"
}