{
  "id": 190263,
  "title": "Question about hidden test set",
  "url": "/competitions/riiid-test-answer-prediction/discussion/190263",
  "author_name": "Ning Jia",
  "post_date": "2020-10-10T22:30:58.105000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<blockquote>\n  <p>Some questions will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling newly introduced questions. Their metadata is still in question.csv as usual. </p>\n</blockquote>\n<p>My understanding is those questions in hidden test set will be in question.csv only but not in the train set.  But my data sanity check shows all the 13523 questions in question.csv have been shown in the train set. What is the hidden test set then?</p>\n<p>Basically, <code>train_df[(train_df['content_type_id']==0)]['content_id'].unique()</code> and <code>question_df['question_id'].unique()</code> are same set.</p>",
  "messages": [
    {
      "id": 1045684,
      "postDate": "2020-10-10T22:30:58.107Z",
      "content": "<blockquote>\n  <p>Some questions will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling newly introduced questions. Their metadata is still in question.csv as usual. </p>\n</blockquote>\n<p>My understanding is those questions in hidden test set will be in question.csv only but not in the train set.  But my data sanity check shows all the 13523 questions in question.csv have been shown in the train set. What is the hidden test set then?</p>\n<p>Basically, <code>train_df[(train_df['content_type_id']==0)]['content_id'].unique()</code> and <code>question_df['question_id'].unique()</code> are same set.</p>",
      "rawMarkdown": ">  Some questions will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling newly introduced questions. Their metadata is still in question.csv as usual. \n\nMy understanding is those questions in hidden test set will be in question.csv only but not in the train set.  But my data sanity check shows all the 13523 questions in question.csv have been shown in the train set. What is the hidden test set then?\n\nBasically, `train_df[(train_df['content_type_id']==0)]['content_id'].unique()` and `question_df['question_id'].unique()` are same set.\n\n",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1045684": ">  Some questions will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling newly introduced questions. Their metadata is still in question.csv as usual. \n\nMy understanding is those questions in hidden test set will be in question.csv only but not in the train set.  But my data sanity check shows all the 13523 questions in question.csv have been shown in the train set. What is the hidden test set then?\n\nBasically, `train_df[(train_df['content_type_id']==0)]['content_id'].unique()` and `question_df['question_id'].unique()` are same set.\n\n"
  }
}