{
  "id": 190415,
  "title": "Will there *really* be new questions in private test set?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/190415",
  "author_name": "",
  "post_date": "2020-10-11T17:45:38.187605900Z",
  "votes": 6,
  "comment_count": 9,
  "views": 0,
  "content": "<p>The description says</p>\n<blockquote>\n  <p>Some questions will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling newly introduced questions. Their metadata is still in question.csv as usual.</p>\n</blockquote>\n<p>This says that test set will have new questions, but their metadata will be in <code>questions.csv</code><br>\nBut I can see that all the questions in <code>questions.csv</code> are also in <code>train.csv</code></p>\n<pre><code>q_not_in_train_data = set(questions['question_id']) - set(train.loc[train['content_type_id'] == 0, 'content_id'])\nprint(f'Questions in Metadata, but not in train data - {q_not_in_train_data}')\n\nq_not_in_metadata = set(train.loc[train['content_type_id'] == 0, 'content_id']) - set(questions['question_id'])\nprint(f'Questions in train, but not in metadata - {q_not_in_metadata}')\n</code></pre>\n<blockquote>\n  <p>Questions in Metadata, but not in train data - set()<br>\n  Questions in train, but not in metadata - set()</p>\n</blockquote>\n<p>Am I missing something obvious?</p>",
  "messages": [
    {
      "id": "1046482",
      "postDate": "10/11/2020 17:45:38",
      "content": "<p>The description says</p>\n<blockquote>\n  <p>Some questions will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling newly introduced questions. Their metadata is still in question.csv as usual.</p>\n</blockquote>\n<p>This says that test set will have new questions, but their metadata will be in <code>questions.csv</code><br>\nBut I can see that all the questions in <code>questions.csv</code> are also in <code>train.csv</code></p>\n<pre><code>q_not_in_train_data = set(questions['question_id']) - set(train.loc[train['content_type_id'] == 0, 'content_id'])\nprint(f'Questions in Metadata, but not in train data - {q_not_in_train_data}')\n\nq_not_in_metadata = set(train.loc[train['content_type_id'] == 0, 'content_id']) - set(questions['question_id'])\nprint(f'Questions in train, but not in metadata - {q_not_in_metadata}')\n</code></pre>\n<blockquote>\n  <p>Questions in Metadata, but not in train data - set()<br>\n  Questions in train, but not in metadata - set()</p>\n</blockquote>\n<p>Am I missing something obvious?</p>",
      "rawMarkdown": "The description says\n\n> Some questions will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling newly introduced questions. Their metadata is still in question.csv as usual.\n\nThis says that test set will have new questions, but their metadata will be in `questions.csv`\nBut I can see that all the questions in `questions.csv` are also in `train.csv`\n\n```\nq_not_in_train_data = set(questions['question_id']) - set(train.loc[train['content_type_id'] == 0, 'content_id'])\nprint(f'Questions in Metadata, but not in train data - {q_not_in_train_data}')\n\nq_not_in_metadata = set(train.loc[train['content_type_id'] == 0, 'content_id']) - set(questions['question_id'])\nprint(f'Questions in train, but not in metadata - {q_not_in_metadata}')\n```\n> Questions in Metadata, but not in train data - set()\nQuestions in train, but not in metadata - set()\n\nAm I missing something obvious?",
      "votes": null
    },
    {
      "id": "1046515",
      "postDate": "10/11/2020 18:19:46",
      "content": "<p>You check questions.csv only in train set.The test set may contain a different file (imho)</p>",
      "rawMarkdown": "You check questions.csv only in train set.The test set may contain a different file (imho)",
      "votes": null
    },
    {
      "id": "1046550",
      "postDate": "10/11/2020 19:05:18",
      "content": "<p>This seems to be the only possibility by the process of elimination.<br>\nIf this is the case it should be stated clearly</p>",
      "rawMarkdown": "This seems to be the only possibility by the process of elimination.\nIf this is the case it should be stated clearly",
      "votes": null
    },
    {
      "id": "1046663",
      "postDate": "10/11/2020 21:45:51",
      "content": "<p><a href=\"https://www.kaggle.com/sapr3s\" target=\"_blank\">@sapr3s</a> I checked, the question.csv, when we submit to run, has the same set of question ids as the question.csv when we commit.</p>",
      "rawMarkdown": "sapr3s I checked, the question.csv, when we submit to run, has the same set of question ids as the question.csv when we commit.",
      "votes": null
    },
    {
      "id": "1046670",
      "postDate": "10/11/2020 22:02:57",
      "content": "<p>I also have same question. <a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190263</a></p>\n<p>So questions.csv may change for test set? I doubt that.</p>",
      "rawMarkdown": "I also have same question. [https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190263](url)\n\nSo questions.csv may change for test set? I doubt that.",
      "votes": null
    },
    {
      "id": "1046680",
      "postDate": "10/11/2020 22:24:01",
      "content": "<p>Check this thread</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188899\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188899</a></p>",
      "rawMarkdown": "Check this thread\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188899",
      "votes": null
    },
    {
      "id": "1046930",
      "postDate": "10/12/2020 05:40:05",
      "content": "<p>So if <code>4) All question IDs queried at test time are already seen in questions.csv?</code> is true. That means all questions in the test set are in <code>train.csv</code>, which is contrary to what is written in the data description, that's why I raised this question.</p>\n<p>If you could make the kernel where you checked it public it will be very helpful.</p>",
      "rawMarkdown": "So if `4) All question IDs queried at test time are already seen in questions.csv?` is true. That means all questions in the test set are in `train.csv`, which is contrary to what is written in the data description, that's why I raised this question.\n\nIf you could make the kernel where you checked it public it will be very helpful.",
      "votes": null
    },
    {
      "id": "1047760",
      "postDate": "10/12/2020 22:43:12",
      "content": "<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190667\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190667</a></p>",
      "rawMarkdown": "https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190667",
      "votes": null
    },
    {
      "id": "1105004",
      "postDate": "12/07/2020 12:56:43",
      "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a>  what can we conclude out of this.. Statement in Data Description is contrary to the data in questions. csv ?</p>",
      "rawMarkdown": "yihdarshieh  what can we conclude out of this.. Statement in Data Description is contrary to the data in questions. csv ?",
      "votes": null
    },
    {
      "id": "1105892",
      "postDate": "12/08/2020 09:54:27",
      "content": "<p>the data description has changed a long time ago …  and now there</p>\n<blockquote>\n  <p>Some users will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling new arrivals to a website.</p>\n</blockquote>",
      "rawMarkdown": "the data description has changed a long time ago ...  and now there\n\n> Some users will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling new arrivals to a website.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1046515,
      "author_name": "sapr3s",
      "author_url": "",
      "post_date": "10/11/2020 18:19:46",
      "content": "<p>You check questions.csv only in train set.The test set may contain a different file (imho)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1046550,
          "author_name": "gautham11",
          "author_url": "",
          "post_date": "10/11/2020 19:05:18",
          "content": "<p>This seems to be the only possibility by the process of elimination.<br>\nIf this is the case it should be stated clearly</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1046663,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "10/11/2020 21:45:51",
          "content": "<p><a href=\"https://www.kaggle.com/sapr3s\" target=\"_blank\">@sapr3s</a> I checked, the question.csv, when we submit to run, has the same set of question ids as the question.csv when we commit.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1105004,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/07/2020 12:56:43",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a>  what can we conclude out of this.. Statement in Data Description is contrary to the data in questions. csv ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1105892,
          "author_name": "sapr3s",
          "author_url": "",
          "post_date": "12/08/2020 09:54:27",
          "content": "<p>the data description has changed a long time ago …  and now there</p>\n<blockquote>\n  <p>Some users will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling new arrivals to a website.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1046670,
      "author_name": "ningjiaca",
      "author_url": "",
      "post_date": "10/11/2020 22:02:57",
      "content": "<p>I also have same question. <a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190263</a></p>\n<p>So questions.csv may change for test set? I doubt that.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1046680,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "10/11/2020 22:24:01",
      "content": "<p>Check this thread</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188899\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188899</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1046930,
          "author_name": "gautham11",
          "author_url": "",
          "post_date": "10/12/2020 05:40:05",
          "content": "<p>So if <code>4) All question IDs queried at test time are already seen in questions.csv?</code> is true. That means all questions in the test set are in <code>train.csv</code>, which is contrary to what is written in the data description, that's why I raised this question.</p>\n<p>If you could make the kernel where you checked it public it will be very helpful.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1047760,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "10/12/2020 22:43:12",
          "content": "<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190667\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190667</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1046482": "The description says\n\n> Some questions will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling newly introduced questions. Their metadata is still in question.csv as usual.\n\nThis says that test set will have new questions, but their metadata will be in `questions.csv`\nBut I can see that all the questions in `questions.csv` are also in `train.csv`\n\n```\nq_not_in_train_data = set(questions['question_id']) - set(train.loc[train['content_type_id'] == 0, 'content_id'])\nprint(f'Questions in Metadata, but not in train data - {q_not_in_train_data}')\n\nq_not_in_metadata = set(train.loc[train['content_type_id'] == 0, 'content_id']) - set(questions['question_id'])\nprint(f'Questions in train, but not in metadata - {q_not_in_metadata}')\n```\n> Questions in Metadata, but not in train data - set()\nQuestions in train, but not in metadata - set()\n\nAm I missing something obvious?",
    "1046515": "You check questions.csv only in train set.The test set may contain a different file (imho)",
    "1046550": "This seems to be the only possibility by the process of elimination.\nIf this is the case it should be stated clearly",
    "1046663": "sapr3s I checked, the question.csv, when we submit to run, has the same set of question ids as the question.csv when we commit.",
    "1046670": "I also have same question. [https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190263](url)\n\nSo questions.csv may change for test set? I doubt that.",
    "1046680": "Check this thread\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188899",
    "1046930": "So if `4) All question IDs queried at test time are already seen in questions.csv?` is true. That means all questions in the test set are in `train.csv`, which is contrary to what is written in the data description, that's why I raised this question.\n\nIf you could make the kernel where you checked it public it will be very helpful.",
    "1047760": "https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190667",
    "1105004": "yihdarshieh  what can we conclude out of this.. Statement in Data Description is contrary to the data in questions. csv ?",
    "1105892": "the data description has changed a long time ago ...  and now there\n\n> Some users will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling new arrivals to a website."
  },
  "source": "meta"
}