{
  "id": 190200,
  "title": "Important questions regarding private test set",
  "url": "/competitions/riiid-test-answer-prediction/discussion/190200",
  "author_name": "",
  "post_date": "2020-10-10T16:06:58.799231400Z",
  "votes": 54,
  "comment_count": 17,
  "views": 0,
  "content": "<p>I think we the competitors are deserved to be answered some very basic but important questions about the private test set. If unanswered, it will leave frustration in model design/feature engineering. I will present 3 typical questions and explain why it is very important to be answered.</p>\n<p>Question 1: Is there any user_id in private test set that overlap with train set?</p>\n<p>Why this question is important:</p>\n<ul>\n<li>If the answer is Yes, then user_id is one important feature that can be aggregated by the average target. However, another question (question 2) will pop up regarding how to engineer this feature.</li>\n</ul>\n<p>Question 2: If the answer to question 1 is Yes (which is very probable), then are the timestamps of that user_id in private test set AFTER the last timestamp of that user_id in train set?</p>\n<p>Why this question is important:</p>\n<ul>\n<li>If the answer is Yes, then we don't have leakage. Otherwise, leakage will present, affecting the engineering pipeline.</li>\n</ul>\n<p>And the MOST IMPORTANT QUESTION is:<br>\nQuestion 3: Suppose the (real) history of user ABC has 100 rows. Suppose this user is NOT present in train set. Are all of these 100 rows present in private test set, or just some of them (for example, 20 rows)? </p>\n<p>Why this question is important:</p>\n<ul>\n<li>If the answer is No, then we cannot use full history of any user_id in the train set to build a sequential model (RNN, Transformer…). In other words, any sequential model will fail in test set because we don't have full history up to any row. And we can't know unless you tell us.</li>\n</ul>\n<p>I think these questions are not anything sensitive to not be answered. They are very basic, but important to the competition. Please give us the answers. </p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>",
  "messages": [
    {
      "id": "1045422",
      "postDate": "10/10/2020 16:06:58",
      "content": "<p>I think we the competitors are deserved to be answered some very basic but important questions about the private test set. If unanswered, it will leave frustration in model design/feature engineering. I will present 3 typical questions and explain why it is very important to be answered.</p>\n<p>Question 1: Is there any user_id in private test set that overlap with train set?</p>\n<p>Why this question is important:</p>\n<ul>\n<li>If the answer is Yes, then user_id is one important feature that can be aggregated by the average target. However, another question (question 2) will pop up regarding how to engineer this feature.</li>\n</ul>\n<p>Question 2: If the answer to question 1 is Yes (which is very probable), then are the timestamps of that user_id in private test set AFTER the last timestamp of that user_id in train set?</p>\n<p>Why this question is important:</p>\n<ul>\n<li>If the answer is Yes, then we don't have leakage. Otherwise, leakage will present, affecting the engineering pipeline.</li>\n</ul>\n<p>And the MOST IMPORTANT QUESTION is:<br>\nQuestion 3: Suppose the (real) history of user ABC has 100 rows. Suppose this user is NOT present in train set. Are all of these 100 rows present in private test set, or just some of them (for example, 20 rows)? </p>\n<p>Why this question is important:</p>\n<ul>\n<li>If the answer is No, then we cannot use full history of any user_id in the train set to build a sequential model (RNN, Transformer…). In other words, any sequential model will fail in test set because we don't have full history up to any row. And we can't know unless you tell us.</li>\n</ul>\n<p>I think these questions are not anything sensitive to not be answered. They are very basic, but important to the competition. Please give us the answers. </p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>",
      "rawMarkdown": "I think we the competitors are deserved to be answered some very basic but important questions about the private test set. If unanswered, it will leave frustration in model design/feature engineering. I will present 3 typical questions and explain why it is very important to be answered.\n\nQuestion 1: Is there any user_id in private test set that overlap with train set?\n\nWhy this question is important:\n+ If the answer is Yes, then user_id is one important feature that can be aggregated by the average target. However, another question (question 2) will pop up regarding how to engineer this feature.\n\nQuestion 2: If the answer to question 1 is Yes (which is very probable), then are the timestamps of that user_id in private test set AFTER the last timestamp of that user_id in train set?\n\nWhy this question is important:\n+ If the answer is Yes, then we don't have leakage. Otherwise, leakage will present, affecting the engineering pipeline.\n\nAnd the MOST IMPORTANT QUESTION is:\nQuestion 3: Suppose the (real) history of user ABC has 100 rows. Suppose this user is NOT present in train set. Are all of these 100 rows present in private test set, or just some of them (for example, 20 rows)? \n\nWhy this question is important:\n+ If the answer is No, then we cannot use full history of any user_id in the train set to build a sequential model (RNN, Transformer...). In other words, any sequential model will fail in test set because we don't have full history up to any row. And we can't know unless you tell us.\n\nI think these questions are not anything sensitive to not be answered. They are very basic, but important to the competition. Please give us the answers. \n\n@sohier",
      "votes": null
    },
    {
      "id": "1045521",
      "postDate": "10/10/2020 18:05:32",
      "content": "<p>I am trying to understand if any Y/N to any questions here changes your will of building a better validated model without leakage?</p>",
      "rawMarkdown": "I am trying to understand if any Y/N to any questions here changes your will of building a better validated model without leakage?",
      "votes": null
    },
    {
      "id": "1045735",
      "postDate": "10/11/2020 01:34:44",
      "content": "<p>Yes, this is important because we don't know how to build a reliable validation scheme without knowing some basic things about the private test.</p>",
      "rawMarkdown": "Yes, this is important because we don't know how to build a reliable validation scheme without knowing some basic things about the private test.",
      "votes": null
    },
    {
      "id": "1045785",
      "postDate": "10/11/2020 03:09:13",
      "content": "<p>Q1: Yes, but there are user_ids not in the train at all.<br>\nQ2: Yes.<br>\nQ3: Doesn't matter. Because even the answer is Yes, we still cannot use the full history. Please take look of the example test file, only <code>prior_question_elapsed_time</code>,<code>prior_question_had_explanation</code>,<code>prior_group_responses</code> and <code>prior_group_answers_correct</code> have some limited history info.</p>",
      "rawMarkdown": "Q1: Yes, but there are user_ids not in the train at all.\nQ2: Yes.\nQ3: Doesn't matter. Because even the answer is Yes, we still cannot use the full history. Please take look of the example test file, only `prior_question_elapsed_time`,`prior_question_had_explanation`,`prior_group_responses ` and `prior_group_answers_correct` have some limited history info.",
      "votes": null
    },
    {
      "id": "1045790",
      "postDate": "10/11/2020 03:40:38",
      "content": "<p>How do you know Q2 is Yes? And I don't understand your point in Q3. If the batches of groups in private test set are increasing and cover all history, we can accumulate full history of any user.</p>",
      "rawMarkdown": "How do you know Q2 is Yes? And I don't understand your point in Q3. If the batches of groups in private test set are increasing and cover all history, we can accumulate full history of any user.",
      "votes": null
    },
    {
      "id": "1045838",
      "postDate": "10/11/2020 05:05:23",
      "content": "<p><a href=\"https://www.kaggle.com/khahuras\" target=\"_blank\">@khahuras</a> <br>\nThere are 2 <code>user_id</code> that overlaps with train and example test daa<br>\nfor the <code>user_id</code> = <code>7792299</code>, its continuous.<br>\nAnd for the other user_id there is a huge gap in their timestamp.</p>\n<p>EDIT : I'm sorry I loaded only like 10**5 rows, that's why i only got 2 user_id overlapped with test data</p>",
      "rawMarkdown": "khahuras \nThere are 2 `user_id` that overlaps with train and example test daa\nfor the `user_id` = `7792299`, its continuous.\nAnd for the other user_id there is a huge gap in their timestamp.\n\n\nEDIT : I'm sorry I loaded only like 10**5 rows, that's why i only got 2 user_id overlapped with test data",
      "votes": null
    },
    {
      "id": "1046026",
      "postDate": "10/11/2020 09:22:52",
      "content": "<p>Plus, I can see people are creating features using group-by etc on every batch, but is that even correct? It's already visible that there can be a user_id in group_num 10 and the same user_id can appear in group_num 100. And we don't even know the order etc nor (as of now) we can club all the test set as well as we have to make a prediction for every group_num. I am not quite sue how can we override the preds that's already made and is not accessible anymore.<br>\nThoughts? Ty!</p>",
      "rawMarkdown": "Plus, I can see people are creating features using group-by etc on every batch, but is that even correct? It's already visible that there can be a user_id in group_num 10 and the same user_id can appear in group_num 100. And we don't even know the order etc nor (as of now) we can club all the test set as well as we have to make a prediction for every group_num. I am not quite sue how can we override the preds that's already made and is not accessible anymore.\nThoughts? Ty!",
      "votes": null
    },
    {
      "id": "1046430",
      "postDate": "10/11/2020 16:47:47",
      "content": "<p>can u name any difference in your validation scheme if in any case difference between yes or no to any of the questions ? <a href=\"https://www.kaggle.com/khahuras\" target=\"_blank\">@khahuras</a> </p>",
      "rawMarkdown": "can u name any difference in your validation scheme if in any case difference between yes or no to any of the questions ? @khahuras",
      "votes": null
    },
    {
      "id": "1046435",
      "postDate": "10/11/2020 16:51:42",
      "content": "<blockquote>\n  <p>How do you know Q2 is Yes? And I don't understand your point in Q3. If the batches of groups in private test set are increasing and cover all history, we can accumulate full history of any user.</p>\n</blockquote>\n<p>I cannot valid Q2 directly, that's my assumption. Because the organizer should consider to avoid basic data leakage. I think the train/test split is done either by users basis or time basis. Both cases cannot result in the data leakage you described in Q2. </p>\n<p>For Q3, do you think it's possible to accumulate the full history in the test submissions notebook?  Considering the limit of running time and memory requirement of the notebook and also the difficulties of debugging in those time series API calls, honestly I don't know for now. Maybe we can only accumulate very short history for each user.  </p>",
      "rawMarkdown": "> How do you know Q2 is Yes? And I don't understand your point in Q3. If the batches of groups in private test set are increasing and cover all history, we can accumulate full history of any user.\n\nI cannot valid Q2 directly, that's my assumption. Because the organizer should consider to avoid basic data leakage. I think the train/test split is done either by users basis or time basis. Both cases cannot result in the data leakage you described in Q2. \n\nFor Q3, do you think it's possible to accumulate the full history in the test submissions notebook?  Considering the limit of running time and memory requirement of the notebook and also the difficulties of debugging in those time series API calls, honestly I don't know for now. Maybe we can only accumulate very short history for each user.",
      "votes": null
    },
    {
      "id": "1046446",
      "postDate": "10/11/2020 17:00:48",
      "content": "<p>We can accumulate for sure but there's no assurance when will we get to see the same user again. It's less likely that you will run out of memory if you are \"purely\" making a \"test-only-submission\". Plus we have to call predict for every group_num irrespective of the same, so at any point you cannot use your accumulated data to full advantage as iterating over is coupled with env.predict. (that's my understanding as of now atleast)</p>",
      "rawMarkdown": "We can accumulate for sure but there's no assurance when will we get to see the same user again. It's less likely that you will run out of memory if you are \"purely\" making a \"test-only-submission\". Plus we have to call predict for every group_num irrespective of the same, so at any point you cannot use your accumulated data to full advantage as iterating over is coupled with env.predict. (that's my understanding as of now atleast)",
      "votes": null
    },
    {
      "id": "1048155",
      "postDate": "10/13/2020 08:38:00",
      "content": "<p>We cannot override the preds that already been made. What we can do is to accumulate test.csv up until the current processing batch.</p>",
      "rawMarkdown": "We cannot override the preds that already been made. What we can do is to accumulate test.csv up until the current processing batch.",
      "votes": null
    },
    {
      "id": "1048160",
      "postDate": "10/13/2020 08:39:56",
      "content": "<p>I agree we cannot overwrite the previous preds but we can certainly make corrections to our next set of preds for sure by using info from those 2 extra columns..</p>",
      "rawMarkdown": "I agree we cannot overwrite the previous preds but we can certainly make corrections to our next set of preds for sure by using info from those 2 extra columns..",
      "votes": null
    },
    {
      "id": "1048161",
      "postDate": "10/13/2020 08:40:17",
      "content": "<p>For example, if all rows in private test set have timestamps after all rows in train (for each user_id), then in training pipeline we should not use any future row (from the perspective of a row sample) to craft features. The validation strategy is thus should be taken carefully.</p>",
      "rawMarkdown": "For example, if all rows in private test set have timestamps after all rows in train (for each user_id), then in training pipeline we should not use any future row (from the perspective of a row sample) to craft features. The validation strategy is thus should be taken carefully.",
      "votes": null
    },
    {
      "id": "1048168",
      "postDate": "10/13/2020 08:44:14",
      "content": "<blockquote>\n  <p>For example, if all rows in private test set have timestamps after all rows in train (for each user_id), then in training pipeline we should not use any future row (from the perspective of a row sample) to craft features. The validation strategy is thus should be taken carefully.</p>\n</blockquote>\n<p>You are speaking from the point of validation scheme but it seems we can here do active/incremental learning as well. (re-training the model based on the latest test-data combined.) It has a lot of caveats as well where our newly trained model is better than the old one etc, then we should use it else keep the old one etc. </p>",
      "rawMarkdown": ">For example, if all rows in private test set have timestamps after all rows in train (for each user_id), then in training pipeline we should not use any future row (from the perspective of a row sample) to craft features. The validation strategy is thus should be taken carefully.\n\nYou are speaking from the point of validation scheme but it seems we can here do active/incremental learning as well. (re-training the model based on the latest test-data combined.) It has a lot of caveats as well where our newly trained model is better than the old one etc, then we should use it else keep the old one etc.",
      "votes": null
    },
    {
      "id": "1048410",
      "postDate": "10/13/2020 13:12:01",
      "content": "<p><a href=\"https://www.kaggle.com/ningjiaca\" target=\"_blank\">@ningjiaca</a><br>\nQ2 might be false and not leak any data (e.g. just missing imputation task).<br>\nBut I agree with you regarding Q3: it's irrelevant.</p>",
      "rawMarkdown": "ningjiaca\nQ2 might be false and not leak any data (e.g. just missing imputation task).\nBut I agree with you regarding Q3: it's irrelevant.",
      "votes": null
    },
    {
      "id": "1048436",
      "postDate": "10/13/2020 13:46:37",
      "content": "<p><a href=\"https://www.kaggle.com/khahuras\" target=\"_blank\">@khahuras</a> That does not have anything to do with the answer, I guess. Ur validation scheme should always be ready for the worst case scenario and thus, u need to use forward chaining here.</p>",
      "rawMarkdown": "khahuras That does not have anything to do with the answer, I guess. Ur validation scheme should always be ready for the worst case scenario and thus, u need to use forward chaining here.",
      "votes": null
    },
    {
      "id": "1050018",
      "postDate": "10/15/2020 01:41:37",
      "content": "<p>All questions here are answered by Kaggle . Thanks a lot. <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106</a></p>",
      "rawMarkdown": "All questions here are answered by Kaggle . Thanks a lot. https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106",
      "votes": null
    },
    {
      "id": "1050334",
      "postDate": "10/15/2020 09:34:58",
      "content": "<p>you are right </p>",
      "rawMarkdown": "you are right",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1045521,
      "author_name": "elvinagammed",
      "author_url": "",
      "post_date": "10/10/2020 18:05:32",
      "content": "<p>I am trying to understand if any Y/N to any questions here changes your will of building a better validated model without leakage?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1045735,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "10/11/2020 01:34:44",
          "content": "<p>Yes, this is important because we don't know how to build a reliable validation scheme without knowing some basic things about the private test.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1046430,
          "author_name": "elvinagammed",
          "author_url": "",
          "post_date": "10/11/2020 16:47:47",
          "content": "<p>can u name any difference in your validation scheme if in any case difference between yes or no to any of the questions ? <a href=\"https://www.kaggle.com/khahuras\" target=\"_blank\">@khahuras</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1048161,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "10/13/2020 08:40:17",
          "content": "<p>For example, if all rows in private test set have timestamps after all rows in train (for each user_id), then in training pipeline we should not use any future row (from the perspective of a row sample) to craft features. The validation strategy is thus should be taken carefully.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1048168,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "10/13/2020 08:44:14",
          "content": "<blockquote>\n  <p>For example, if all rows in private test set have timestamps after all rows in train (for each user_id), then in training pipeline we should not use any future row (from the perspective of a row sample) to craft features. The validation strategy is thus should be taken carefully.</p>\n</blockquote>\n<p>You are speaking from the point of validation scheme but it seems we can here do active/incremental learning as well. (re-training the model based on the latest test-data combined.) It has a lot of caveats as well where our newly trained model is better than the old one etc, then we should use it else keep the old one etc. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1048436,
          "author_name": "elvinagammed",
          "author_url": "",
          "post_date": "10/13/2020 13:46:37",
          "content": "<p><a href=\"https://www.kaggle.com/khahuras\" target=\"_blank\">@khahuras</a> That does not have anything to do with the answer, I guess. Ur validation scheme should always be ready for the worst case scenario and thus, u need to use forward chaining here.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1045785,
      "author_name": "ningjiaca",
      "author_url": "",
      "post_date": "10/11/2020 03:09:13",
      "content": "<p>Q1: Yes, but there are user_ids not in the train at all.<br>\nQ2: Yes.<br>\nQ3: Doesn't matter. Because even the answer is Yes, we still cannot use the full history. Please take look of the example test file, only <code>prior_question_elapsed_time</code>,<code>prior_question_had_explanation</code>,<code>prior_group_responses</code> and <code>prior_group_answers_correct</code> have some limited history info.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1045790,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "10/11/2020 03:40:38",
          "content": "<p>How do you know Q2 is Yes? And I don't understand your point in Q3. If the batches of groups in private test set are increasing and cover all history, we can accumulate full history of any user.</p>",
          "votes": null,
          "replies": [
            {
              "id": 1046435,
              "author_name": "ningjiaca",
              "author_url": "",
              "post_date": "10/11/2020 16:51:42",
              "content": "<blockquote>\n  <p>How do you know Q2 is Yes? And I don't understand your point in Q3. If the batches of groups in private test set are increasing and cover all history, we can accumulate full history of any user.</p>\n</blockquote>\n<p>I cannot valid Q2 directly, that's my assumption. Because the organizer should consider to avoid basic data leakage. I think the train/test split is done either by users basis or time basis. Both cases cannot result in the data leakage you described in Q2. </p>\n<p>For Q3, do you think it's possible to accumulate the full history in the test submissions notebook?  Considering the limit of running time and memory requirement of the notebook and also the difficulties of debugging in those time series API calls, honestly I don't know for now. Maybe we can only accumulate very short history for each user.  </p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 1046446,
              "author_name": "adityaecdrid",
              "author_url": "",
              "post_date": "10/11/2020 17:00:48",
              "content": "<p>We can accumulate for sure but there's no assurance when will we get to see the same user again. It's less likely that you will run out of memory if you are \"purely\" making a \"test-only-submission\". Plus we have to call predict for every group_num irrespective of the same, so at any point you cannot use your accumulated data to full advantage as iterating over is coupled with env.predict. (that's my understanding as of now atleast)</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 1048410,
              "author_name": "carlossouza",
              "author_url": "",
              "post_date": "10/13/2020 13:12:01",
              "content": "<p><a href=\"https://www.kaggle.com/ningjiaca\" target=\"_blank\">@ningjiaca</a><br>\nQ2 might be false and not leak any data (e.g. just missing imputation task).<br>\nBut I agree with you regarding Q3: it's irrelevant.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1045838,
          "author_name": "msharuk589",
          "author_url": "",
          "post_date": "10/11/2020 05:05:23",
          "content": "<p><a href=\"https://www.kaggle.com/khahuras\" target=\"_blank\">@khahuras</a> <br>\nThere are 2 <code>user_id</code> that overlaps with train and example test daa<br>\nfor the <code>user_id</code> = <code>7792299</code>, its continuous.<br>\nAnd for the other user_id there is a huge gap in their timestamp.</p>\n<p>EDIT : I'm sorry I loaded only like 10**5 rows, that's why i only got 2 user_id overlapped with test data</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1046026,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "10/11/2020 09:22:52",
      "content": "<p>Plus, I can see people are creating features using group-by etc on every batch, but is that even correct? It's already visible that there can be a user_id in group_num 10 and the same user_id can appear in group_num 100. And we don't even know the order etc nor (as of now) we can club all the test set as well as we have to make a prediction for every group_num. I am not quite sue how can we override the preds that's already made and is not accessible anymore.<br>\nThoughts? Ty!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1048155,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "10/13/2020 08:38:00",
          "content": "<p>We cannot override the preds that already been made. What we can do is to accumulate test.csv up until the current processing batch.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1048160,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "10/13/2020 08:39:56",
          "content": "<p>I agree we cannot overwrite the previous preds but we can certainly make corrections to our next set of preds for sure by using info from those 2 extra columns..</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1050018,
      "author_name": "khahuras",
      "author_url": "",
      "post_date": "10/15/2020 01:41:37",
      "content": "<p>All questions here are answered by Kaggle . Thanks a lot. <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1050334,
      "author_name": "ayusharma18",
      "author_url": "",
      "post_date": "10/15/2020 09:34:58",
      "content": "<p>you are right </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1045422": "I think we the competitors are deserved to be answered some very basic but important questions about the private test set. If unanswered, it will leave frustration in model design/feature engineering. I will present 3 typical questions and explain why it is very important to be answered.\n\nQuestion 1: Is there any user_id in private test set that overlap with train set?\n\nWhy this question is important:\n+ If the answer is Yes, then user_id is one important feature that can be aggregated by the average target. However, another question (question 2) will pop up regarding how to engineer this feature.\n\nQuestion 2: If the answer to question 1 is Yes (which is very probable), then are the timestamps of that user_id in private test set AFTER the last timestamp of that user_id in train set?\n\nWhy this question is important:\n+ If the answer is Yes, then we don't have leakage. Otherwise, leakage will present, affecting the engineering pipeline.\n\nAnd the MOST IMPORTANT QUESTION is:\nQuestion 3: Suppose the (real) history of user ABC has 100 rows. Suppose this user is NOT present in train set. Are all of these 100 rows present in private test set, or just some of them (for example, 20 rows)? \n\nWhy this question is important:\n+ If the answer is No, then we cannot use full history of any user_id in the train set to build a sequential model (RNN, Transformer...). In other words, any sequential model will fail in test set because we don't have full history up to any row. And we can't know unless you tell us.\n\nI think these questions are not anything sensitive to not be answered. They are very basic, but important to the competition. Please give us the answers. \n\n@sohier",
    "1045521": "I am trying to understand if any Y/N to any questions here changes your will of building a better validated model without leakage?",
    "1045735": "Yes, this is important because we don't know how to build a reliable validation scheme without knowing some basic things about the private test.",
    "1045785": "Q1: Yes, but there are user_ids not in the train at all.\nQ2: Yes.\nQ3: Doesn't matter. Because even the answer is Yes, we still cannot use the full history. Please take look of the example test file, only `prior_question_elapsed_time`,`prior_question_had_explanation`,`prior_group_responses ` and `prior_group_answers_correct` have some limited history info.",
    "1045790": "How do you know Q2 is Yes? And I don't understand your point in Q3. If the batches of groups in private test set are increasing and cover all history, we can accumulate full history of any user.",
    "1045838": "khahuras \nThere are 2 `user_id` that overlaps with train and example test daa\nfor the `user_id` = `7792299`, its continuous.\nAnd for the other user_id there is a huge gap in their timestamp.\n\n\nEDIT : I'm sorry I loaded only like 10**5 rows, that's why i only got 2 user_id overlapped with test data",
    "1046026": "Plus, I can see people are creating features using group-by etc on every batch, but is that even correct? It's already visible that there can be a user_id in group_num 10 and the same user_id can appear in group_num 100. And we don't even know the order etc nor (as of now) we can club all the test set as well as we have to make a prediction for every group_num. I am not quite sue how can we override the preds that's already made and is not accessible anymore.\nThoughts? Ty!",
    "1046430": "can u name any difference in your validation scheme if in any case difference between yes or no to any of the questions ? @khahuras",
    "1046435": "> How do you know Q2 is Yes? And I don't understand your point in Q3. If the batches of groups in private test set are increasing and cover all history, we can accumulate full history of any user.\n\nI cannot valid Q2 directly, that's my assumption. Because the organizer should consider to avoid basic data leakage. I think the train/test split is done either by users basis or time basis. Both cases cannot result in the data leakage you described in Q2. \n\nFor Q3, do you think it's possible to accumulate the full history in the test submissions notebook?  Considering the limit of running time and memory requirement of the notebook and also the difficulties of debugging in those time series API calls, honestly I don't know for now. Maybe we can only accumulate very short history for each user.",
    "1046446": "We can accumulate for sure but there's no assurance when will we get to see the same user again. It's less likely that you will run out of memory if you are \"purely\" making a \"test-only-submission\". Plus we have to call predict for every group_num irrespective of the same, so at any point you cannot use your accumulated data to full advantage as iterating over is coupled with env.predict. (that's my understanding as of now atleast)",
    "1048155": "We cannot override the preds that already been made. What we can do is to accumulate test.csv up until the current processing batch.",
    "1048160": "I agree we cannot overwrite the previous preds but we can certainly make corrections to our next set of preds for sure by using info from those 2 extra columns..",
    "1048161": "For example, if all rows in private test set have timestamps after all rows in train (for each user_id), then in training pipeline we should not use any future row (from the perspective of a row sample) to craft features. The validation strategy is thus should be taken carefully.",
    "1048168": ">For example, if all rows in private test set have timestamps after all rows in train (for each user_id), then in training pipeline we should not use any future row (from the perspective of a row sample) to craft features. The validation strategy is thus should be taken carefully.\n\nYou are speaking from the point of validation scheme but it seems we can here do active/incremental learning as well. (re-training the model based on the latest test-data combined.) It has a lot of caveats as well where our newly trained model is better than the old one etc, then we should use it else keep the old one etc.",
    "1048410": "ningjiaca\nQ2 might be false and not leak any data (e.g. just missing imputation task).\nBut I agree with you regarding Q3: it's irrelevant.",
    "1048436": "khahuras That does not have anything to do with the answer, I guess. Ur validation scheme should always be ready for the worst case scenario and thus, u need to use forward chaining here.",
    "1050018": "All questions here are answered by Kaggle . Thanks a lot. https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106",
    "1050334": "you are right"
  },
  "source": "meta"
}