{
  "id": 190430,
  "title": "Can anyone help me understand \"prior_group_answers_correct\" column in the test set?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/190430",
  "author_name": "",
  "post_date": "2020-10-11T19:08:26.055933300Z",
  "votes": 6,
  "comment_count": 12,
  "views": 0,
  "content": "<p>It seems to me that the \"prior_group_answers_correct\" column contains the truth for the previous group of questions, which we can use to compare to the prediction we made in the previous group.<br>\nIt's like we made predictions, then the organizers told us the truth immediately.  Maybe I totally misunderstood the meaning of this column.</p>",
  "messages": [
    {
      "id": "1046553",
      "postDate": "10/11/2020 19:08:26",
      "content": "<p>It seems to me that the \"prior_group_answers_correct\" column contains the truth for the previous group of questions, which we can use to compare to the prediction we made in the previous group.<br>\nIt's like we made predictions, then the organizers told us the truth immediately.  Maybe I totally misunderstood the meaning of this column.</p>",
      "rawMarkdown": "It seems to me that the \"prior_group_answers_correct\" column contains the truth for the previous group of questions, which we can use to compare to the prediction we made in the previous group.\nIt's like we made predictions, then the organizers told us the truth immediately.  Maybe I totally misunderstood the meaning of this column.",
      "votes": null
    },
    {
      "id": "1046574",
      "postDate": "10/11/2020 19:39:41",
      "content": "<p>No, it is like that. Once you submitted your predictions for a group, you get the next test_df off the iterator and that immediately tells you whether you were right or not. You can use this information to improve your model before continuing with going through the test set, or you can just ignore it.<br>\nAs you can't submit predictions for the same group twice, you can't cheat with it. It's just meant to be used for improving your prediction algorithm as you get more information, as is typical for realtime applications.</p>",
      "rawMarkdown": "No, it is like that. Once you submitted your predictions for a group, you get the next test_df off the iterator and that immediately tells you whether you were right or not. You can use this information to improve your model before continuing with going through the test set, or you can just ignore it.\nAs you can't submit predictions for the same group twice, you can't cheat with it. It's just meant to be used for improving your prediction algorithm as you get more information, as is typical for realtime applications.",
      "votes": null
    },
    {
      "id": "1046610",
      "postDate": "10/11/2020 20:42:24",
      "content": "<p>Yes it is. Ideally, you can retrain your model …. only first you need to make a model that can be retrained on a small amount of new data</p>",
      "rawMarkdown": "Yes it is. Ideally, you can retrain your model .... only first you need to make a model that can be retrained on a small amount of new data",
      "votes": null
    },
    {
      "id": "1046676",
      "postDate": "10/11/2020 22:14:00",
      "content": "<p>Thank you so much! <a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a> , <a href=\"https://www.kaggle.com/sapr3s\" target=\"_blank\">@sapr3s</a> </p>",
      "rawMarkdown": "Thank you so much! @spacelx , @sapr3s",
      "votes": null
    },
    {
      "id": "1047683",
      "postDate": "10/12/2020 21:00:29",
      "content": "<p>Thanks for asking, it was in me mind too</p>",
      "rawMarkdown": "Thanks for asking, it was in me mind too",
      "votes": null
    },
    {
      "id": "1047899",
      "postDate": "10/13/2020 03:37:23",
      "content": "<p>I think the main use of that column is to update user based features as you go.(e.g mean correct answer per user). Also can be used to start building stats for new users</p>",
      "rawMarkdown": "I think the main use of that column is to update user based features as you go.(e.g mean correct answer per user). Also can be used to start building stats for new users",
      "votes": null
    },
    {
      "id": "1047963",
      "postDate": "10/13/2020 05:08:49",
      "content": "<p>So, you mean to say that we will have the real predictions (the actual labels for the test set for the previous batch we submitted in the current test-set chunk)?</p>",
      "rawMarkdown": "So, you mean to say that we will have the real predictions (the actual labels for the test set for the previous batch we submitted in the current test-set chunk)?",
      "votes": null
    },
    {
      "id": "1047992",
      "postDate": "10/13/2020 05:36:02",
      "content": "<p>That seems to be the case</p>",
      "rawMarkdown": "That seems to be the case",
      "votes": null
    },
    {
      "id": "1048055",
      "postDate": "10/13/2020 06:50:26",
      "content": "<p>Yeah that's how I understand it… otherwise there would be no way of getting a user's full previous performance.</p>",
      "rawMarkdown": "Yeah that's how I understand it... otherwise there would be no way of getting a user's full previous performance.",
      "votes": null
    },
    {
      "id": "1048083",
      "postDate": "10/13/2020 07:06:47",
      "content": "<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190748\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190748</a> Makes sense to me now! This means that we can use test set to train as well incrementally…</p>",
      "rawMarkdown": "https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190748 Makes sense to me now! This means that we can use test set to train as well incrementally...",
      "votes": null
    },
    {
      "id": "1048117",
      "postDate": "10/13/2020 08:01:33",
      "content": "<p>Yep, that's exactly what my demo kernel <a href=\"https://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme\" target=\"_blank\">https://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme</a> does and what could make incremental learning a viable strategy here! Really liking this competition so far.</p>",
      "rawMarkdown": "Yep, that's exactly what my demo kernel https://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme does and what could make incremental learning a viable strategy here! Really liking this competition so far.",
      "votes": null
    },
    {
      "id": "1048121",
      "postDate": "10/13/2020 08:04:31",
      "content": "<p>Well I wish they worded the desc properly and gave us little more example_test rows to play around as it's crucial to understand the data first and throw in whatever you want :) This comp is a SDE design challenge as well, in 4 years here, seeing it for the first time!</p>",
      "rawMarkdown": "Well I wish they worded the desc properly and gave us little more example_test rows to play around as it's crucial to understand the data first and throw in whatever you want :) This comp is a SDE design challenge as well, in 4 years here, seeing it for the first time!",
      "votes": null
    },
    {
      "id": "1048125",
      "postDate": "10/13/2020 08:10:16",
      "content": "<p>Having a fully representative public test set would certainly reduce the frustration in finding minor bugs yeah… it's annoying enough that you have to restart to kernel every time you run the test submission loop to fix bugs, but it's really bad that some bugs only show up in the private test set because the public one isn't complete and you have to make actual submissions with certain lines commented out just to find out where it fails.</p>",
      "rawMarkdown": "Having a fully representative public test set would certainly reduce the frustration in finding minor bugs yeah... it's annoying enough that you have to restart to kernel every time you run the test submission loop to fix bugs, but it's really bad that some bugs only show up in the private test set because the public one isn't complete and you have to make actual submissions with certain lines commented out just to find out where it fails.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1046574,
      "author_name": "spacelx",
      "author_url": "",
      "post_date": "10/11/2020 19:39:41",
      "content": "<p>No, it is like that. Once you submitted your predictions for a group, you get the next test_df off the iterator and that immediately tells you whether you were right or not. You can use this information to improve your model before continuing with going through the test set, or you can just ignore it.<br>\nAs you can't submit predictions for the same group twice, you can't cheat with it. It's just meant to be used for improving your prediction algorithm as you get more information, as is typical for realtime applications.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1047963,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "10/13/2020 05:08:49",
          "content": "<p>So, you mean to say that we will have the real predictions (the actual labels for the test set for the previous batch we submitted in the current test-set chunk)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1047992,
          "author_name": "abhimanyud",
          "author_url": "",
          "post_date": "10/13/2020 05:36:02",
          "content": "<p>That seems to be the case</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1048055,
          "author_name": "spacelx",
          "author_url": "",
          "post_date": "10/13/2020 06:50:26",
          "content": "<p>Yeah that's how I understand it… otherwise there would be no way of getting a user's full previous performance.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1048083,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "10/13/2020 07:06:47",
          "content": "<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190748\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190748</a> Makes sense to me now! This means that we can use test set to train as well incrementally…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1048117,
          "author_name": "spacelx",
          "author_url": "",
          "post_date": "10/13/2020 08:01:33",
          "content": "<p>Yep, that's exactly what my demo kernel <a href=\"https://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme\" target=\"_blank\">https://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme</a> does and what could make incremental learning a viable strategy here! Really liking this competition so far.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1048121,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "10/13/2020 08:04:31",
          "content": "<p>Well I wish they worded the desc properly and gave us little more example_test rows to play around as it's crucial to understand the data first and throw in whatever you want :) This comp is a SDE design challenge as well, in 4 years here, seeing it for the first time!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1048125,
          "author_name": "spacelx",
          "author_url": "",
          "post_date": "10/13/2020 08:10:16",
          "content": "<p>Having a fully representative public test set would certainly reduce the frustration in finding minor bugs yeah… it's annoying enough that you have to restart to kernel every time you run the test submission loop to fix bugs, but it's really bad that some bugs only show up in the private test set because the public one isn't complete and you have to make actual submissions with certain lines commented out just to find out where it fails.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1046610,
      "author_name": "sapr3s",
      "author_url": "",
      "post_date": "10/11/2020 20:42:24",
      "content": "<p>Yes it is. Ideally, you can retrain your model …. only first you need to make a model that can be retrained on a small amount of new data</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1046676,
      "author_name": "lrtmonkey",
      "author_url": "",
      "post_date": "10/11/2020 22:14:00",
      "content": "<p>Thank you so much! <a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a> , <a href=\"https://www.kaggle.com/sapr3s\" target=\"_blank\">@sapr3s</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1047683,
      "author_name": "domizianostingi",
      "author_url": "",
      "post_date": "10/12/2020 21:00:29",
      "content": "<p>Thanks for asking, it was in me mind too</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1047899,
      "author_name": "abhimanyud",
      "author_url": "",
      "post_date": "10/13/2020 03:37:23",
      "content": "<p>I think the main use of that column is to update user based features as you go.(e.g mean correct answer per user). Also can be used to start building stats for new users</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1046553": "It seems to me that the \"prior_group_answers_correct\" column contains the truth for the previous group of questions, which we can use to compare to the prediction we made in the previous group.\nIt's like we made predictions, then the organizers told us the truth immediately.  Maybe I totally misunderstood the meaning of this column.",
    "1046574": "No, it is like that. Once you submitted your predictions for a group, you get the next test_df off the iterator and that immediately tells you whether you were right or not. You can use this information to improve your model before continuing with going through the test set, or you can just ignore it.\nAs you can't submit predictions for the same group twice, you can't cheat with it. It's just meant to be used for improving your prediction algorithm as you get more information, as is typical for realtime applications.",
    "1046610": "Yes it is. Ideally, you can retrain your model .... only first you need to make a model that can be retrained on a small amount of new data",
    "1046676": "Thank you so much! @spacelx , @sapr3s",
    "1047683": "Thanks for asking, it was in me mind too",
    "1047899": "I think the main use of that column is to update user based features as you go.(e.g mean correct answer per user). Also can be used to start building stats for new users",
    "1047963": "So, you mean to say that we will have the real predictions (the actual labels for the test set for the previous batch we submitted in the current test-set chunk)?",
    "1047992": "That seems to be the case",
    "1048055": "Yeah that's how I understand it... otherwise there would be no way of getting a user's full previous performance.",
    "1048083": "https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190748 Makes sense to me now! This means that we can use test set to train as well incrementally...",
    "1048117": "Yep, that's exactly what my demo kernel https://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme does and what could make incremental learning a viable strategy here! Really liking this competition so far.",
    "1048121": "Well I wish they worded the desc properly and gave us little more example_test rows to play around as it's crucial to understand the data first and throw in whatever you want :) This comp is a SDE design challenge as well, in 4 years here, seeing it for the first time!",
    "1048125": "Having a fully representative public test set would certainly reduce the frustration in finding minor bugs yeah... it's annoying enough that you have to restart to kernel every time you run the test submission loop to fix bugs, but it's really bad that some bugs only show up in the private test set because the public one isn't complete and you have to make actual submissions with certain lines commented out just to find out where it fails."
  },
  "source": "meta"
}