{
  "id": 190552,
  "title": "how lectures in training set will help?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/190552",
  "author_name": "",
  "post_date": "2020-10-12T10:03:54.377753800Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am little confused by the data description:</p>\n<p>*user_answer: (int8) the user's answer to the question, if any. Read -1 as null, for lectures.</p>\n<p>answered_correctly: (int8) if the user responded correctly. Read -1 as null, for lectures.*</p>\n<p>if -1 meaning that is not a question, and the target feature is  -1 as well.</p>\n<p>why we have these records, seems they not helping the model to predict an answer is correct or not.</p>\n<p>or we are going to predict in the test set if a record is a lecture?</p>",
  "messages": [
    {
      "id": "1047146",
      "postDate": "10/12/2020 10:03:54",
      "content": "<p>I am little confused by the data description:</p>\n<p>*user_answer: (int8) the user's answer to the question, if any. Read -1 as null, for lectures.</p>\n<p>answered_correctly: (int8) if the user responded correctly. Read -1 as null, for lectures.*</p>\n<p>if -1 meaning that is not a question, and the target feature is  -1 as well.</p>\n<p>why we have these records, seems they not helping the model to predict an answer is correct or not.</p>\n<p>or we are going to predict in the test set if a record is a lecture?</p>",
      "rawMarkdown": "I am little confused by the data description:\n\n*user_answer: (int8) the user's answer to the question, if any. Read -1 as null, for lectures.\n\nanswered_correctly: (int8) if the user responded correctly. Read -1 as null, for lectures.*\n\nif -1 meaning that is not a question, and the target feature is  -1 as well.\n\nwhy we have these records, seems they not helping the model to predict an answer is correct or not.\n\nor we are going to predict in the test set if a record is a lecture?",
      "votes": null
    },
    {
      "id": "1047152",
      "postDate": "10/12/2020 10:09:53",
      "content": "<p>because these are raw data that need to be processed by us. The request of the competition is to make prediction over questions, then we know for sure that all the lectures sessions are not important, this information is useful for the preprocessing analysis</p>",
      "rawMarkdown": "because these are raw data that need to be processed by us. The request of the competition is to make prediction over questions, then we know for sure that all the lectures sessions are not important, this information is useful for the preprocessing analysis",
      "votes": null
    },
    {
      "id": "1047184",
      "postDate": "10/12/2020 10:57:27",
      "content": "<p>No you should not try to predict the lectures. You have to remove them. But you can use them create create feature (e.g you could create a feature that describe how many lectures one student attended).</p>",
      "rawMarkdown": "No you should not try to predict the lectures. You have to remove them. But you can use them create create feature (e.g you could create a feature that describe how many lectures one student attended).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1047152,
      "author_name": "domizianostingi",
      "author_url": "",
      "post_date": "10/12/2020 10:09:53",
      "content": "<p>because these are raw data that need to be processed by us. The request of the competition is to make prediction over questions, then we know for sure that all the lectures sessions are not important, this information is useful for the preprocessing analysis</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1047184,
      "author_name": "yohannwattiez",
      "author_url": "",
      "post_date": "10/12/2020 10:57:27",
      "content": "<p>No you should not try to predict the lectures. You have to remove them. But you can use them create create feature (e.g you could create a feature that describe how many lectures one student attended).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1047146": "I am little confused by the data description:\n\n*user_answer: (int8) the user's answer to the question, if any. Read -1 as null, for lectures.\n\nanswered_correctly: (int8) if the user responded correctly. Read -1 as null, for lectures.*\n\nif -1 meaning that is not a question, and the target feature is  -1 as well.\n\nwhy we have these records, seems they not helping the model to predict an answer is correct or not.\n\nor we are going to predict in the test set if a record is a lecture?",
    "1047152": "because these are raw data that need to be processed by us. The request of the competition is to make prediction over questions, then we know for sure that all the lectures sessions are not important, this information is useful for the preprocessing analysis",
    "1047184": "No you should not try to predict the lectures. You have to remove them. But you can use them create create feature (e.g you could create a feature that describe how many lectures one student attended)."
  },
  "source": "meta"
}