{
  "id": 190389,
  "title": "Question about prior_question_elapsed_time and prior_question_had_explanation",
  "url": "/competitions/riiid-test-answer-prediction/discussion/190389",
  "author_name": "",
  "post_date": "2020-10-11T15:02:25.596408800Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>prior_question_elapsed_time and prior_question_had_explanation, both are related to prior question bundle. So, they explains the situation of prior question. If I am right, shouldn't we should move one index above. Please tell me if I am wrong. Btw, if that is true can you help me with how to do that, I mean for loop is not helping (some error like pandas copying issue)</p>",
  "messages": [
    {
      "id": "1046343",
      "postDate": "10/11/2020 15:02:25",
      "content": "<p>prior_question_elapsed_time and prior_question_had_explanation, both are related to prior question bundle. So, they explains the situation of prior question. If I am right, shouldn't we should move one index above. Please tell me if I am wrong. Btw, if that is true can you help me with how to do that, I mean for loop is not helping (some error like pandas copying issue)</p>",
      "rawMarkdown": "prior_question_elapsed_time and prior_question_had_explanation, both are related to prior question bundle. So, they explains the situation of prior question. If I am right, shouldn't we should move one index above. Please tell me if I am wrong. Btw, if that is true can you help me with how to do that, I mean for loop is not helping (some error like pandas copying issue)",
      "votes": null
    },
    {
      "id": "1064538",
      "postDate": "10/30/2020 08:18:41",
      "content": "<p>I think it's not as simple as that, since questions with same <code>bundle_id</code> all have the same <code>prior_question_elapsed_time</code>. When you want to know the elapsed time for the current question you have to </p>\n<ol>\n<li>make sure your dataframe is sorted by <code>timestamp</code></li>\n<li>iterate over dataframes that only contain one speciffic <code>user_id</code> </li>\n<li>shift <code>prior_question_elapsed_time</code> one above (as you mentioned)</li>\n<li>finally always take the last entry for a section of questions sharing the same <code>bundle_id</code>.</li>\n</ol>\n<p>However, (that is at least what I think) it should not be allowed to do this (or let's say you should not do this). The reason is, that knowing the time it takes to answer a question can only be observed after answering the question. I guess the test data batches are organized in a way that each batch contains (mainly) different users and therefore you can not observe the current elapsed question time (and you have to fill this missing value by a method of your choice)</p>",
      "rawMarkdown": "I think it's not as simple as that, since questions with same `bundle_id` all have the same `prior_question_elapsed_time`. When you want to know the elapsed time for the current question you have to \n\n1. make sure your dataframe is sorted by `timestamp`\n2. iterate over dataframes that only contain one speciffic `user_id` \n3. shift `prior_question_elapsed_time` one above (as you mentioned)\n4. finally always take the last entry for a section of questions sharing the same `bundle_id`.\n\nHowever, (that is at least what I think) it should not be allowed to do this (or let's say you should not do this). The reason is, that knowing the time it takes to answer a question can only be observed after answering the question. I guess the test data batches are organized in a way that each batch contains (mainly) different users and therefore you can not observe the current elapsed question time (and you have to fill this missing value by a method of your choice)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1064538,
      "author_name": "markuslill",
      "author_url": "",
      "post_date": "10/30/2020 08:18:41",
      "content": "<p>I think it's not as simple as that, since questions with same <code>bundle_id</code> all have the same <code>prior_question_elapsed_time</code>. When you want to know the elapsed time for the current question you have to </p>\n<ol>\n<li>make sure your dataframe is sorted by <code>timestamp</code></li>\n<li>iterate over dataframes that only contain one speciffic <code>user_id</code> </li>\n<li>shift <code>prior_question_elapsed_time</code> one above (as you mentioned)</li>\n<li>finally always take the last entry for a section of questions sharing the same <code>bundle_id</code>.</li>\n</ol>\n<p>However, (that is at least what I think) it should not be allowed to do this (or let's say you should not do this). The reason is, that knowing the time it takes to answer a question can only be observed after answering the question. I guess the test data batches are organized in a way that each batch contains (mainly) different users and therefore you can not observe the current elapsed question time (and you have to fill this missing value by a method of your choice)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1046343": "prior_question_elapsed_time and prior_question_had_explanation, both are related to prior question bundle. So, they explains the situation of prior question. If I am right, shouldn't we should move one index above. Please tell me if I am wrong. Btw, if that is true can you help me with how to do that, I mean for loop is not helping (some error like pandas copying issue)",
    "1064538": "I think it's not as simple as that, since questions with same `bundle_id` all have the same `prior_question_elapsed_time`. When you want to know the elapsed time for the current question you have to \n\n1. make sure your dataframe is sorted by `timestamp`\n2. iterate over dataframes that only contain one speciffic `user_id` \n3. shift `prior_question_elapsed_time` one above (as you mentioned)\n4. finally always take the last entry for a section of questions sharing the same `bundle_id`.\n\nHowever, (that is at least what I think) it should not be allowed to do this (or let's say you should not do this). The reason is, that knowing the time it takes to answer a question can only be observed after answering the question. I guess the test data batches are organized in a way that each batch contains (mainly) different users and therefore you can not observe the current elapsed question time (and you have to fill this missing value by a method of your choice)"
  },
  "source": "meta"
}