{
  "id": 196879,
  "title": "[Find a Bug?]`content_type_id == 0` & `timstamp == 0` but `prior_question_elapsed_time` is not NaN?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/196879",
  "author_name": "",
  "post_date": "2020-11-13T07:59:22.617841400Z",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p><code>prior_question_elapsed_time</code>: (float32) The average time in milliseconds it took a user to answer each question in the previous question bundle, ignoring any lectures in between. <strong>Is null for a user's first question bundle or lecture</strong>. Note that the time is the average time a user took to solve each question in the previous bundle.<br>\nBut when I check it up,</p>\n<pre><code>cond1 = train_df['content_type_id'] == 0\ncond2 = train_df['timestamp'] == 0\nb = train_df.loc[cond1 &amp; cond2]\nb[b[\"prior_question_elapsed_time\"].notna()]\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F999983%2Fd69ef1222c653d8b58f20dfbe41c4fc8%2FScreen%20Shot%202020-11-15%20at%2016.10.16.png?generation=1605424279321157&amp;alt=media\" alt=\"\"><br>\n3896 data does have integer values(not NaN values) even though they have <code>timestamp</code> of 0 and <code>content_type_id</code> of 0. Why is it?</p>",
  "messages": [
    {
      "id": "1077065",
      "postDate": "11/13/2020 07:59:22",
      "content": "<p><code>prior_question_elapsed_time</code>: (float32) The average time in milliseconds it took a user to answer each question in the previous question bundle, ignoring any lectures in between. <strong>Is null for a user's first question bundle or lecture</strong>. Note that the time is the average time a user took to solve each question in the previous bundle.<br>\nBut when I check it up,</p>\n<pre><code>cond1 = train_df['content_type_id'] == 0\ncond2 = train_df['timestamp'] == 0\nb = train_df.loc[cond1 &amp; cond2]\nb[b[\"prior_question_elapsed_time\"].notna()]\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F999983%2Fd69ef1222c653d8b58f20dfbe41c4fc8%2FScreen%20Shot%202020-11-15%20at%2016.10.16.png?generation=1605424279321157&amp;alt=media\" alt=\"\"><br>\n3896 data does have integer values(not NaN values) even though they have <code>timestamp</code> of 0 and <code>content_type_id</code> of 0. Why is it?</p>",
      "rawMarkdown": "`prior_question_elapsed_time`: (float32) The average time in milliseconds it took a user to answer each question in the previous question bundle, ignoring any lectures in between. **Is null for a user's first question bundle or lecture**. Note that the time is the average time a user took to solve each question in the previous bundle.\n\nBut when I check it up,\n\n```\ncond1 = train_df['content_type_id'] == 0\ncond2 = train_df['timestamp'] == 0\nb = train_df.loc[cond1 & cond2]\nb[b[\"prior_question_elapsed_time\"].notna()]\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F999983%2Fd69ef1222c653d8b58f20dfbe41c4fc8%2FScreen%20Shot%202020-11-15%20at%2016.10.16.png?generation=1605424279321157&alt=media)\n\n\n3896 data does have integer values(not NaN values) even though they have `timestamp` of 0 and `content_type_id` of 0. Why is it?",
      "votes": null
    },
    {
      "id": "1078585",
      "postDate": "11/15/2020 03:26:05",
      "content": "<p>That's an interesting find. I double checked if their <code>task_container_id</code> (indicator of the order in which the student <em>saw</em> the questions) are not 0s just in case, but found their task_container_ids to be 0 as well. May be a bug with the data.</p>",
      "rawMarkdown": "That's an interesting find. I double checked if their `task_container_id` (indicator of the order in which the student *saw* the questions) are not 0s just in case, but found their task_container_ids to be 0 as well. May be a bug with the data.",
      "votes": null
    },
    {
      "id": "1078690",
      "postDate": "11/15/2020 07:15:37",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <a href=\"https://www.kaggle.com/jineonjinbaek\" target=\"_blank\">@jineonjinbaek</a> Could you guys check this out, please?</p>",
      "rawMarkdown": "sohier @jineonjinbaek Could you guys check this out, please?",
      "votes": null
    },
    {
      "id": "1126929",
      "postDate": "12/26/2020 05:08:27",
      "content": "<p><a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a>  i referred your other discussion related to prio and task container id. <br>\nDefault data is sorted by TS. Ideally Task container id should also be increasing assuming   system is serving question task container id in increasing order.  In a offline assesment student may jump to any other question at any time ,but in online system provides the bundle in seq and hence task container id should also be generated in 1 after the other. <br>\nOverall still trying to understand why Prior elapsed time in any row may not always refer to Task container id interaction in previous Time stamp when sorted by TS. </p>",
      "rawMarkdown": "doctorkael  i referred your other discussion related to prio and task container id. \nDefault data is sorted by TS. Ideally Task container id should also be increasing assuming   system is serving question task container id in increasing order.  In a offline assesment student may jump to any other question at any time ,but in online system provides the bundle in seq and hence task container id should also be generated in 1 after the other. \nOverall still trying to understand why Prior elapsed time in any row may not always refer to Task container id interaction in previous Time stamp when sorted by TS.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1078585,
      "author_name": "doctorkael",
      "author_url": "",
      "post_date": "11/15/2020 03:26:05",
      "content": "<p>That's an interesting find. I double checked if their <code>task_container_id</code> (indicator of the order in which the student <em>saw</em> the questions) are not 0s just in case, but found their task_container_ids to be 0 as well. May be a bug with the data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1126929,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/26/2020 05:08:27",
          "content": "<p><a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a>  i referred your other discussion related to prio and task container id. <br>\nDefault data is sorted by TS. Ideally Task container id should also be increasing assuming   system is serving question task container id in increasing order.  In a offline assesment student may jump to any other question at any time ,but in online system provides the bundle in seq and hence task container id should also be generated in 1 after the other. <br>\nOverall still trying to understand why Prior elapsed time in any row may not always refer to Task container id interaction in previous Time stamp when sorted by TS. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1078690,
      "author_name": "choisgood",
      "author_url": "",
      "post_date": "11/15/2020 07:15:37",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <a href=\"https://www.kaggle.com/jineonjinbaek\" target=\"_blank\">@jineonjinbaek</a> Could you guys check this out, please?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1077065": "`prior_question_elapsed_time`: (float32) The average time in milliseconds it took a user to answer each question in the previous question bundle, ignoring any lectures in between. **Is null for a user's first question bundle or lecture**. Note that the time is the average time a user took to solve each question in the previous bundle.\n\nBut when I check it up,\n\n```\ncond1 = train_df['content_type_id'] == 0\ncond2 = train_df['timestamp'] == 0\nb = train_df.loc[cond1 & cond2]\nb[b[\"prior_question_elapsed_time\"].notna()]\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F999983%2Fd69ef1222c653d8b58f20dfbe41c4fc8%2FScreen%20Shot%202020-11-15%20at%2016.10.16.png?generation=1605424279321157&alt=media)\n\n\n3896 data does have integer values(not NaN values) even though they have `timestamp` of 0 and `content_type_id` of 0. Why is it?",
    "1078585": "That's an interesting find. I double checked if their `task_container_id` (indicator of the order in which the student *saw* the questions) are not 0s just in case, but found their task_container_ids to be 0 as well. May be a bug with the data.",
    "1078690": "sohier @jineonjinbaek Could you guys check this out, please?",
    "1126929": "doctorkael  i referred your other discussion related to prio and task container id. \nDefault data is sorted by TS. Ideally Task container id should also be increasing assuming   system is serving question task container id in increasing order.  In a offline assesment student may jump to any other question at any time ,but in online system provides the bundle in seq and hence task container id should also be generated in 1 after the other. \nOverall still trying to understand why Prior elapsed time in any row may not always refer to Task container id interaction in previous Time stamp when sorted by TS."
  },
  "source": "meta"
}