{
  "id": 194184,
  "title": "Mapping `prior_question_*` features with their respective rows",
  "url": "/competitions/riiid-test-answer-prediction/discussion/194184",
  "author_name": "",
  "post_date": "2020-10-31T06:58:36.824683500Z",
  "votes": 23,
  "comment_count": 33,
  "views": 0,
  "content": "<p><code>prior_question_elapsed_time</code> and <code>prior_question_had_explanation</code> seem kind of important to me. No discussion so far however, seems specify how these features need to be mapped with their corresponding rows.  </p>\n<p>The description for pq_had_explanation reads: </p>\n<pre><code>Whether or not the user saw an explanation and the correct response(s) after answering the previous question bundle, ignoring any lectures in between. The value is shared across a single question bundle, and is null for a user's first question bundle or lecture.\n</code></pre>\n<p>If by the term <em>question bundle</em> they refer to the <code>bundle_id</code>, I'm not really sure what <em>previous</em> question bundle would mean. Say for bundle_id 99 the <code>prior_question_had_explanation</code> is True. Then does it mean that the <code>bundle_id</code> 98 had explanation? Or does it refer to the bundle_id in the <em>previous position</em>? We should also note that the <code>bundle_id</code> is not served in any order, so the bundle_id in the previous position would be something random like 1902 or 0.</p>\n<p>If however the term <em>previous question bundle</em> refers to the previous <code>task_container_id</code> (I strongly suspect this to the case) the task to map with corresponding row gets much simpler. </p>\n<p>For the same student, the same questions were asked repeatedly maybe to verify if the student didn't get it right merely by <em>chance</em>. Some questions if answered wrong were asked until the student got it right. The <code>prior_question_had_explanation</code> if mapped correctly could be an important feature to determine if the students gets it right the second time around.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4175266%2Fa793109d08210eef931f22f1a3832ef8%2F1.png?generation=1604127572115869&amp;alt=media\" alt=\"\"></p>\n<p>It would be really nice if someone could help out with this! </p>",
  "messages": [
    {
      "id": "1065331",
      "postDate": "10/31/2020 06:58:36",
      "content": "<p><code>prior_question_elapsed_time</code> and <code>prior_question_had_explanation</code> seem kind of important to me. No discussion so far however, seems specify how these features need to be mapped with their corresponding rows.  </p>\n<p>The description for pq_had_explanation reads: </p>\n<pre><code>Whether or not the user saw an explanation and the correct response(s) after answering the previous question bundle, ignoring any lectures in between. The value is shared across a single question bundle, and is null for a user's first question bundle or lecture.\n</code></pre>\n<p>If by the term <em>question bundle</em> they refer to the <code>bundle_id</code>, I'm not really sure what <em>previous</em> question bundle would mean. Say for bundle_id 99 the <code>prior_question_had_explanation</code> is True. Then does it mean that the <code>bundle_id</code> 98 had explanation? Or does it refer to the bundle_id in the <em>previous position</em>? We should also note that the <code>bundle_id</code> is not served in any order, so the bundle_id in the previous position would be something random like 1902 or 0.</p>\n<p>If however the term <em>previous question bundle</em> refers to the previous <code>task_container_id</code> (I strongly suspect this to the case) the task to map with corresponding row gets much simpler. </p>\n<p>For the same student, the same questions were asked repeatedly maybe to verify if the student didn't get it right merely by <em>chance</em>. Some questions if answered wrong were asked until the student got it right. The <code>prior_question_had_explanation</code> if mapped correctly could be an important feature to determine if the students gets it right the second time around.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4175266%2Fa793109d08210eef931f22f1a3832ef8%2F1.png?generation=1604127572115869&amp;alt=media\" alt=\"\"></p>\n<p>It would be really nice if someone could help out with this! </p>",
      "rawMarkdown": "`prior_question_elapsed_time` and `prior_question_had_explanation` seem kind of important to me. No discussion so far however, seems specify how these features need to be mapped with their corresponding rows.  \n\nThe description for pq_had_explanation reads: \n```\nWhether or not the user saw an explanation and the correct response(s) after answering the previous question bundle, ignoring any lectures in between. The value is shared across a single question bundle, and is null for a user's first question bundle or lecture.\n```\n\nIf by the term *question bundle* they refer to the `bundle_id`, I'm not really sure what *previous* question bundle would mean. Say for bundle_id 99 the `prior_question_had_explanation` is True. Then does it mean that the `bundle_id` 98 had explanation? Or does it refer to the bundle_id in the *previous position*? We should also note that the `bundle_id` is not served in any order, so the bundle_id in the previous position would be something random like 1902 or 0.\n\nIf however the term *previous question bundle* refers to the previous `task_container_id` (I strongly suspect this to the case) the task to map with corresponding row gets much simpler. \n\nFor the same student, the same questions were asked repeatedly maybe to verify if the student didn't get it right merely by *chance*. Some questions if answered wrong were asked until the student got it right. The `prior_question_had_explanation` if mapped correctly could be an important feature to determine if the students gets it right the second time around.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4175266%2Fa793109d08210eef931f22f1a3832ef8%2F1.png?generation=1604127572115869&alt=media)\n\nIt would be really nice if someone could help out with this!",
      "votes": null
    },
    {
      "id": "1065422",
      "postDate": "10/31/2020 09:14:59",
      "content": "<blockquote>\n  <p>If however the term previous question bundle refers to the previous task_container_id (I strongly suspect this to the case) the task to map with corresponding row gets much simpler.</p>\n</blockquote>\n<p>I believe this is correct and how I've interpreted it.</p>",
      "rawMarkdown": "> If however the term previous question bundle refers to the previous task_container_id (I strongly suspect this to the case) the task to map with corresponding row gets much simpler.\n\nI believe this is correct and how I've interpreted it.",
      "votes": null
    },
    {
      "id": "1065758",
      "postDate": "10/31/2020 17:46:48",
      "content": "<p><a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a>  did u get clarity to it.. <br>\nPrevious means any question or series of questions asked to user  just prior to this   <br>\nSo if suppose<br>\nq1 - Correct<br>\nq2-Correct<br>\nq3-Incorrect<br>\nq4-correct<br>\nq5-Incorrect- prior explaination =True</p>\n<p>So   in this prior could be  Q4   as  may be student was trying to apply similar thought process to solve q5 but  response was incorrect so he thought let me check the explaination for q5..</p>\n<p><code>For the same student, the same questions were asked repeatedly maybe to verify if the student didn't get it right merely by chance.</code> how you know this</p>",
      "rawMarkdown": "doctorkael  did u get clarity to it.. \nPrevious means any question or series of questions asked to user  just prior to this   \nSo if suppose\nq1 - Correct\nq2-Correct\nq3-Incorrect\nq4-correct\nq5-Incorrect- prior explaination =True\n\nSo   in this prior could be  Q4   as  may be student was trying to apply similar thought process to solve q5 but  response was incorrect so he thought let me check the explaination for q5..\n\n`For the same student, the same questions were asked repeatedly maybe to verify if the student didn't get it right merely by chance.` how you know this",
      "votes": null
    },
    {
      "id": "1066011",
      "postDate": "11/01/2020 07:29:42",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a>! <code>prior_questions_had_explanation</code> seems like the questions <em>either came with explanations at the end or they didn't</em>. It would make sense however for the user to be able to check the answers for questions he needed more clarity on.  </p>\n<p>Just like <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> told I have started using <code>task_container_id</code> to map the <code>prior_question_*</code> features with the corresponding rows. A student seeing the explanation for a question that was answered incorrectly would certainly do better when the same question gets asked again. Verifying this assumption for a few users by manual inspection has indeed proved correct.</p>\n<p>Some students were asked the same question a lot of times to the point that it feels like there might have been a mistake. Take user_id <code>15632472</code>. He was asked this question_id <code>1796</code> for 30 times. This student seems to be more of an anomaly though.</p>\n<p>The general trend however is that roughly <em>11% of the entire questions asked</em> were questions that had already been asked to the same user (More repeated questions for some users &amp; some with 0 repeated questions). The rationale for asking these questions again and again might have been to test the user's knowledge thoroughly not leaving things upto <em>chance</em>.</p>\n<p>PS: My analysis was only done on the first 1e6 rows. But I assume that this percentage to be roughly same for the entire dataset.</p>",
      "rawMarkdown": "Hey @jaideepvalani! `prior_questions_had_explanation` seems like the questions *either came with explanations at the end or they didn't*. It would make sense however for the user to be able to check the answers for questions he needed more clarity on.  \n\nJust like @rohanrao told I have started using `task_container_id` to map the `prior_question_*` features with the corresponding rows. A student seeing the explanation for a question that was answered incorrectly would certainly do better when the same question gets asked again. Verifying this assumption for a few users by manual inspection has indeed proved correct.\n\nSome students were asked the same question a lot of times to the point that it feels like there might have been a mistake. Take user_id `15632472`. He was asked this question_id `1796` for 30 times. This student seems to be more of an anomaly though.\n\nThe general trend however is that roughly *11% of the entire questions asked* were questions that had already been asked to the same user (More repeated questions for some users & some with 0 repeated questions). The rationale for asking these questions again and again might have been to test the user's knowledge thoroughly not leaving things upto *chance*.\n\nPS: My analysis was only done on the first 1e6 rows. But I assume that this percentage to be roughly same for the entire dataset.",
      "votes": null
    },
    {
      "id": "1067743",
      "postDate": "11/02/2020 17:06:13",
      "content": "<p>I've also understood it in the same manner. <br>\n<a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a> did you get any progress on the mapping ? </p>",
      "rawMarkdown": "I've also understood it in the same manner. \n@doctorkael did you get any progress on the mapping ?",
      "votes": null
    },
    {
      "id": "1068211",
      "postDate": "11/03/2020 07:18:25",
      "content": "<p>I haven't incorporated any of those features into a model yet. The huge data and the special API are intimidating to say the least. </p>\n<p>But after further explorations, I found those features to be predictive of the answer correctness for repeated questions. It is indeed previous <code>task_container_id</code> that the features need to be shifted on.</p>",
      "rawMarkdown": "I haven't incorporated any of those features into a model yet. The huge data and the special API are intimidating to say the least. \n\nBut after further explorations, I found those features to be predictive of the answer correctness for repeated questions. It is indeed previous `task_container_id ` that the features need to be shifted on.",
      "votes": null
    },
    {
      "id": "1068321",
      "postDate": "11/03/2020 09:23:37",
      "content": "<p>The question is : how to do it without exploding the processing time </p>",
      "rawMarkdown": "The question is : how to do it without exploding the processing time",
      "votes": null
    },
    {
      "id": "1092674",
      "postDate": "11/27/2020 04:45:56",
      "content": "<p>Sorry, when you all say \"previous <code>task_container_id</code>\", what do you mean by \"previous\"? <code>task_container_id - 1</code>? Or whichever one appeared last in timestamp order for that user? because apparently, according to other posts, those are not necessarily the same (the <code>task_container_id</code> doesn't always monotonically increase with the timestamp).</p>\n<p>I'm trying to figure out how to decide whether a user saw feedback <em>for a given question</em>, since it's kind of important in modeling that user's state. </p>",
      "rawMarkdown": "Sorry, when you all say \"previous `task_container_id`\", what do you mean by \"previous\"? `task_container_id - 1`? Or whichever one appeared last in timestamp order for that user? because apparently, according to other posts, those are not necessarily the same (the `task_container_id` doesn't always monotonically increase with the timestamp).\n\nI'm trying to figure out how to decide whether a user saw feedback *for a given question*, since it's kind of important in modeling that user's state.",
      "votes": null
    },
    {
      "id": "1092743",
      "postDate": "11/27/2020 06:07:19",
      "content": "<p>\"Previous\" <code>task_container_id</code> refers to <code>task_container_id - 1</code>. This is the where the features such as <code>prior_question_elapsed_time</code> and <code>prior_question_had_explanation</code> needs to be mapped to (ignoring any lectures that may come in between). </p>\n<p>Case 1: Without any lectures in between<br>\nAssume a user has answered a question with <code>task_container_id</code> <strong>1</strong>. The same user had seen the explanation if prior_question_had_explanation for <code>task_container_id</code> <strong>2</strong> is True. </p>\n<p>Case 2: User saw a few lectures in between<br>\nAssume a user has answered a question with <code>task_container_id</code> <strong>1</strong>. Suppose now that the user had seen two lectures in between with <code>task_container_id</code>'s 2 &amp; 3. The same user had seen the explanation for the question (of task_container_id 1) if the prior_question_had_explanation for <code>task_container_id</code> <strong>4</strong> is True. <em>Note that prior_question_had_explanation for task_container_ids 2 and 3 (the lectures) would be NA and they can be safely dropped or ignored.</em></p>",
      "rawMarkdown": "\"Previous\" `task_container_id` refers to `task_container_id - 1`. This is the where the features such as `prior_question_elapsed_time` and `prior_question_had_explanation` needs to be mapped to (ignoring any lectures that may come in between). \n\nCase 1: Without any lectures in between\nAssume a user has answered a question with `task_container_id` **1**. The same user had seen the explanation if prior_question_had_explanation for `task_container_id` **2** is True. \n\nCase 2: User saw a few lectures in between\nAssume a user has answered a question with `task_container_id` **1**. Suppose now that the user had seen two lectures in between with `task_container_id`'s 2 & 3. The same user had seen the explanation for the question (of task_container_id 1) if the prior_question_had_explanation for `task_container_id` **4** is True. _Note that prior_question_had_explanation for task_container_ids 2 and 3 (the lectures) would be NA and they can be safely dropped or ignored._",
      "votes": null
    },
    {
      "id": "1096922",
      "postDate": "11/30/2020 22:03:48",
      "content": "<p>Are we sure that the mapping is from bundle with <strong>task container id</strong> to <strong>task container id -1</strong>.</p>\n<p>Since there are some issues in the order of task container id. I believe the mapping should be from <strong>timestamp at current bundle</strong> to <strong>timestamp of previous bundle</strong>.</p>",
      "rawMarkdown": "Are we sure that the mapping is from bundle with **task container id** to **task container id -1**.\n\nSince there are some issues in the order of task container id. I believe the mapping should be from **timestamp at current bundle** to **timestamp of previous bundle**.",
      "votes": null
    },
    {
      "id": "1099034",
      "postDate": "12/02/2020 03:20:37",
      "content": "<blockquote>\n  <p>Since there are some issues in the order of task container id. I believe the mapping should be from <strong>timestamp at current bundle</strong> to <strong>timestamp of previous bundle</strong>.</p>\n</blockquote>\n<p>After some experimentation, I believe that \"previous\" is actually based on the timestamp order, not task container id order.</p>\n<p>I did manage to extract the elapsed time and had_explanation flag for the current question! And I'm pretty sure it's even correct. Here is my notebook (it also has some more text about why I think it's timestamp and not task_container_id that determines order):<br>\n<a href=\"https://www.kaggle.com/yanamal/explanation-and-elapsed-time-for-current-question\" target=\"_blank\">https://www.kaggle.com/yanamal/explanation-and-elapsed-time-for-current-question</a></p>",
      "rawMarkdown": "> Since there are some issues in the order of task container id. I believe the mapping should be from **timestamp at current bundle** to **timestamp of previous bundle**.\n\nAfter some experimentation, I believe that \"previous\" is actually based on the timestamp order, not task container id order.\n\nI did manage to extract the elapsed time and had_explanation flag for the current question! And I'm pretty sure it's even correct. Here is my notebook (it also has some more text about why I think it's timestamp and not task_container_id that determines order):\nhttps://www.kaggle.com/yanamal/explanation-and-elapsed-time-for-current-question",
      "votes": null
    },
    {
      "id": "1099202",
      "postDate": "12/02/2020 07:05:24",
      "content": "<p>Is task container id relevant other than as the number and order of bundles delivered to users?</p>",
      "rawMarkdown": "Is task container id relevant other than as the number and order of bundles delivered to users?",
      "votes": null
    },
    {
      "id": "1099275",
      "postDate": "12/02/2020 08:26:17",
      "content": "<p>It's relevant for the purpose of determining what \"previous\" means for \"previous question\", because \"previous question\" seems to actually be \"the question(s) in the task_container which was chronologically (by timestamp) served to the user most recently before this task_container\". If that makes any sense.</p>\n<p>p.s. the notebook I mentioned above is slightly broken (it only identifies the elapsed time/had_expanation for the first thing in a task container), but I'm having a lot of trouble implementing a fix that will actually fit in RAM and not crash the kernel.</p>",
      "rawMarkdown": "It's relevant for the purpose of determining what \"previous\" means for \"previous question\", because \"previous question\" seems to actually be \"the question(s) in the task_container which was chronologically (by timestamp) served to the user most recently before this task_container\". If that makes any sense.\n\np.s. the notebook I mentioned above is slightly broken (it only identifies the elapsed time/had_expanation for the first thing in a task container), but I'm having a lot of trouble implementing a fix that will actually fit in RAM and not crash the kernel.",
      "votes": null
    },
    {
      "id": "1099434",
      "postDate": "12/02/2020 11:06:49",
      "content": "<p>I wasn't able to do this for all users at the same time (RAM issues)<br>\nI ended up looping over user_ids and updating one user at a time. It works but takes about 30-40 mins</p>",
      "rawMarkdown": "I wasn't able to do this for all users at the same time (RAM issues)\nI ended up looping over user_ids and updating one user at a time. It works but takes about 30-40 mins",
      "votes": null
    },
    {
      "id": "1100220",
      "postDate": "12/02/2020 23:16:09",
      "content": "<p>Got it. That's what I was thinking as well, but wanted to make sure I wasn't missing some other important aspect. I've been having good luck with sqlite3 to keep processing times short.</p>",
      "rawMarkdown": "Got it. That's what I was thinking as well, but wanted to make sure I wasn't missing some other important aspect. I've been having good luck with sqlite3 to keep processing times short.",
      "votes": null
    },
    {
      "id": "1100584",
      "postDate": "12/03/2020 07:04:02",
      "content": "<p><a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a>  how do we know which interaction is chronologically just prior to the Row which prior explaination flag as True ,does sorting by time stamp for the users ensures that any row having prior explaination flag as True refers to interaction appearing before it after sorted by TS </p>",
      "rawMarkdown": "doctorkael  how do we know which interaction is chronologically just prior to the Row which prior explaination flag as True ,does sorting by time stamp for the users ensures that any row having prior explaination flag as True refers to interaction appearing before it after sorted by TS",
      "votes": null
    },
    {
      "id": "1100590",
      "postDate": "12/03/2020 07:09:54",
      "content": "<p><a href=\"https://www.kaggle.com/yanamal\" target=\"_blank\">@yanamal</a> <br>\nwhen you say <br>\n<code>question\" seems to actually be \"the question(s) in the task_container which was chronologically (by timestamp) served to the user most recently before this task_container\". If that makes any sense.</code></p>\n<p>then here does chronology means sorting by timestamp for each user not whole data set offcourse<br>\n i too felt that when we sort timestamp we get the interactions in the order in which user attempted those ..</p>",
      "rawMarkdown": "yanamal \nwhen you say \n`question\" seems to actually be \"the question(s) in the task_container which was chronologically (by timestamp) served to the user most recently before this task_container\". If that makes any sense.`\n\nthen here does chronology means sorting by timestamp for each user not whole data set offcourse\n i too felt that when we sort timestamp we get the interactions in the order in which user attempted those ..",
      "votes": null
    },
    {
      "id": "1100678",
      "postDate": "12/03/2020 08:44:11",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/calebeverett\" target=\"_blank\">@calebeverett</a>! <code>task_container_id</code> AFAIK doesn't seem to be that important by itself besides serving as a means to determine the order in which the users where served questions. We could use this feature to create better features that correlate to the target.</p>\n<hr>\n<p><a href=\"https://www.kaggle.com/yanamal\" target=\"_blank\">@yanamal</a> <a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> The issues with ordering is relatively minor. The vast majority of the users can be safely mapped using the <code>task_container_id</code>. So I think its an easier option to simply shift using user_id and task_container_id.  Steps to get this feature might be something like this: </p>\n<ol>\n<li>Sort the data by user_id and task_container_id (By default the order is user_id, timestamp).</li>\n<li>Groupby user_id, task_container_id, calculate the mean (Some questions come as bundles, mean will be the same for all user_id, task_container_id pair).</li>\n<li>Do a shift operation after doing a groupby on user_id alone. The sort operation we did before should ensure that the shift happens to task_container_id - 1.</li>\n<li>Merge the above result on user_id and task_container_id.</li>\n</ol>\n<p>This feature does indeed have a high importance on the target. Sadly it cannot be used during prediction since during prediction, we will see <em>utmost one bundle for a single user</em> which means that we will not see next bundle's <code>prior_question_elapsed_time</code> until we had already finished making our predictions for this bundle.</p>\n<hr>\n<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a>, the rows are already sorted on user_id and timestamp. You can check this by doing:</p>\n<pre><code>data.groupby(['user_id'])['timestamp'].is_monotonic_increasing.all()\n&gt;&gt;&gt; True\n</code></pre>",
      "rawMarkdown": "Hey @calebeverett! `task_container_id` AFAIK doesn't seem to be that important by itself besides serving as a means to determine the order in which the users where served questions. We could use this feature to create better features that correlate to the target.\n\n***\n\n@yanamal @abdurrafae The issues with ordering is relatively minor. The vast majority of the users can be safely mapped using the `task_container_id`. So I think its an easier option to simply shift using user_id and task_container_id.  Steps to get this feature might be something like this: \n\n1. Sort the data by user_id and task_container_id (By default the order is user_id, timestamp).\n2. Groupby user_id, task_container_id, calculate the mean (Some questions come as bundles, mean will be the same for all user_id, task_container_id pair).\n3. Do a shift operation after doing a groupby on user_id alone. The sort operation we did before should ensure that the shift happens to task_container_id - 1.\n4. Merge the above result on user_id and task_container_id.\n\nThis feature does indeed have a high importance on the target. Sadly it cannot be used during prediction since during prediction, we will see *utmost one bundle for a single user* which means that we will not see next bundle's `prior_question_elapsed_time` until we had already finished making our predictions for this bundle.\n\n***\n\n@jaideepvalani, the rows are already sorted on user_id and timestamp. You can check this by doing:\n```\ndata.groupby(['user_id'])['timestamp'].is_monotonic_increasing.all()\n>>> True\n```",
      "votes": null
    },
    {
      "id": "1100995",
      "postDate": "12/03/2020 14:35:41",
      "content": "<p>Thanks - even if <code>prior_question_elapsed_time</code> isn't available for the current question, perhaps there is something in an historical aggregation - by <code>user_id</code>, <code>user_id</code>-<code>content_id</code>, <code>user_id</code>-<code>tag</code> or  <code>user_id</code>-<code>part</code>.</p>",
      "rawMarkdown": "Thanks - even if `prior_question_elapsed_time` isn't available for the current question, perhaps there is something in an historical aggregation - by `user_id`, `user_id`-`content_id`, `user_id`-`tag` or  `user_id`-`part`.",
      "votes": null
    },
    {
      "id": "1101571",
      "postDate": "12/04/2020 03:11:12",
      "content": "<p>Yes! That would certainly be possible. Better to do an aggregation based on content_id since user based features require the hassle of updating them as new users arrive.</p>",
      "rawMarkdown": "Yes! That would certainly be possible. Better to do an aggregation based on content_id since user based features require the hassle of updating them as new users arrive.",
      "votes": null
    },
    {
      "id": "1103203",
      "postDate": "12/05/2020 18:01:13",
      "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> yeah, I meant sort by timestamp per user (so either group or sort by user)<br>\n(also, that's already how the data comes sorted)</p>\n<p><a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a> </p>\n<blockquote>\n  <p>Do a shift operation after doing a groupby on user_id alone. The sort operation we did before should ensure that the shift happens to task_container_id - 1.</p>\n</blockquote>\n<p>This won't actually achieve the desired result 100% of the time with bundles of size &gt; 1, since you're shifting by 1, and therefore shifting to the next question, which may still be in the same bundle.</p>\n<p>My current process is more like:</p>\n<ol>\n<li>don't bother sorting, because the existing sorting by user + timestamp works well enough (arguably better)</li>\n<li>don't bother averaging either, because I'm 99% sure the description says elapsed time is the same for every single question in the task_container, anyway.</li>\n<li>extract (into separate df) unique rows of <code>[user_id, task_container_id, prior_question_had_explanation, prior_question_elapsed_time]</code></li>\n<li>shift <code>prior_question*</code> in <em>that</em> dataframe to create <code>this_question*</code></li>\n<li>add these new rows back into the original dataframe (the indices will match up with first entry for each <code>[user, task_container_id]</code> combo)</li>\n<li>(can't do because of RAM-hugry pandas) group by <code>[user_id, task_container_id]</code> and ffill this_question_* to propagate the values within task_container</li>\n</ol>\n<p>I'm going to try numpy next, when I get around to it.</p>\n<blockquote>\n  <p>Sadly it cannot be used during prediction since during prediction, we will see <em>utmost one bundle for a single user</em> which means that we will not see next bundle's <code>prior_question_elapsed_time</code> until we had already finished making our predictions for this bundle.</p>\n</blockquote>\n<p>I'm pretty sure this is on purpose, and exactly why they went with this prior_question_* scheme. I think knowing how long it took to answer a particular question would be considered \"data leakage\", since in a real-life setting you would be trying to predict this before you even gave the question to the user. </p>\n<p>(nevermind that our predictions would only be trained to predict this for the kinds of questions that <strong>the app's algorithm chooses for that particular user</strong>, and so would probably not generalize usefully to questions overall…) </p>",
      "rawMarkdown": "jaideepvalani yeah, I meant sort by timestamp per user (so either group or sort by user)\n(also, that's already how the data comes sorted)\n\n@doctorkael \n\n> Do a shift operation after doing a groupby on user_id alone. The sort operation we did before should ensure that the shift happens to task_container_id - 1.\n\nThis won't actually achieve the desired result 100% of the time with bundles of size > 1, since you're shifting by 1, and therefore shifting to the next question, which may still be in the same bundle.\n\nMy current process is more like:\n1. don't bother sorting, because the existing sorting by user + timestamp works well enough (arguably better)\n2. don't bother averaging either, because I'm 99% sure the description says elapsed time is the same for every single question in the task_container, anyway.\n3. extract (into separate df) unique rows of `[user_id, task_container_id, prior_question_had_explanation, prior_question_elapsed_time]`\n4. shift `prior_question*` in *that* dataframe to create `this_question*`\n5. add these new rows back into the original dataframe (the indices will match up with first entry for each `[user, task_container_id]` combo)\n6. (can't do because of RAM-hugry pandas) group by `[user_id, task_container_id]` and ffill this_question_* to propagate the values within task_container\n\nI'm going to try numpy next, when I get around to it.\n\n> Sadly it cannot be used during prediction since during prediction, we will see *utmost one bundle for a single user* which means that we will not see next bundle's `prior_question_elapsed_time` until we had already finished making our predictions for this bundle.\n\nI'm pretty sure this is on purpose, and exactly why they went with this prior_question_* scheme. I think knowing how long it took to answer a particular question would be considered \"data leakage\", since in a real-life setting you would be trying to predict this before you even gave the question to the user. \n\n(nevermind that our predictions would only be trained to predict this for the kinds of questions that **the app's algorithm chooses for that particular user**, and so would probably not generalize usefully to questions overall...)",
      "votes": null
    },
    {
      "id": "1103923",
      "postDate": "12/06/2020 12:28:53",
      "content": "<p><a href=\"https://www.kaggle.com/yanamal\" target=\"_blank\">@yanamal</a> </p>\n<pre><code>This won't actually achieve the desired result 100% of the time with bundles of size &gt; 1, since you're shifting by 1, and therefore shifting to the next question, which may still be in the same bundle.\n</code></pre>\n<p>We do an average precisely to deal with those bundles with size &gt; 1. </p>\n<p>Given that for a single bundle, the contents can have only one unique pq_elapsed_time and one unique pq_had_explanation, doing an average (n times a value, divided by n gives the same value), shifting it on task_container_id and merging it back on user_id and task_container_id would always ensure that the pq_* are properly mapped (based on task_container_id in this case).</p>",
      "rawMarkdown": "yanamal \n\n```\nThis won't actually achieve the desired result 100% of the time with bundles of size > 1, since you're shifting by 1, and therefore shifting to the next question, which may still be in the same bundle.\n```\n\nWe do an average precisely to deal with those bundles with size > 1. \n\nGiven that for a single bundle, the contents can have only one unique pq_elapsed_time and one unique pq_had_explanation, doing an average (n times a value, divided by n gives the same value), shifting it on task_container_id and merging it back on user_id and task_container_id would always ensure that the pq_* are properly mapped (based on task_container_id in this case).",
      "votes": null
    },
    {
      "id": "1104084",
      "postDate": "12/06/2020 15:38:27",
      "content": "<p><a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a> <br>\n1) i dint understand how this constraint putting bottleneck in knowing prior time elapsed details</p>\n<p><code>next bundle's prior_question_elapsed_time until we had already finished making our predictions for this bundle.</code><br>\nI know that at a time we will have only group or one bundle(correct me if it can be more than one bundle also at a time in iteration)  to be predicted, so when we go to next group we have got now the time elapsed for previous bundle to use </p>\n<p>2) If i know correctly Set of users between Train/Valid and Test going to be same ?</p>\n<p>3) Question ids /Topics etc could always be different right ? so in that case one cant use the TIme elapsed data from the Train set right ? if its case one will have to accumulate this data as the prediction happens  for subsequent group ?</p>\n<p>Sorry if i sound very redudant ,its been only few days i got back to competition, currently working SAINT base not plus (partial features).. have achieved some success  in terms of  better CV score than SAKT 74.71 compared to 74.56 of SAKT. </p>",
      "rawMarkdown": "doctorkael \n1) i dint understand how this constraint putting bottleneck in knowing prior time elapsed details\n\n`next bundle's prior_question_elapsed_time until we had already finished making our predictions for this bundle.`\nI know that at a time we will have only group or one bundle(correct me if it can be more than one bundle also at a time in iteration)  to be predicted, so when we go to next group we have got now the time elapsed for previous bundle to use \n\n2) If i know correctly Set of users between Train/Valid and Test going to be same ?\n\n3) Question ids /Topics etc could always be different right ? so in that case one cant use the TIme elapsed data from the Train set right ? if its case one will have to accumulate this data as the prediction happens  for subsequent group ?\n\nSorry if i sound very redudant ,its been only few days i got back to competition, currently working SAINT base not plus (partial features).. have achieved some success  in terms of  better CV score than SAKT 74.71 compared to 74.56 of SAKT.",
      "votes": null
    },
    {
      "id": "1104308",
      "postDate": "12/06/2020 19:46:07",
      "content": "<p><a href=\"https://www.kaggle.com/doktorkael\" target=\"_blank\">@doktorkael</a> Then I don't understand what you mean by \"Do a shift operation after doing a groupby on user_id alone\". How are you \"shifting it on task_container_id\" if you're only grouping by user_id?</p>\n<p>Maybe we're actually talking about doing the same thing, after all (in terms of the shifting, anyway)</p>\n<p>Also, like I said, averaging will not change the data in any way at all, because it's already averaged across the bundles/task_containers. The data tab of the challenge says:</p>\n<blockquote>\n  <p><code>prior_question_elapsed_time</code>: (float32) The <strong>average</strong> time in milliseconds it took a user to answer each question <strong>in the previous question bundle</strong></p>\n</blockquote>\n<p>Granted, their use of the word \"bundle\" is rather confusing, but elsewhere they've clarified that they are talking about the user+task_conainer groupings. </p>",
      "rawMarkdown": "doktorkael Then I don't understand what you mean by \"Do a shift operation after doing a groupby on user_id alone\". How are you \"shifting it on task_container_id\" if you're only grouping by user_id?\n\nMaybe we're actually talking about doing the same thing, after all (in terms of the shifting, anyway)\n\nAlso, like I said, averaging will not change the data in any way at all, because it's already averaged across the bundles/task_containers. The data tab of the challenge says:\n\n> `prior_question_elapsed_time`: (float32) The **average** time in milliseconds it took a user to answer each question **in the previous question bundle**\n\nGranted, their use of the word \"bundle\" is rather confusing, but elsewhere they've clarified that they are talking about the user+task_conainer groupings.",
      "votes": null
    },
    {
      "id": "1105775",
      "postDate": "12/08/2020 07:29:34",
      "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a>, from the discussion <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106\" target=\"_blank\">here</a>:</p>\n<pre><code>Correction- the hidden test set contains new users but not new questions.\n</code></pre>\n<p>For this reason, content based features would require very little updation. It is user based feature that would need to be updated. And yes we would only be seeing one bundle per user during the test inference. Still it doesn't mean pq_* features are entirely useless. Like <a href=\"https://www.kaggle.com/calebeverett\" target=\"_blank\">@calebeverett</a> mentioned, we can perform some aggregation and craft useful features out of them :)</p>\n<hr>\n<p><a href=\"https://www.kaggle.com/yanamal\" target=\"_blank\">@yanamal</a>, <code>How are you \"shifting it on task_container_id\" if you're only grouping by user_id?</code></p>\n<p>Groupby by default sorts the dataframe after the aggregation is done. So doing a groupby using user_id and task_container_id would ensure that the resulting dataframe would be sorted by task_container_id per user. After this it's simply a matter of shifting them and re-merging them back to the dataframe.</p>\n<pre><code>Also, like I said, averaging will not change the data in any way at all, because it's already averaged across the bundles/task_containers.\n</code></pre>\n<p>The purpose of averaging is to ensure that there is only one entry per user_id and task_container_id. Agreed, it doesn't change the data in any meaningful way. This is done just to make the shift operation easier. Hope this helps :)</p>",
      "rawMarkdown": "jaideepvalani, from the discussion [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106):\n\n```\nCorrection- the hidden test set contains new users but not new questions.\n```\n\nFor this reason, content based features would require very little updation. It is user based feature that would need to be updated. And yes we would only be seeing one bundle per user during the test inference. Still it doesn't mean pq_* features are entirely useless. Like @calebeverett mentioned, we can perform some aggregation and craft useful features out of them :)\n\n***\n@yanamal, ```How are you \"shifting it on task_container_id\" if you're only grouping by user_id?```\n\nGroupby by default sorts the dataframe after the aggregation is done. So doing a groupby using user_id and task_container_id would ensure that the resulting dataframe would be sorted by task_container_id per user. After this it's simply a matter of shifting them and re-merging them back to the dataframe.\n\n```\nAlso, like I said, averaging will not change the data in any way at all, because it's already averaged across the bundles/task_containers.\n```\n\nThe purpose of averaging is to ensure that there is only one entry per user_id and task_container_id. Agreed, it doesn't change the data in any meaningful way. This is done just to make the shift operation easier. Hope this helps :)",
      "votes": null
    },
    {
      "id": "1105806",
      "postDate": "12/08/2020 08:08:36",
      "content": "<p>Ohh, I see now, thanks. So I think that is pretty similar to what I was doing, except I did unique() instead of averaging, which hopefully would produce the same results, as long as the prior_* fields are actually the same within the task_container.</p>\n<p>…. and also apparently I forgot to group by user in my notebook. details, details.</p>",
      "rawMarkdown": "Ohh, I see now, thanks. So I think that is pretty similar to what I was doing, except I did unique() instead of averaging, which hopefully would produce the same results, as long as the prior_* fields are actually the same within the task_container.\n\n.... and also apparently I forgot to group by user in my notebook. details, details.",
      "votes": null
    },
    {
      "id": "1105903",
      "postDate": "12/08/2020 10:10:58",
      "content": "<p><a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a> thanku.. if one is not using Aggregation features.. like following approach as there in SAINT /SAKT paper . <br>\nI was thinking how to handle the new users  in that case .Below is excerpt from SAKT inference part.</p>\n<p>if user_id in self.samples.index:<br>\n            q_, qa_ = self.samples[user_id]</p>\n<pre><code>        seq_len = len(q_)\n\n        if seq_len &gt;= self.max_seq:\n            q = q_[-self.max_seq:]\n            qa = qa_[-self.max_seq:]\n        else:\n            q[-seq_len:] = q_\n            qa[-seq_len:] = qa_          \n\n    x = np.zeros(self.max_seq-1, dtype=int)\n    x = q[1:].copy()\n    x += (qa[1:] == 1) * self.n_skill\n\n    questions = np.append(q[2:], [target_id])\n</code></pre>\n<p>Here if  suppose user id not present  then ?</p>\n<p>And i dint understand indexings 1:,2:  used , compared to :-1 for x and 1: for training</p>",
      "rawMarkdown": "doctorkael thanku.. if one is not using Aggregation features.. like following approach as there in SAINT /SAKT paper . \nI was thinking how to handle the new users  in that case .Below is excerpt from SAKT inference part.\n\n if user_id in self.samples.index:\n            q_, qa_ = self.samples[user_id]\n            \n            seq_len = len(q_)\n\n            if seq_len >= self.max_seq:\n                q = q_[-self.max_seq:]\n                qa = qa_[-self.max_seq:]\n            else:\n                q[-seq_len:] = q_\n                qa[-seq_len:] = qa_          \n        \n        x = np.zeros(self.max_seq-1, dtype=int)\n        x = q[1:].copy()\n        x += (qa[1:] == 1) * self.n_skill\n        \n        questions = np.append(q[2:], [target_id])\n\n\nHere if  suppose user id not present  then ?\n\nAnd i dint understand indexings 1:,2:  used , compared to :-1 for x and 1: for training",
      "votes": null
    },
    {
      "id": "1106099",
      "postDate": "12/08/2020 14:17:59",
      "content": "<blockquote>\n  <p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a>, from the discussion <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106\" target=\"_blank\">here</a>:</p>\n<pre><code>Correction- the hidden test set contains new users but not new questions.\n</code></pre>\n  <p>For this reason, content based features would require very little updation. It is user based feature that would need to be updated. And yes we would only be seeing one bundle per user during the test inference. Still it doesn't mean pq_* features are entirely useless. Like <a href=\"https://www.kaggle.com/calebeverett\" target=\"_blank\">@calebeverett</a> mentioned, we can perform some aggregation and craft useful features out of them :)</p>\n  <hr>\n  <p><a href=\"https://www.kaggle.com/yanamal\" target=\"_blank\">@yanamal</a>, <code>How are you \"shifting it on task_container_id\" if you're only grouping by user_id?</code></p>\n  <p>Groupby by default sorts the dataframe after the aggregation is done. So doing a groupby using user_id and task_container_id would ensure that the resulting dataframe would be sorted by task_container_id per user. After this it's simply a matter of shifting them and re-merging them back to the dataframe.</p>\n<pre><code>Also, like I said, averaging will not change the data in any way at all, because it's already averaged across the bundles/task_containers.\n</code></pre>\n  <p>The purpose of averaging is to ensure that there is only one entry per user_id and task_container_id. Agreed, it doesn't change the data in any meaningful way. This is done just to make the shift operation easier. Hope this helps :)</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a> <br>\nDo you mean of this sort<br>\n<code>train_df.groupby(['user_id','task_container_id']).agg(mpelt=('prior_question_elapsed_time','mean')).reset_index()</code> <br>\nfollowed by shift on Task container id and joining it back ?</p>\n<p>if its so then not sure data for this user complying to the assumption</p>\n<pre><code>user_id    task_container_id   elapsed time\n115    0   55000.0\n1    NaN\n2    37000.0\n3    19000.0\n4    11000.0\n</code></pre>",
      "rawMarkdown": "> @jaideepvalani, from the discussion [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106):\n> \n> ```\n> Correction- the hidden test set contains new users but not new questions.\n> ```\n> \n> For this reason, content based features would require very little updation. It is user based feature that would need to be updated. And yes we would only be seeing one bundle per user during the test inference. Still it doesn't mean pq_* features are entirely useless. Like @calebeverett mentioned, we can perform some aggregation and craft useful features out of them :)\n> \n> ***\n> @yanamal, ```How are you \"shifting it on task_container_id\" if you're only grouping by user_id?```\n> \n> Groupby by default sorts the dataframe after the aggregation is done. So doing a groupby using user_id and task_container_id would ensure that the resulting dataframe would be sorted by task_container_id per user. After this it's simply a matter of shifting them and re-merging them back to the dataframe.\n> \n> ```\n> Also, like I said, averaging will not change the data in any way at all, because it's already averaged across the bundles/task_containers.\n> ```\n> \n> The purpose of averaging is to ensure that there is only one entry per user_id and task_container_id. Agreed, it doesn't change the data in any meaningful way. This is done just to make the shift operation easier. Hope this helps :)\n\n@doctorkael \nDo you mean of this sort\n`train_df.groupby(['user_id','task_container_id']).agg(mpelt=('prior_question_elapsed_time','mean')).reset_index() ` \nfollowed by shift on Task container id and joining it back ?\n\nif its so then not sure data for this user complying to the assumption\n\n```\nuser_id\ttask_container_id\telapsed time\n115\t0\t55000.0\n1\tNaN\n2\t37000.0\n3\t19000.0\n4\t11000.0\n```",
      "votes": null
    },
    {
      "id": "1126983",
      "postDate": "12/26/2020 06:28:56",
      "content": "<blockquote>\n  <p>Got it. That's what I was thinking as well, but wanted to make sure I wasn't missing some other important aspect. I've been having good luck with sqlite3 to keep processing times short.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/calebeverett\" target=\"_blank\">@calebeverett</a> <br>\nDoes  it  discussion here suggests that suppose  if order of question in timestamp order<br>\nis q1,q2,q3,q4<br>\nq1 T2<br>\nQ2 T2<br>\nQ3 T2<br>\nQ4 T4<br>\nQ5-T4<br>\nprior time elapsed time  in Q3,Q4,Q5  refers to time student took to finish questions in Task container id T3 and not tast container T2 ?</p>\n<p>2) If above is case ,can not the user just sort by task container id  each user , and then serve to model ,that way prior time is aligned automatically</p>",
      "rawMarkdown": "> Got it. That's what I was thinking as well, but wanted to make sure I wasn't missing some other important aspect. I've been having good luck with sqlite3 to keep processing times short.\n\n@calebeverett \nDoes  it  discussion here suggests that suppose  if order of question in timestamp order\nis q1,q2,q3,q4\nq1 T2\nQ2 T2\nQ3 T2\nQ4 T4\nQ5-T4\nprior time elapsed time  in Q3,Q4,Q5  refers to time student took to finish questions in Task container id T3 and not tast container T2 ?\n\n2) If above is case ,can not the user just sort by task container id  each user , and then serve to model ,that way prior time is aligned automatically",
      "votes": null
    },
    {
      "id": "1127503",
      "postDate": "12/26/2020 15:32:43",
      "content": "<p>Hi,</p>\n<p>I believe from prior discussion the way to think about timestamp is that it when the user starts the bundle of questions and prior_question_elapsed_time is how long they took to complete it.</p>\n<p>There are cases where the task_container_id is not in the same order as timestamp (I counted approximately 2.5 million instances over the full training set), apparently because users may start questions on different devices and complete them in a different order than they started them. I ended up renumbering the task_container_ids so they were in the same order as the timestamps.</p>",
      "rawMarkdown": "Hi,\n\nI believe from prior discussion the way to think about timestamp is that it when the user starts the bundle of questions and prior_question_elapsed_time is how long they took to complete it.\n\nThere are cases where the task_container_id is not in the same order as timestamp (I counted approximately 2.5 million instances over the full training set), apparently because users may start questions on different devices and complete them in a different order than they started them. I ended up renumbering the task_container_ids so they were in the same order as the timestamps.",
      "votes": null
    },
    {
      "id": "1127511",
      "postDate": "12/26/2020 15:46:16",
      "content": "<blockquote>\n  <p>Hi,</p>\n  <p>I believe from prior discussion the way to think about timestamp is that it when the user starts the bundle of questions and prior_question_elapsed_time is how long they took to complete it.</p>\n  <p>There are cases where the task_container_id is not in the same order as timestamp (I counted approximately 2.5 million instances over the full training set), apparently because users may start questions on different devices and complete them in a different order than they started them. I ended up renumbering the task_container_ids so they were in the same order as the timestamps.</p>\n</blockquote>\n<p>Thanks , if suppose  we have this series <br>\nTs1 Q1 t3  t1   <br>\nTs2 Q2 t5  t3  pet2</p>\n<p>How renumbering help here </p>",
      "rawMarkdown": "> Hi,\n> \n> I believe from prior discussion the way to think about timestamp is that it when the user starts the bundle of questions and prior_question_elapsed_time is how long they took to complete it.\n> \n> There are cases where the task_container_id is not in the same order as timestamp (I counted approximately 2.5 million instances over the full training set), apparently because users may start questions on different devices and complete them in a different order than they started them. I ended up renumbering the task_container_ids so they were in the same order as the timestamps.\n\nThanks , if suppose  we have this series \nTs1 Q1 t3  t1   \nTs2 Q2 t5  t3  pet2\n\nHow renumbering help here",
      "votes": null
    },
    {
      "id": "1127548",
      "postDate": "12/26/2020 16:24:46",
      "content": "<p>Hi, not sure I totally followed the abbreviations, but those look like they are in order. I was doing some calculations based on differences in task container ids that required them to be in order to return the correct differences. If you just wanted the rows in the right order, I think you could just use timestamp.</p>",
      "rawMarkdown": "Hi, not sure I totally followed the abbreviations, but those look like they are in order. I was doing some calculations based on differences in task container ids that required them to be in order to return the correct differences. If you just wanted the rows in the right order, I think you could just use timestamp.",
      "votes": null
    },
    {
      "id": "1127556",
      "postDate": "12/26/2020 16:31:13",
      "content": "<p><a href=\"https://www.kaggle.com/calebeverett\" target=\"_blank\">@calebeverett</a> at time seq 1 completed task container id T3  and at ts2  completed T5 .. prior elapsed time  which t5 says is for T4 or T3..<br>\nis it the case that falls in 2.5 Million interactions.. ?</p>\n<p>basically if i pass Q1,q2,q3 seq to model then prior elapsed time in their respective r  interaction rows should refer to  that of question appearing  prior  in above  sequence . does time sort ensures same ?</p>",
      "rawMarkdown": "calebeverett at time seq 1 completed task container id T3  and at ts2  completed T5 .. prior elapsed time  which t5 says is for T4 or T3..\nis it the case that falls in 2.5 Million interactions.. ?\n\nbasically if i pass Q1,q2,q3 seq to model then prior elapsed time in their respective r  interaction rows should refer to  that of question appearing  prior  in above  sequence . does time sort ensures same ?",
      "votes": null
    },
    {
      "id": "1127705",
      "postDate": "12/26/2020 19:08:04",
      "content": "<p>Ok, I see what you are asking - good question and I don't know the answer, i.e., does prior_question_elapsed_time relate to the interaction with the immediately preceding task_container_id or the immediately preceding timestamp?</p>",
      "rawMarkdown": "Ok, I see what you are asking - good question and I don't know the answer, i.e., does prior_question_elapsed_time relate to the interaction with the immediately preceding task_container_id or the immediately preceding timestamp?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1065422,
      "author_name": "rohanrao",
      "author_url": "",
      "post_date": "10/31/2020 09:14:59",
      "content": "<blockquote>\n  <p>If however the term previous question bundle refers to the previous task_container_id (I strongly suspect this to the case) the task to map with corresponding row gets much simpler.</p>\n</blockquote>\n<p>I believe this is correct and how I've interpreted it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1067743,
          "author_name": "alexj21",
          "author_url": "",
          "post_date": "11/02/2020 17:06:13",
          "content": "<p>I've also understood it in the same manner. <br>\n<a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a> did you get any progress on the mapping ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1068211,
          "author_name": "doctorkael",
          "author_url": "",
          "post_date": "11/03/2020 07:18:25",
          "content": "<p>I haven't incorporated any of those features into a model yet. The huge data and the special API are intimidating to say the least. </p>\n<p>But after further explorations, I found those features to be predictive of the answer correctness for repeated questions. It is indeed previous <code>task_container_id</code> that the features need to be shifted on.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1068321,
          "author_name": "alexj21",
          "author_url": "",
          "post_date": "11/03/2020 09:23:37",
          "content": "<p>The question is : how to do it without exploding the processing time </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1092674,
          "author_name": "yanamal",
          "author_url": "",
          "post_date": "11/27/2020 04:45:56",
          "content": "<p>Sorry, when you all say \"previous <code>task_container_id</code>\", what do you mean by \"previous\"? <code>task_container_id - 1</code>? Or whichever one appeared last in timestamp order for that user? because apparently, according to other posts, those are not necessarily the same (the <code>task_container_id</code> doesn't always monotonically increase with the timestamp).</p>\n<p>I'm trying to figure out how to decide whether a user saw feedback <em>for a given question</em>, since it's kind of important in modeling that user's state. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1092743,
          "author_name": "doctorkael",
          "author_url": "",
          "post_date": "11/27/2020 06:07:19",
          "content": "<p>\"Previous\" <code>task_container_id</code> refers to <code>task_container_id - 1</code>. This is the where the features such as <code>prior_question_elapsed_time</code> and <code>prior_question_had_explanation</code> needs to be mapped to (ignoring any lectures that may come in between). </p>\n<p>Case 1: Without any lectures in between<br>\nAssume a user has answered a question with <code>task_container_id</code> <strong>1</strong>. The same user had seen the explanation if prior_question_had_explanation for <code>task_container_id</code> <strong>2</strong> is True. </p>\n<p>Case 2: User saw a few lectures in between<br>\nAssume a user has answered a question with <code>task_container_id</code> <strong>1</strong>. Suppose now that the user had seen two lectures in between with <code>task_container_id</code>'s 2 &amp; 3. The same user had seen the explanation for the question (of task_container_id 1) if the prior_question_had_explanation for <code>task_container_id</code> <strong>4</strong> is True. <em>Note that prior_question_had_explanation for task_container_ids 2 and 3 (the lectures) would be NA and they can be safely dropped or ignored.</em></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1096922,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "11/30/2020 22:03:48",
          "content": "<p>Are we sure that the mapping is from bundle with <strong>task container id</strong> to <strong>task container id -1</strong>.</p>\n<p>Since there are some issues in the order of task container id. I believe the mapping should be from <strong>timestamp at current bundle</strong> to <strong>timestamp of previous bundle</strong>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1099034,
          "author_name": "yanamal",
          "author_url": "",
          "post_date": "12/02/2020 03:20:37",
          "content": "<blockquote>\n  <p>Since there are some issues in the order of task container id. I believe the mapping should be from <strong>timestamp at current bundle</strong> to <strong>timestamp of previous bundle</strong>.</p>\n</blockquote>\n<p>After some experimentation, I believe that \"previous\" is actually based on the timestamp order, not task container id order.</p>\n<p>I did manage to extract the elapsed time and had_explanation flag for the current question! And I'm pretty sure it's even correct. Here is my notebook (it also has some more text about why I think it's timestamp and not task_container_id that determines order):<br>\n<a href=\"https://www.kaggle.com/yanamal/explanation-and-elapsed-time-for-current-question\" target=\"_blank\">https://www.kaggle.com/yanamal/explanation-and-elapsed-time-for-current-question</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1099202,
          "author_name": "calebeverett",
          "author_url": "",
          "post_date": "12/02/2020 07:05:24",
          "content": "<p>Is task container id relevant other than as the number and order of bundles delivered to users?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1099275,
          "author_name": "yanamal",
          "author_url": "",
          "post_date": "12/02/2020 08:26:17",
          "content": "<p>It's relevant for the purpose of determining what \"previous\" means for \"previous question\", because \"previous question\" seems to actually be \"the question(s) in the task_container which was chronologically (by timestamp) served to the user most recently before this task_container\". If that makes any sense.</p>\n<p>p.s. the notebook I mentioned above is slightly broken (it only identifies the elapsed time/had_expanation for the first thing in a task container), but I'm having a lot of trouble implementing a fix that will actually fit in RAM and not crash the kernel.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1099434,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/02/2020 11:06:49",
          "content": "<p>I wasn't able to do this for all users at the same time (RAM issues)<br>\nI ended up looping over user_ids and updating one user at a time. It works but takes about 30-40 mins</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1100220,
          "author_name": "calebeverett",
          "author_url": "",
          "post_date": "12/02/2020 23:16:09",
          "content": "<p>Got it. That's what I was thinking as well, but wanted to make sure I wasn't missing some other important aspect. I've been having good luck with sqlite3 to keep processing times short.</p>",
          "votes": null,
          "replies": [
            {
              "id": 1126983,
              "author_name": "jaideepvalani",
              "author_url": "",
              "post_date": "12/26/2020 06:28:56",
              "content": "<blockquote>\n  <p>Got it. That's what I was thinking as well, but wanted to make sure I wasn't missing some other important aspect. I've been having good luck with sqlite3 to keep processing times short.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/calebeverett\" target=\"_blank\">@calebeverett</a> <br>\nDoes  it  discussion here suggests that suppose  if order of question in timestamp order<br>\nis q1,q2,q3,q4<br>\nq1 T2<br>\nQ2 T2<br>\nQ3 T2<br>\nQ4 T4<br>\nQ5-T4<br>\nprior time elapsed time  in Q3,Q4,Q5  refers to time student took to finish questions in Task container id T3 and not tast container T2 ?</p>\n<p>2) If above is case ,can not the user just sort by task container id  each user , and then serve to model ,that way prior time is aligned automatically</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1100584,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/03/2020 07:04:02",
          "content": "<p><a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a>  how do we know which interaction is chronologically just prior to the Row which prior explaination flag as True ,does sorting by time stamp for the users ensures that any row having prior explaination flag as True refers to interaction appearing before it after sorted by TS </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1100590,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/03/2020 07:09:54",
          "content": "<p><a href=\"https://www.kaggle.com/yanamal\" target=\"_blank\">@yanamal</a> <br>\nwhen you say <br>\n<code>question\" seems to actually be \"the question(s) in the task_container which was chronologically (by timestamp) served to the user most recently before this task_container\". If that makes any sense.</code></p>\n<p>then here does chronology means sorting by timestamp for each user not whole data set offcourse<br>\n i too felt that when we sort timestamp we get the interactions in the order in which user attempted those ..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1100678,
          "author_name": "doctorkael",
          "author_url": "",
          "post_date": "12/03/2020 08:44:11",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/calebeverett\" target=\"_blank\">@calebeverett</a>! <code>task_container_id</code> AFAIK doesn't seem to be that important by itself besides serving as a means to determine the order in which the users where served questions. We could use this feature to create better features that correlate to the target.</p>\n<hr>\n<p><a href=\"https://www.kaggle.com/yanamal\" target=\"_blank\">@yanamal</a> <a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> The issues with ordering is relatively minor. The vast majority of the users can be safely mapped using the <code>task_container_id</code>. So I think its an easier option to simply shift using user_id and task_container_id.  Steps to get this feature might be something like this: </p>\n<ol>\n<li>Sort the data by user_id and task_container_id (By default the order is user_id, timestamp).</li>\n<li>Groupby user_id, task_container_id, calculate the mean (Some questions come as bundles, mean will be the same for all user_id, task_container_id pair).</li>\n<li>Do a shift operation after doing a groupby on user_id alone. The sort operation we did before should ensure that the shift happens to task_container_id - 1.</li>\n<li>Merge the above result on user_id and task_container_id.</li>\n</ol>\n<p>This feature does indeed have a high importance on the target. Sadly it cannot be used during prediction since during prediction, we will see <em>utmost one bundle for a single user</em> which means that we will not see next bundle's <code>prior_question_elapsed_time</code> until we had already finished making our predictions for this bundle.</p>\n<hr>\n<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a>, the rows are already sorted on user_id and timestamp. You can check this by doing:</p>\n<pre><code>data.groupby(['user_id'])['timestamp'].is_monotonic_increasing.all()\n&gt;&gt;&gt; True\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1100995,
          "author_name": "calebeverett",
          "author_url": "",
          "post_date": "12/03/2020 14:35:41",
          "content": "<p>Thanks - even if <code>prior_question_elapsed_time</code> isn't available for the current question, perhaps there is something in an historical aggregation - by <code>user_id</code>, <code>user_id</code>-<code>content_id</code>, <code>user_id</code>-<code>tag</code> or  <code>user_id</code>-<code>part</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1101571,
          "author_name": "doctorkael",
          "author_url": "",
          "post_date": "12/04/2020 03:11:12",
          "content": "<p>Yes! That would certainly be possible. Better to do an aggregation based on content_id since user based features require the hassle of updating them as new users arrive.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1103203,
          "author_name": "yanamal",
          "author_url": "",
          "post_date": "12/05/2020 18:01:13",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> yeah, I meant sort by timestamp per user (so either group or sort by user)<br>\n(also, that's already how the data comes sorted)</p>\n<p><a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a> </p>\n<blockquote>\n  <p>Do a shift operation after doing a groupby on user_id alone. The sort operation we did before should ensure that the shift happens to task_container_id - 1.</p>\n</blockquote>\n<p>This won't actually achieve the desired result 100% of the time with bundles of size &gt; 1, since you're shifting by 1, and therefore shifting to the next question, which may still be in the same bundle.</p>\n<p>My current process is more like:</p>\n<ol>\n<li>don't bother sorting, because the existing sorting by user + timestamp works well enough (arguably better)</li>\n<li>don't bother averaging either, because I'm 99% sure the description says elapsed time is the same for every single question in the task_container, anyway.</li>\n<li>extract (into separate df) unique rows of <code>[user_id, task_container_id, prior_question_had_explanation, prior_question_elapsed_time]</code></li>\n<li>shift <code>prior_question*</code> in <em>that</em> dataframe to create <code>this_question*</code></li>\n<li>add these new rows back into the original dataframe (the indices will match up with first entry for each <code>[user, task_container_id]</code> combo)</li>\n<li>(can't do because of RAM-hugry pandas) group by <code>[user_id, task_container_id]</code> and ffill this_question_* to propagate the values within task_container</li>\n</ol>\n<p>I'm going to try numpy next, when I get around to it.</p>\n<blockquote>\n  <p>Sadly it cannot be used during prediction since during prediction, we will see <em>utmost one bundle for a single user</em> which means that we will not see next bundle's <code>prior_question_elapsed_time</code> until we had already finished making our predictions for this bundle.</p>\n</blockquote>\n<p>I'm pretty sure this is on purpose, and exactly why they went with this prior_question_* scheme. I think knowing how long it took to answer a particular question would be considered \"data leakage\", since in a real-life setting you would be trying to predict this before you even gave the question to the user. </p>\n<p>(nevermind that our predictions would only be trained to predict this for the kinds of questions that <strong>the app's algorithm chooses for that particular user</strong>, and so would probably not generalize usefully to questions overall…) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1103923,
          "author_name": "doctorkael",
          "author_url": "",
          "post_date": "12/06/2020 12:28:53",
          "content": "<p><a href=\"https://www.kaggle.com/yanamal\" target=\"_blank\">@yanamal</a> </p>\n<pre><code>This won't actually achieve the desired result 100% of the time with bundles of size &gt; 1, since you're shifting by 1, and therefore shifting to the next question, which may still be in the same bundle.\n</code></pre>\n<p>We do an average precisely to deal with those bundles with size &gt; 1. </p>\n<p>Given that for a single bundle, the contents can have only one unique pq_elapsed_time and one unique pq_had_explanation, doing an average (n times a value, divided by n gives the same value), shifting it on task_container_id and merging it back on user_id and task_container_id would always ensure that the pq_* are properly mapped (based on task_container_id in this case).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1104084,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/06/2020 15:38:27",
          "content": "<p><a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a> <br>\n1) i dint understand how this constraint putting bottleneck in knowing prior time elapsed details</p>\n<p><code>next bundle's prior_question_elapsed_time until we had already finished making our predictions for this bundle.</code><br>\nI know that at a time we will have only group or one bundle(correct me if it can be more than one bundle also at a time in iteration)  to be predicted, so when we go to next group we have got now the time elapsed for previous bundle to use </p>\n<p>2) If i know correctly Set of users between Train/Valid and Test going to be same ?</p>\n<p>3) Question ids /Topics etc could always be different right ? so in that case one cant use the TIme elapsed data from the Train set right ? if its case one will have to accumulate this data as the prediction happens  for subsequent group ?</p>\n<p>Sorry if i sound very redudant ,its been only few days i got back to competition, currently working SAINT base not plus (partial features).. have achieved some success  in terms of  better CV score than SAKT 74.71 compared to 74.56 of SAKT. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1104308,
          "author_name": "yanamal",
          "author_url": "",
          "post_date": "12/06/2020 19:46:07",
          "content": "<p><a href=\"https://www.kaggle.com/doktorkael\" target=\"_blank\">@doktorkael</a> Then I don't understand what you mean by \"Do a shift operation after doing a groupby on user_id alone\". How are you \"shifting it on task_container_id\" if you're only grouping by user_id?</p>\n<p>Maybe we're actually talking about doing the same thing, after all (in terms of the shifting, anyway)</p>\n<p>Also, like I said, averaging will not change the data in any way at all, because it's already averaged across the bundles/task_containers. The data tab of the challenge says:</p>\n<blockquote>\n  <p><code>prior_question_elapsed_time</code>: (float32) The <strong>average</strong> time in milliseconds it took a user to answer each question <strong>in the previous question bundle</strong></p>\n</blockquote>\n<p>Granted, their use of the word \"bundle\" is rather confusing, but elsewhere they've clarified that they are talking about the user+task_conainer groupings. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1105775,
          "author_name": "doctorkael",
          "author_url": "",
          "post_date": "12/08/2020 07:29:34",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a>, from the discussion <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106\" target=\"_blank\">here</a>:</p>\n<pre><code>Correction- the hidden test set contains new users but not new questions.\n</code></pre>\n<p>For this reason, content based features would require very little updation. It is user based feature that would need to be updated. And yes we would only be seeing one bundle per user during the test inference. Still it doesn't mean pq_* features are entirely useless. Like <a href=\"https://www.kaggle.com/calebeverett\" target=\"_blank\">@calebeverett</a> mentioned, we can perform some aggregation and craft useful features out of them :)</p>\n<hr>\n<p><a href=\"https://www.kaggle.com/yanamal\" target=\"_blank\">@yanamal</a>, <code>How are you \"shifting it on task_container_id\" if you're only grouping by user_id?</code></p>\n<p>Groupby by default sorts the dataframe after the aggregation is done. So doing a groupby using user_id and task_container_id would ensure that the resulting dataframe would be sorted by task_container_id per user. After this it's simply a matter of shifting them and re-merging them back to the dataframe.</p>\n<pre><code>Also, like I said, averaging will not change the data in any way at all, because it's already averaged across the bundles/task_containers.\n</code></pre>\n<p>The purpose of averaging is to ensure that there is only one entry per user_id and task_container_id. Agreed, it doesn't change the data in any meaningful way. This is done just to make the shift operation easier. Hope this helps :)</p>",
          "votes": null,
          "replies": [
            {
              "id": 1106099,
              "author_name": "jaideepvalani",
              "author_url": "",
              "post_date": "12/08/2020 14:17:59",
              "content": "<blockquote>\n  <p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a>, from the discussion <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106\" target=\"_blank\">here</a>:</p>\n<pre><code>Correction- the hidden test set contains new users but not new questions.\n</code></pre>\n  <p>For this reason, content based features would require very little updation. It is user based feature that would need to be updated. And yes we would only be seeing one bundle per user during the test inference. Still it doesn't mean pq_* features are entirely useless. Like <a href=\"https://www.kaggle.com/calebeverett\" target=\"_blank\">@calebeverett</a> mentioned, we can perform some aggregation and craft useful features out of them :)</p>\n  <hr>\n  <p><a href=\"https://www.kaggle.com/yanamal\" target=\"_blank\">@yanamal</a>, <code>How are you \"shifting it on task_container_id\" if you're only grouping by user_id?</code></p>\n  <p>Groupby by default sorts the dataframe after the aggregation is done. So doing a groupby using user_id and task_container_id would ensure that the resulting dataframe would be sorted by task_container_id per user. After this it's simply a matter of shifting them and re-merging them back to the dataframe.</p>\n<pre><code>Also, like I said, averaging will not change the data in any way at all, because it's already averaged across the bundles/task_containers.\n</code></pre>\n  <p>The purpose of averaging is to ensure that there is only one entry per user_id and task_container_id. Agreed, it doesn't change the data in any meaningful way. This is done just to make the shift operation easier. Hope this helps :)</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a> <br>\nDo you mean of this sort<br>\n<code>train_df.groupby(['user_id','task_container_id']).agg(mpelt=('prior_question_elapsed_time','mean')).reset_index()</code> <br>\nfollowed by shift on Task container id and joining it back ?</p>\n<p>if its so then not sure data for this user complying to the assumption</p>\n<pre><code>user_id    task_container_id   elapsed time\n115    0   55000.0\n1    NaN\n2    37000.0\n3    19000.0\n4    11000.0\n</code></pre>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1105806,
          "author_name": "yanamal",
          "author_url": "",
          "post_date": "12/08/2020 08:08:36",
          "content": "<p>Ohh, I see now, thanks. So I think that is pretty similar to what I was doing, except I did unique() instead of averaging, which hopefully would produce the same results, as long as the prior_* fields are actually the same within the task_container.</p>\n<p>…. and also apparently I forgot to group by user in my notebook. details, details.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1105903,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/08/2020 10:10:58",
          "content": "<p><a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a> thanku.. if one is not using Aggregation features.. like following approach as there in SAINT /SAKT paper . <br>\nI was thinking how to handle the new users  in that case .Below is excerpt from SAKT inference part.</p>\n<p>if user_id in self.samples.index:<br>\n            q_, qa_ = self.samples[user_id]</p>\n<pre><code>        seq_len = len(q_)\n\n        if seq_len &gt;= self.max_seq:\n            q = q_[-self.max_seq:]\n            qa = qa_[-self.max_seq:]\n        else:\n            q[-seq_len:] = q_\n            qa[-seq_len:] = qa_          \n\n    x = np.zeros(self.max_seq-1, dtype=int)\n    x = q[1:].copy()\n    x += (qa[1:] == 1) * self.n_skill\n\n    questions = np.append(q[2:], [target_id])\n</code></pre>\n<p>Here if  suppose user id not present  then ?</p>\n<p>And i dint understand indexings 1:,2:  used , compared to :-1 for x and 1: for training</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1127503,
          "author_name": "calebeverett",
          "author_url": "",
          "post_date": "12/26/2020 15:32:43",
          "content": "<p>Hi,</p>\n<p>I believe from prior discussion the way to think about timestamp is that it when the user starts the bundle of questions and prior_question_elapsed_time is how long they took to complete it.</p>\n<p>There are cases where the task_container_id is not in the same order as timestamp (I counted approximately 2.5 million instances over the full training set), apparently because users may start questions on different devices and complete them in a different order than they started them. I ended up renumbering the task_container_ids so they were in the same order as the timestamps.</p>",
          "votes": null,
          "replies": [
            {
              "id": 1127511,
              "author_name": "jaideepvalani",
              "author_url": "",
              "post_date": "12/26/2020 15:46:16",
              "content": "<blockquote>\n  <p>Hi,</p>\n  <p>I believe from prior discussion the way to think about timestamp is that it when the user starts the bundle of questions and prior_question_elapsed_time is how long they took to complete it.</p>\n  <p>There are cases where the task_container_id is not in the same order as timestamp (I counted approximately 2.5 million instances over the full training set), apparently because users may start questions on different devices and complete them in a different order than they started them. I ended up renumbering the task_container_ids so they were in the same order as the timestamps.</p>\n</blockquote>\n<p>Thanks , if suppose  we have this series <br>\nTs1 Q1 t3  t1   <br>\nTs2 Q2 t5  t3  pet2</p>\n<p>How renumbering help here </p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1127548,
          "author_name": "calebeverett",
          "author_url": "",
          "post_date": "12/26/2020 16:24:46",
          "content": "<p>Hi, not sure I totally followed the abbreviations, but those look like they are in order. I was doing some calculations based on differences in task container ids that required them to be in order to return the correct differences. If you just wanted the rows in the right order, I think you could just use timestamp.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1127556,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/26/2020 16:31:13",
          "content": "<p><a href=\"https://www.kaggle.com/calebeverett\" target=\"_blank\">@calebeverett</a> at time seq 1 completed task container id T3  and at ts2  completed T5 .. prior elapsed time  which t5 says is for T4 or T3..<br>\nis it the case that falls in 2.5 Million interactions.. ?</p>\n<p>basically if i pass Q1,q2,q3 seq to model then prior elapsed time in their respective r  interaction rows should refer to  that of question appearing  prior  in above  sequence . does time sort ensures same ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1127705,
          "author_name": "calebeverett",
          "author_url": "",
          "post_date": "12/26/2020 19:08:04",
          "content": "<p>Ok, I see what you are asking - good question and I don't know the answer, i.e., does prior_question_elapsed_time relate to the interaction with the immediately preceding task_container_id or the immediately preceding timestamp?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1065758,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "10/31/2020 17:46:48",
      "content": "<p><a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a>  did u get clarity to it.. <br>\nPrevious means any question or series of questions asked to user  just prior to this   <br>\nSo if suppose<br>\nq1 - Correct<br>\nq2-Correct<br>\nq3-Incorrect<br>\nq4-correct<br>\nq5-Incorrect- prior explaination =True</p>\n<p>So   in this prior could be  Q4   as  may be student was trying to apply similar thought process to solve q5 but  response was incorrect so he thought let me check the explaination for q5..</p>\n<p><code>For the same student, the same questions were asked repeatedly maybe to verify if the student didn't get it right merely by chance.</code> how you know this</p>",
      "votes": null,
      "replies": [
        {
          "id": 1066011,
          "author_name": "doctorkael",
          "author_url": "",
          "post_date": "11/01/2020 07:29:42",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a>! <code>prior_questions_had_explanation</code> seems like the questions <em>either came with explanations at the end or they didn't</em>. It would make sense however for the user to be able to check the answers for questions he needed more clarity on.  </p>\n<p>Just like <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> told I have started using <code>task_container_id</code> to map the <code>prior_question_*</code> features with the corresponding rows. A student seeing the explanation for a question that was answered incorrectly would certainly do better when the same question gets asked again. Verifying this assumption for a few users by manual inspection has indeed proved correct.</p>\n<p>Some students were asked the same question a lot of times to the point that it feels like there might have been a mistake. Take user_id <code>15632472</code>. He was asked this question_id <code>1796</code> for 30 times. This student seems to be more of an anomaly though.</p>\n<p>The general trend however is that roughly <em>11% of the entire questions asked</em> were questions that had already been asked to the same user (More repeated questions for some users &amp; some with 0 repeated questions). The rationale for asking these questions again and again might have been to test the user's knowledge thoroughly not leaving things upto <em>chance</em>.</p>\n<p>PS: My analysis was only done on the first 1e6 rows. But I assume that this percentage to be roughly same for the entire dataset.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1065331": "`prior_question_elapsed_time` and `prior_question_had_explanation` seem kind of important to me. No discussion so far however, seems specify how these features need to be mapped with their corresponding rows.  \n\nThe description for pq_had_explanation reads: \n```\nWhether or not the user saw an explanation and the correct response(s) after answering the previous question bundle, ignoring any lectures in between. The value is shared across a single question bundle, and is null for a user's first question bundle or lecture.\n```\n\nIf by the term *question bundle* they refer to the `bundle_id`, I'm not really sure what *previous* question bundle would mean. Say for bundle_id 99 the `prior_question_had_explanation` is True. Then does it mean that the `bundle_id` 98 had explanation? Or does it refer to the bundle_id in the *previous position*? We should also note that the `bundle_id` is not served in any order, so the bundle_id in the previous position would be something random like 1902 or 0.\n\nIf however the term *previous question bundle* refers to the previous `task_container_id` (I strongly suspect this to the case) the task to map with corresponding row gets much simpler. \n\nFor the same student, the same questions were asked repeatedly maybe to verify if the student didn't get it right merely by *chance*. Some questions if answered wrong were asked until the student got it right. The `prior_question_had_explanation` if mapped correctly could be an important feature to determine if the students gets it right the second time around.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4175266%2Fa793109d08210eef931f22f1a3832ef8%2F1.png?generation=1604127572115869&alt=media)\n\nIt would be really nice if someone could help out with this!",
    "1065422": "> If however the term previous question bundle refers to the previous task_container_id (I strongly suspect this to the case) the task to map with corresponding row gets much simpler.\n\nI believe this is correct and how I've interpreted it.",
    "1065758": "doctorkael  did u get clarity to it.. \nPrevious means any question or series of questions asked to user  just prior to this   \nSo if suppose\nq1 - Correct\nq2-Correct\nq3-Incorrect\nq4-correct\nq5-Incorrect- prior explaination =True\n\nSo   in this prior could be  Q4   as  may be student was trying to apply similar thought process to solve q5 but  response was incorrect so he thought let me check the explaination for q5..\n\n`For the same student, the same questions were asked repeatedly maybe to verify if the student didn't get it right merely by chance.` how you know this",
    "1066011": "Hey @jaideepvalani! `prior_questions_had_explanation` seems like the questions *either came with explanations at the end or they didn't*. It would make sense however for the user to be able to check the answers for questions he needed more clarity on.  \n\nJust like @rohanrao told I have started using `task_container_id` to map the `prior_question_*` features with the corresponding rows. A student seeing the explanation for a question that was answered incorrectly would certainly do better when the same question gets asked again. Verifying this assumption for a few users by manual inspection has indeed proved correct.\n\nSome students were asked the same question a lot of times to the point that it feels like there might have been a mistake. Take user_id `15632472`. He was asked this question_id `1796` for 30 times. This student seems to be more of an anomaly though.\n\nThe general trend however is that roughly *11% of the entire questions asked* were questions that had already been asked to the same user (More repeated questions for some users & some with 0 repeated questions). The rationale for asking these questions again and again might have been to test the user's knowledge thoroughly not leaving things upto *chance*.\n\nPS: My analysis was only done on the first 1e6 rows. But I assume that this percentage to be roughly same for the entire dataset.",
    "1067743": "I've also understood it in the same manner. \n@doctorkael did you get any progress on the mapping ?",
    "1068211": "I haven't incorporated any of those features into a model yet. The huge data and the special API are intimidating to say the least. \n\nBut after further explorations, I found those features to be predictive of the answer correctness for repeated questions. It is indeed previous `task_container_id ` that the features need to be shifted on.",
    "1068321": "The question is : how to do it without exploding the processing time",
    "1092674": "Sorry, when you all say \"previous `task_container_id`\", what do you mean by \"previous\"? `task_container_id - 1`? Or whichever one appeared last in timestamp order for that user? because apparently, according to other posts, those are not necessarily the same (the `task_container_id` doesn't always monotonically increase with the timestamp).\n\nI'm trying to figure out how to decide whether a user saw feedback *for a given question*, since it's kind of important in modeling that user's state.",
    "1092743": "\"Previous\" `task_container_id` refers to `task_container_id - 1`. This is the where the features such as `prior_question_elapsed_time` and `prior_question_had_explanation` needs to be mapped to (ignoring any lectures that may come in between). \n\nCase 1: Without any lectures in between\nAssume a user has answered a question with `task_container_id` **1**. The same user had seen the explanation if prior_question_had_explanation for `task_container_id` **2** is True. \n\nCase 2: User saw a few lectures in between\nAssume a user has answered a question with `task_container_id` **1**. Suppose now that the user had seen two lectures in between with `task_container_id`'s 2 & 3. The same user had seen the explanation for the question (of task_container_id 1) if the prior_question_had_explanation for `task_container_id` **4** is True. _Note that prior_question_had_explanation for task_container_ids 2 and 3 (the lectures) would be NA and they can be safely dropped or ignored._",
    "1096922": "Are we sure that the mapping is from bundle with **task container id** to **task container id -1**.\n\nSince there are some issues in the order of task container id. I believe the mapping should be from **timestamp at current bundle** to **timestamp of previous bundle**.",
    "1099034": "> Since there are some issues in the order of task container id. I believe the mapping should be from **timestamp at current bundle** to **timestamp of previous bundle**.\n\nAfter some experimentation, I believe that \"previous\" is actually based on the timestamp order, not task container id order.\n\nI did manage to extract the elapsed time and had_explanation flag for the current question! And I'm pretty sure it's even correct. Here is my notebook (it also has some more text about why I think it's timestamp and not task_container_id that determines order):\nhttps://www.kaggle.com/yanamal/explanation-and-elapsed-time-for-current-question",
    "1099202": "Is task container id relevant other than as the number and order of bundles delivered to users?",
    "1099275": "It's relevant for the purpose of determining what \"previous\" means for \"previous question\", because \"previous question\" seems to actually be \"the question(s) in the task_container which was chronologically (by timestamp) served to the user most recently before this task_container\". If that makes any sense.\n\np.s. the notebook I mentioned above is slightly broken (it only identifies the elapsed time/had_expanation for the first thing in a task container), but I'm having a lot of trouble implementing a fix that will actually fit in RAM and not crash the kernel.",
    "1099434": "I wasn't able to do this for all users at the same time (RAM issues)\nI ended up looping over user_ids and updating one user at a time. It works but takes about 30-40 mins",
    "1100220": "Got it. That's what I was thinking as well, but wanted to make sure I wasn't missing some other important aspect. I've been having good luck with sqlite3 to keep processing times short.",
    "1100584": "doctorkael  how do we know which interaction is chronologically just prior to the Row which prior explaination flag as True ,does sorting by time stamp for the users ensures that any row having prior explaination flag as True refers to interaction appearing before it after sorted by TS",
    "1100590": "yanamal \nwhen you say \n`question\" seems to actually be \"the question(s) in the task_container which was chronologically (by timestamp) served to the user most recently before this task_container\". If that makes any sense.`\n\nthen here does chronology means sorting by timestamp for each user not whole data set offcourse\n i too felt that when we sort timestamp we get the interactions in the order in which user attempted those ..",
    "1100678": "Hey @calebeverett! `task_container_id` AFAIK doesn't seem to be that important by itself besides serving as a means to determine the order in which the users where served questions. We could use this feature to create better features that correlate to the target.\n\n***\n\n@yanamal @abdurrafae The issues with ordering is relatively minor. The vast majority of the users can be safely mapped using the `task_container_id`. So I think its an easier option to simply shift using user_id and task_container_id.  Steps to get this feature might be something like this: \n\n1. Sort the data by user_id and task_container_id (By default the order is user_id, timestamp).\n2. Groupby user_id, task_container_id, calculate the mean (Some questions come as bundles, mean will be the same for all user_id, task_container_id pair).\n3. Do a shift operation after doing a groupby on user_id alone. The sort operation we did before should ensure that the shift happens to task_container_id - 1.\n4. Merge the above result on user_id and task_container_id.\n\nThis feature does indeed have a high importance on the target. Sadly it cannot be used during prediction since during prediction, we will see *utmost one bundle for a single user* which means that we will not see next bundle's `prior_question_elapsed_time` until we had already finished making our predictions for this bundle.\n\n***\n\n@jaideepvalani, the rows are already sorted on user_id and timestamp. You can check this by doing:\n```\ndata.groupby(['user_id'])['timestamp'].is_monotonic_increasing.all()\n>>> True\n```",
    "1100995": "Thanks - even if `prior_question_elapsed_time` isn't available for the current question, perhaps there is something in an historical aggregation - by `user_id`, `user_id`-`content_id`, `user_id`-`tag` or  `user_id`-`part`.",
    "1101571": "Yes! That would certainly be possible. Better to do an aggregation based on content_id since user based features require the hassle of updating them as new users arrive.",
    "1103203": "jaideepvalani yeah, I meant sort by timestamp per user (so either group or sort by user)\n(also, that's already how the data comes sorted)\n\n@doctorkael \n\n> Do a shift operation after doing a groupby on user_id alone. The sort operation we did before should ensure that the shift happens to task_container_id - 1.\n\nThis won't actually achieve the desired result 100% of the time with bundles of size > 1, since you're shifting by 1, and therefore shifting to the next question, which may still be in the same bundle.\n\nMy current process is more like:\n1. don't bother sorting, because the existing sorting by user + timestamp works well enough (arguably better)\n2. don't bother averaging either, because I'm 99% sure the description says elapsed time is the same for every single question in the task_container, anyway.\n3. extract (into separate df) unique rows of `[user_id, task_container_id, prior_question_had_explanation, prior_question_elapsed_time]`\n4. shift `prior_question*` in *that* dataframe to create `this_question*`\n5. add these new rows back into the original dataframe (the indices will match up with first entry for each `[user, task_container_id]` combo)\n6. (can't do because of RAM-hugry pandas) group by `[user_id, task_container_id]` and ffill this_question_* to propagate the values within task_container\n\nI'm going to try numpy next, when I get around to it.\n\n> Sadly it cannot be used during prediction since during prediction, we will see *utmost one bundle for a single user* which means that we will not see next bundle's `prior_question_elapsed_time` until we had already finished making our predictions for this bundle.\n\nI'm pretty sure this is on purpose, and exactly why they went with this prior_question_* scheme. I think knowing how long it took to answer a particular question would be considered \"data leakage\", since in a real-life setting you would be trying to predict this before you even gave the question to the user. \n\n(nevermind that our predictions would only be trained to predict this for the kinds of questions that **the app's algorithm chooses for that particular user**, and so would probably not generalize usefully to questions overall...)",
    "1103923": "yanamal \n\n```\nThis won't actually achieve the desired result 100% of the time with bundles of size > 1, since you're shifting by 1, and therefore shifting to the next question, which may still be in the same bundle.\n```\n\nWe do an average precisely to deal with those bundles with size > 1. \n\nGiven that for a single bundle, the contents can have only one unique pq_elapsed_time and one unique pq_had_explanation, doing an average (n times a value, divided by n gives the same value), shifting it on task_container_id and merging it back on user_id and task_container_id would always ensure that the pq_* are properly mapped (based on task_container_id in this case).",
    "1104084": "doctorkael \n1) i dint understand how this constraint putting bottleneck in knowing prior time elapsed details\n\n`next bundle's prior_question_elapsed_time until we had already finished making our predictions for this bundle.`\nI know that at a time we will have only group or one bundle(correct me if it can be more than one bundle also at a time in iteration)  to be predicted, so when we go to next group we have got now the time elapsed for previous bundle to use \n\n2) If i know correctly Set of users between Train/Valid and Test going to be same ?\n\n3) Question ids /Topics etc could always be different right ? so in that case one cant use the TIme elapsed data from the Train set right ? if its case one will have to accumulate this data as the prediction happens  for subsequent group ?\n\nSorry if i sound very redudant ,its been only few days i got back to competition, currently working SAINT base not plus (partial features).. have achieved some success  in terms of  better CV score than SAKT 74.71 compared to 74.56 of SAKT.",
    "1104308": "doktorkael Then I don't understand what you mean by \"Do a shift operation after doing a groupby on user_id alone\". How are you \"shifting it on task_container_id\" if you're only grouping by user_id?\n\nMaybe we're actually talking about doing the same thing, after all (in terms of the shifting, anyway)\n\nAlso, like I said, averaging will not change the data in any way at all, because it's already averaged across the bundles/task_containers. The data tab of the challenge says:\n\n> `prior_question_elapsed_time`: (float32) The **average** time in milliseconds it took a user to answer each question **in the previous question bundle**\n\nGranted, their use of the word \"bundle\" is rather confusing, but elsewhere they've clarified that they are talking about the user+task_conainer groupings.",
    "1105775": "jaideepvalani, from the discussion [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106):\n\n```\nCorrection- the hidden test set contains new users but not new questions.\n```\n\nFor this reason, content based features would require very little updation. It is user based feature that would need to be updated. And yes we would only be seeing one bundle per user during the test inference. Still it doesn't mean pq_* features are entirely useless. Like @calebeverett mentioned, we can perform some aggregation and craft useful features out of them :)\n\n***\n@yanamal, ```How are you \"shifting it on task_container_id\" if you're only grouping by user_id?```\n\nGroupby by default sorts the dataframe after the aggregation is done. So doing a groupby using user_id and task_container_id would ensure that the resulting dataframe would be sorted by task_container_id per user. After this it's simply a matter of shifting them and re-merging them back to the dataframe.\n\n```\nAlso, like I said, averaging will not change the data in any way at all, because it's already averaged across the bundles/task_containers.\n```\n\nThe purpose of averaging is to ensure that there is only one entry per user_id and task_container_id. Agreed, it doesn't change the data in any meaningful way. This is done just to make the shift operation easier. Hope this helps :)",
    "1105806": "Ohh, I see now, thanks. So I think that is pretty similar to what I was doing, except I did unique() instead of averaging, which hopefully would produce the same results, as long as the prior_* fields are actually the same within the task_container.\n\n.... and also apparently I forgot to group by user in my notebook. details, details.",
    "1105903": "doctorkael thanku.. if one is not using Aggregation features.. like following approach as there in SAINT /SAKT paper . \nI was thinking how to handle the new users  in that case .Below is excerpt from SAKT inference part.\n\n if user_id in self.samples.index:\n            q_, qa_ = self.samples[user_id]\n            \n            seq_len = len(q_)\n\n            if seq_len >= self.max_seq:\n                q = q_[-self.max_seq:]\n                qa = qa_[-self.max_seq:]\n            else:\n                q[-seq_len:] = q_\n                qa[-seq_len:] = qa_          \n        \n        x = np.zeros(self.max_seq-1, dtype=int)\n        x = q[1:].copy()\n        x += (qa[1:] == 1) * self.n_skill\n        \n        questions = np.append(q[2:], [target_id])\n\n\nHere if  suppose user id not present  then ?\n\nAnd i dint understand indexings 1:,2:  used , compared to :-1 for x and 1: for training",
    "1106099": "> @jaideepvalani, from the discussion [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106):\n> \n> ```\n> Correction- the hidden test set contains new users but not new questions.\n> ```\n> \n> For this reason, content based features would require very little updation. It is user based feature that would need to be updated. And yes we would only be seeing one bundle per user during the test inference. Still it doesn't mean pq_* features are entirely useless. Like @calebeverett mentioned, we can perform some aggregation and craft useful features out of them :)\n> \n> ***\n> @yanamal, ```How are you \"shifting it on task_container_id\" if you're only grouping by user_id?```\n> \n> Groupby by default sorts the dataframe after the aggregation is done. So doing a groupby using user_id and task_container_id would ensure that the resulting dataframe would be sorted by task_container_id per user. After this it's simply a matter of shifting them and re-merging them back to the dataframe.\n> \n> ```\n> Also, like I said, averaging will not change the data in any way at all, because it's already averaged across the bundles/task_containers.\n> ```\n> \n> The purpose of averaging is to ensure that there is only one entry per user_id and task_container_id. Agreed, it doesn't change the data in any meaningful way. This is done just to make the shift operation easier. Hope this helps :)\n\n@doctorkael \nDo you mean of this sort\n`train_df.groupby(['user_id','task_container_id']).agg(mpelt=('prior_question_elapsed_time','mean')).reset_index() ` \nfollowed by shift on Task container id and joining it back ?\n\nif its so then not sure data for this user complying to the assumption\n\n```\nuser_id\ttask_container_id\telapsed time\n115\t0\t55000.0\n1\tNaN\n2\t37000.0\n3\t19000.0\n4\t11000.0\n```",
    "1126983": "> Got it. That's what I was thinking as well, but wanted to make sure I wasn't missing some other important aspect. I've been having good luck with sqlite3 to keep processing times short.\n\n@calebeverett \nDoes  it  discussion here suggests that suppose  if order of question in timestamp order\nis q1,q2,q3,q4\nq1 T2\nQ2 T2\nQ3 T2\nQ4 T4\nQ5-T4\nprior time elapsed time  in Q3,Q4,Q5  refers to time student took to finish questions in Task container id T3 and not tast container T2 ?\n\n2) If above is case ,can not the user just sort by task container id  each user , and then serve to model ,that way prior time is aligned automatically",
    "1127503": "Hi,\n\nI believe from prior discussion the way to think about timestamp is that it when the user starts the bundle of questions and prior_question_elapsed_time is how long they took to complete it.\n\nThere are cases where the task_container_id is not in the same order as timestamp (I counted approximately 2.5 million instances over the full training set), apparently because users may start questions on different devices and complete them in a different order than they started them. I ended up renumbering the task_container_ids so they were in the same order as the timestamps.",
    "1127511": "> Hi,\n> \n> I believe from prior discussion the way to think about timestamp is that it when the user starts the bundle of questions and prior_question_elapsed_time is how long they took to complete it.\n> \n> There are cases where the task_container_id is not in the same order as timestamp (I counted approximately 2.5 million instances over the full training set), apparently because users may start questions on different devices and complete them in a different order than they started them. I ended up renumbering the task_container_ids so they were in the same order as the timestamps.\n\nThanks , if suppose  we have this series \nTs1 Q1 t3  t1   \nTs2 Q2 t5  t3  pet2\n\nHow renumbering help here",
    "1127548": "Hi, not sure I totally followed the abbreviations, but those look like they are in order. I was doing some calculations based on differences in task container ids that required them to be in order to return the correct differences. If you just wanted the rows in the right order, I think you could just use timestamp.",
    "1127556": "calebeverett at time seq 1 completed task container id T3  and at ts2  completed T5 .. prior elapsed time  which t5 says is for T4 or T3..\nis it the case that falls in 2.5 Million interactions.. ?\n\nbasically if i pass Q1,q2,q3 seq to model then prior elapsed time in their respective r  interaction rows should refer to  that of question appearing  prior  in above  sequence . does time sort ensures same ?",
    "1127705": "Ok, I see what you are asking - good question and I don't know the answer, i.e., does prior_question_elapsed_time relate to the interaction with the immediately preceding task_container_id or the immediately preceding timestamp?"
  },
  "source": "meta"
}