{
  "id": 207434,
  "title": "Temporal Features pt2",
  "url": "/competitions/riiid-test-answer-prediction/discussion/207434",
  "author_name": "",
  "post_date": "2020-12-29T17:47:10.806104500Z",
  "votes": 13,
  "comment_count": 32,
  "views": 0,
  "content": "<p>Very early on in this competition <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> posted this thread: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189351\" target=\"_blank\">Questions about timestamp and prior_question_elapsed_time</a>, which has been on my mind non-stop. Earlier, I made <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/205515\" target=\"_blank\">this posting</a> where I argued it wasn't possible to compute <code>lag_time</code> as defined in the SAINT+ paper. But after calmly investing more time into it, I've discovered that it actually is quite possible to calculate <code>lag_time</code>, at least at the <code>{user_id, task_container_id}</code> resolution. it's still a bit noisy for a few records, but the number of discrepancies in the final features is still on the order of 1-2% of 100M.</p>\n<p>The way that it works is this:</p>\n<pre><code>s e    s e     s e    s e\n</code></pre>\n<p>Imagine you have a series of start-end pairs, corresponding to when the exercise was first opened by the student and when the student completed their submission.</p>\n<p>We know that <code>e</code> = <code>timestamp</code>, as defined in the dataset, the time when the interaction concluded.</p>\n<p>We are also given prior_question_elapsed_time (which is actually average prior_task_container_elapsed_time, per-user), which is equivalent to <code>e-s</code> for any given single question task_container_id. In the case there are multiple items, just multiply <code>prior_question_elapsed_time</code> by the number of items in the user's task container id.</p>\n<p>The value we are trying to compute is <code>s-e</code>. We have all the pieces.</p>\n<p>Assuming we maintain a feature ts_diff = ts - ts_previous, which is corrected s.t. all values in the bundle share the first ts_diff, then:</p>\n<p>The we can calculate for the <strong>current</strong> container a feature: <code>prior_container_lag_time = prior_container_ts_diff - prior_container_elapsed_time * prior_container_num_items</code></p>\n<p>Visually, that would be (e_x - e_(x-1)) - (e_x - s_x). Just make sure all units are ms, or s.</p>\n<p>To my great surprise, adding this feature did not result in a score boost. Rather, it leads to over-fitting. Not leakage, but overfitting. How can having the lag time of the previous container/bundle result in overfitting of the train set? This is what the distribution looks like on 80M rows after being log transformed and standard scaled. There is a faux peak at 0 because I'm filling some nans. I've tried nan filling with min, max and mean (0).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2F4e76492a9e3aab0603ff19fee1e91cbb%2Findex.png?generation=1609264028033112&amp;alt=media\" alt=\"\"></p>\n<p>Any ideas?</p>",
  "messages": [
    {
      "id": "1131395",
      "postDate": "12/29/2020 17:47:10",
      "content": "<p>Very early on in this competition <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> posted this thread: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189351\" target=\"_blank\">Questions about timestamp and prior_question_elapsed_time</a>, which has been on my mind non-stop. Earlier, I made <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/205515\" target=\"_blank\">this posting</a> where I argued it wasn't possible to compute <code>lag_time</code> as defined in the SAINT+ paper. But after calmly investing more time into it, I've discovered that it actually is quite possible to calculate <code>lag_time</code>, at least at the <code>{user_id, task_container_id}</code> resolution. it's still a bit noisy for a few records, but the number of discrepancies in the final features is still on the order of 1-2% of 100M.</p>\n<p>The way that it works is this:</p>\n<pre><code>s e    s e     s e    s e\n</code></pre>\n<p>Imagine you have a series of start-end pairs, corresponding to when the exercise was first opened by the student and when the student completed their submission.</p>\n<p>We know that <code>e</code> = <code>timestamp</code>, as defined in the dataset, the time when the interaction concluded.</p>\n<p>We are also given prior_question_elapsed_time (which is actually average prior_task_container_elapsed_time, per-user), which is equivalent to <code>e-s</code> for any given single question task_container_id. In the case there are multiple items, just multiply <code>prior_question_elapsed_time</code> by the number of items in the user's task container id.</p>\n<p>The value we are trying to compute is <code>s-e</code>. We have all the pieces.</p>\n<p>Assuming we maintain a feature ts_diff = ts - ts_previous, which is corrected s.t. all values in the bundle share the first ts_diff, then:</p>\n<p>The we can calculate for the <strong>current</strong> container a feature: <code>prior_container_lag_time = prior_container_ts_diff - prior_container_elapsed_time * prior_container_num_items</code></p>\n<p>Visually, that would be (e_x - e_(x-1)) - (e_x - s_x). Just make sure all units are ms, or s.</p>\n<p>To my great surprise, adding this feature did not result in a score boost. Rather, it leads to over-fitting. Not leakage, but overfitting. How can having the lag time of the previous container/bundle result in overfitting of the train set? This is what the distribution looks like on 80M rows after being log transformed and standard scaled. There is a faux peak at 0 because I'm filling some nans. I've tried nan filling with min, max and mean (0).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2F4e76492a9e3aab0603ff19fee1e91cbb%2Findex.png?generation=1609264028033112&amp;alt=media\" alt=\"\"></p>\n<p>Any ideas?</p>",
      "rawMarkdown": "Very early on in this competition @gunesevitan posted this thread: [Questions about timestamp and prior_question_elapsed_time](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189351), which has been on my mind non-stop. Earlier, I made [this posting](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/205515) where I argued it wasn't possible to compute `lag_time` as defined in the SAINT+ paper. But after calmly investing more time into it, I've discovered that it actually is quite possible to calculate `lag_time`, at least at the `{user_id, task_container_id}` resolution. it's still a bit noisy for a few records, but the number of discrepancies in the final features is still on the order of 1-2% of 100M.\n\nThe way that it works is this:\n\n```\ns e    s e     s e    s e\n```\n\nImagine you have a series of start-end pairs, corresponding to when the exercise was first opened by the student and when the student completed their submission.\n\nWe know that `e` = `timestamp`, as defined in the dataset, the time when the interaction concluded.\n\nWe are also given prior_question_elapsed_time (which is actually average prior_task_container_elapsed_time, per-user), which is equivalent to `e-s` for any given single question task_container_id. In the case there are multiple items, just multiply `prior_question_elapsed_time` by the number of items in the user's task container id.\n\nThe value we are trying to compute is `s-e`. We have all the pieces.\n\nAssuming we maintain a feature ts_diff = ts - ts_previous, which is corrected s.t. all values in the bundle share the first ts_diff, then:\n\nThe we can calculate for the **current** container a feature: `prior_container_lag_time = prior_container_ts_diff - prior_container_elapsed_time * prior_container_num_items`\n\nVisually, that would be (e_x - e_(x-1)) - (e_x - s_x). Just make sure all units are ms, or s.\n\nTo my great surprise, adding this feature did not result in a score boost. Rather, it leads to over-fitting. Not leakage, but overfitting. How can having the lag time of the previous container/bundle result in overfitting of the train set? This is what the distribution looks like on 80M rows after being log transformed and standard scaled. There is a faux peak at 0 because I'm filling some nans. I've tried nan filling with min, max and mean (0).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2F4e76492a9e3aab0603ff19fee1e91cbb%2Findex.png?generation=1609264028033112&alt=media)\n\nAny ideas?",
      "votes": null
    },
    {
      "id": "1132093",
      "postDate": "12/30/2020 06:58:39",
      "content": "<p>I also tried to calculate the lag_time using the way you mentioned. When I saw the negative values, I dropped it and started only using timestamp(t) - timestamp(t-1) as lag_time/response_time. I am yet to debug my SAINT+ model since it's not training. If it works, I will share my results. </p>",
      "rawMarkdown": "I also tried to calculate the lag_time using the way you mentioned. When I saw the negative values, I dropped it and started only using timestamp(t) - timestamp(t-1) as lag_time/response_time. I am yet to debug my SAINT+ model since it's not training. If it works, I will share my results.",
      "votes": null
    },
    {
      "id": "1132101",
      "postDate": "12/30/2020 07:02:52",
      "content": "<p>claverru has discussed about the <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632#1107391\" target=\"_blank\">same here</a>. He is only using timestamp(t) - timestamp(t-1), capping it to 1 day and scaling it to [0, 1]</p>",
      "rawMarkdown": "claverru has discussed about the [same here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632#1107391). He is only using timestamp(t) - timestamp(t-1), capping it to 1 day and scaling it to [0, 1]",
      "votes": null
    },
    {
      "id": "1132293",
      "postDate": "12/30/2020 09:40:33",
      "content": "<p>After changing that approach I had a boost in my score. Tip: train with negatives.</p>",
      "rawMarkdown": "After changing that approach I had a boost in my score. Tip: train with negatives.",
      "votes": null
    },
    {
      "id": "1132393",
      "postDate": "12/30/2020 11:26:59",
      "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>   getting confused did  first approach gaveu more lb score or second one i </p>\n<p>secondly.. <br>\nhow many class minutes should one consider for categorizing the lag time durations if in minutes</p>",
      "rawMarkdown": "claverru   getting confused did  first approach gaveu more lb score or second one i \n \n\nsecondly.. \nhow many class minutes should one consider for categorizing the lag time durations if in minutes",
      "votes": null
    },
    {
      "id": "1132407",
      "postDate": "12/30/2020 11:43:50",
      "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> nice observation . Check our your distriution with previous best how do they look like. You should  get some insight.<br>\nSecondly one thing i still understand  if we are given only stop time stamp or completion time stamp. Why first task container solved by student is having Time stamp as 0  ,rather then its actual completion timestamp. Ours lag time approach is also simple so far but gets decent  saint plus score of 79.3</p>",
      "rawMarkdown": "authman nice observation . Check our your distriution with previous best how do they look like. You should  get some insight.\nSecondly one thing i still understand  if we are given only stop time stamp or completion time stamp. Why first task container solved by student is having Time stamp as 0  ,rather then its actual completion timestamp. Ours lag time approach is also simple so far but gets decent  saint plus score of 79.3",
      "votes": null
    },
    {
      "id": "1132455",
      "postDate": "12/30/2020 12:22:11",
      "content": "<p>If you are getting negatives in your timestamp diff, it is because you are not handleing the fact that you have some simultaneous timestamps due to questions in the same bundle. </p>\n<p>If after solving the prior bug, you try to compute the lag, you will still have negatives. This is something counterintuitive that has been dicussed in other posts. I suggest you to think about it. </p>\n<p>You have a good position in the LB, maybe playing around this leads you to gold or higher.</p>",
      "rawMarkdown": "If you are getting negatives in your timestamp diff, it is because you are not handleing the fact that you have some simultaneous timestamps due to questions in the same bundle. \n\nIf after solving the prior bug, you try to compute the lag, you will still have negatives. This is something counterintuitive that has been dicussed in other posts. I suggest you to think about it. \n\nYou have a good position in the LB, maybe playing around this leads you to gold or higher.",
      "votes": null
    },
    {
      "id": "1132473",
      "postDate": "12/30/2020 12:35:32",
      "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>   thanks.. may be i remember one post about the calculation that talked how we should handle repeat Tss for q in bundle. </p>",
      "rawMarkdown": "claverru   thanks.. may be i remember one post about the calculation that talked how we should handle repeat Tss for q in bundle.",
      "votes": null
    },
    {
      "id": "1132634",
      "postDate": "12/30/2020 14:56:17",
      "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> are you embedding this lag feature categorically or with a dense? If categorically, are you simply windsorizing the long tails?</p>",
      "rawMarkdown": "claverru are you embedding this lag feature categorically or with a dense? If categorically, are you simply windsorizing the long tails?",
      "votes": null
    },
    {
      "id": "1132785",
      "postDate": "12/30/2020 17:27:13",
      "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> </p>\n<p><code>The we can calculate for the current container a feature: prior_container_lag_time = prior_container_ts_diff - prior_container_elapsed_time * prior_container_num_items</code></p>\n<p>I again get confused looking at paper time line diagram</p>\n<p>it says<br>\nE1- start time T1 - We dont know<br>\nR1 (Response submitted time) - TIme-T2 - We know it as TS1</p>\n<p>Lagtime Lt1-  To calculate</p>\n<p>E2-Start TIme T3- We dont know </p>\n<p>R2-Stop Time T4 - We know TS2 </p>\n<p>APplying your formula</p>\n<p>Ts2-Ts1 or T4-T2 = Will give T2 to T4 .. in between we dont know T3</p>\n<p>How your formula gives  LT correctly with one known factor</p>",
      "rawMarkdown": "authman \n\n`The we can calculate for the current container a feature: prior_container_lag_time = prior_container_ts_diff - prior_container_elapsed_time * prior_container_num_items`\n\nI again get confused looking at paper time line diagram\n\nit says\nE1- start time T1 - We dont know\nR1 (Response submitted time) - TIme-T2 - We know it as TS1\n\nLagtime Lt1-  To calculate\n\nE2-Start TIme T3- We dont know \n\nR2-Stop Time T4 - We know TS2 \n\nAPplying your formula\n\nTs2-Ts1 or T4-T2 = Will give T2 to T4 .. in between we dont know T3\n\nHow your formula gives  LT correctly with one known factor",
      "votes": null
    },
    {
      "id": "1132839",
      "postDate": "12/30/2020 18:00:51",
      "content": "<p>No worries. Give me a second I'll draw it on paper.</p>",
      "rawMarkdown": "No worries. Give me a second I'll draw it on paper.",
      "votes": null
    },
    {
      "id": "1132843",
      "postDate": "12/30/2020 18:05:01",
      "content": "<p>Sure  taking this info also </p>\n<p><code>The timestamp column shows when an activity is finished, not when it started. A user could have spent an arbitrary amount of time working on the first problem before the first timestamp value would get logged.</code></p>",
      "rawMarkdown": "Sure  taking this info also \n\n`The timestamp column shows when an activity is finished, not when it started. A user could have spent an arbitrary amount of time working on the first problem before the first timestamp value would get logged.`",
      "votes": null
    },
    {
      "id": "1132858",
      "postDate": "12/30/2020 18:22:02",
      "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> Continuous time features</p>",
      "rawMarkdown": "authman Continuous time features",
      "votes": null
    },
    {
      "id": "1132888",
      "postDate": "12/30/2020 19:00:51",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2Fe0602f275397f69e833cd9c65825680e%2FIMG_6084.jpg?generation=1609354808793957&amp;alt=media\" alt=\"\"></p>\n<p>This is the case when the previous task container has just item. If they have multiple items, just multiply previous_elapsed_time by the number of items in the bundle.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2Fe0602f275397f69e833cd9c65825680e%2FIMG_6084.jpg?generation=1609354808793957&alt=media)\n\nThis is the case when the previous task container has just item. If they have multiple items, just multiply previous_elapsed_time by the number of items in the bundle.",
      "votes": null
    },
    {
      "id": "1133188",
      "postDate": "12/31/2020 02:34:10",
      "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a>  thanks elapsed time which you subtract  is then  of task c id 3 ?</p>\n<p>So from you paper draw  lag of last task container<br>\n= last time stamp  - 2nd last tstamp - last elapsed time <br>\n? </p>\n<p>If this is final case then we need a info  one time step back , and one time step after ( elt of current)  current time step to know current tc lag time . <br>\n<a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>  During inference how we manage if above is  method to get right lagtime for test q  </p>",
      "rawMarkdown": "authman  thanks elapsed time which you subtract  is then  of task c id 3 ?\n\nSo from you paper draw  lag of last task container\n= last time stamp  - 2nd last tstamp - last elapsed time \n? \n\nIf this is final case then we need a info  one time step back , and one time step after ( elt of current)  current time step to know current tc lag time . \n@claverru  During inference how we manage if above is  method to get right lagtime for test q",
      "votes": null
    },
    {
      "id": "1133200",
      "postDate": "12/31/2020 02:46:11",
      "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> That's a great tip. Looking at my score vs gold zone solutions I cant help but feel im doing something very wrong.. But after what you metioned I think I'm beginning to understand what i am doing wrong</p>",
      "rawMarkdown": "claverru That's a great tip. Looking at my score vs gold zone solutions I cant help but feel im doing something very wrong.. But after what you metioned I think I'm beginning to understand what i am doing wrong",
      "votes": null
    },
    {
      "id": "1133221",
      "postDate": "12/31/2020 03:13:26",
      "content": "<blockquote>\n  <p>= last time stamp - 2nd last tstamp - last elapsed time </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> correct. And elapsed time should be of the previous container, and luckily, they actually give us exactly that at time \"we are here!\".</p>\n<blockquote>\n  <p>If this is final case then we need a info one time step back , and one time step after ( elt of current) current time step to know current tc lag time . </p>\n</blockquote>\n<p>We will never know the current lag time for a bundle at inference time, similar to how we will never know the current start time of a bundle at inference. Only when we move to the next bundle is it possible for us to rectify previous bundle: start time, lag time, elapsed time. And that shouldn't <em>really</em> be an issue anyway since saint records are shifted due to the starter tag. But even if they weren't, it just means gotta use what you got.</p>",
      "rawMarkdown": "> = last time stamp - 2nd last tstamp - last elapsed time \n\n@jaideepvalani correct. And elapsed time should be of the previous container, and luckily, they actually give us exactly that at time \"we are here!\".\n\n> If this is final case then we need a info one time step back , and one time step after ( elt of current) current time step to know current tc lag time . \n\nWe will never know the current lag time for a bundle at inference time, similar to how we will never know the current start time of a bundle at inference. Only when we move to the next bundle is it possible for us to rectify previous bundle: start time, lag time, elapsed time. And that shouldn't *really* be an issue anyway since saint records are shifted due to the starter tag. But even if they weren't, it just means gotta use what you got.",
      "votes": null
    },
    {
      "id": "1133260",
      "postDate": "12/31/2020 04:09:25",
      "content": "<p>yes but while predictin ans for test q, it will give us elt of last q of users train data series.. we will have to dynamically calculate lag time for  last q, for model to see lag time while predicting ans correctness for test q at tail</p>\n<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>  how many negative values approx one has for user id/tc id</p>",
      "rawMarkdown": "yes but while predictin ans for test q, it will give us elt of last q of users train data series.. we will have to dynamically calculate lag time for  last q, for model to see lag time while predicting ans correctness for test q at tail\n\n@claverru  how many negative values approx one has for user id/tc id",
      "votes": null
    },
    {
      "id": "1134282",
      "postDate": "01/01/2021 05:16:34",
      "content": "<blockquote>\n  <blockquote>\n    <p>= last time stamp - 2nd last tstamp - last elapsed time </p>\n  </blockquote>\n  <p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> correct. And elapsed time should be of the previous container, and luckily, they actually give us exactly that at time \"we are here!\".</p>\n  <blockquote>\n    <p>If this is final case then we need a info one time step back , and one time step after ( elt of current) current time step to know current tc lag time . </p>\n  </blockquote>\n  <p>We will never know the current lag time for a bundle at inference time, similar to how we will never know the current start time of a bundle at inference. Only when we move to the next bundle is it possible for us to rectify previous bundle: start time, lag time, elapsed time. And that shouldn't <em>really</em> be an issue anyway since saint records are shifted due to the starter tag. But even if they weren't, it just means gotta use what you got.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> <br>\n<code>train_df[(train_df.user_id==13134) &amp; (train_df.timestamp&gt;12948521851-10000)]</code></p>\n<p>Will we get negative lag time <br>\nthis user id &amp; task container id 53., applying your formula</p>\n<pre><code>    timestamp   user_id content_id  content_type_id task_container_id   user_answer answered_correctly  prior_question_elapsed_time prior_question_had_explanation  bundle_count\n672    12948517301 13134   868 False   54  0   1   12000.0 1.0 1\n673    12948521851 13134   868 False   53  3   0   7000.0  1.0 1\n674    12948556717 13134   1136False   55  0   1   23000.0\n</code></pre>\n<p><br>\n12948521851-12948517301    -23000</p>",
      "rawMarkdown": "> > = last time stamp - 2nd last tstamp - last elapsed time \n> \n> @jaideepvalani correct. And elapsed time should be of the previous container, and luckily, they actually give us exactly that at time \"we are here!\".\n> \n> > If this is final case then we need a info one time step back , and one time step after ( elt of current) current time step to know current tc lag time . \n> \n> We will never know the current lag time for a bundle at inference time, similar to how we will never know the current start time of a bundle at inference. Only when we move to the next bundle is it possible for us to rectify previous bundle: start time, lag time, elapsed time. And that shouldn't *really* be an issue anyway since saint records are shifted due to the starter tag. But even if they weren't, it just means gotta use what you got.\n\n@authman \n`train_df[(train_df.user_id==13134) & (train_df.timestamp>12948521851-10000)]`\n\nWill we get negative lag time \nthis user id & task container id 53., applying your formula\n\n\n\n```\n\ttimestamp\tuser_id\tcontent_id\tcontent_type_id\ttask_container_id\tuser_answer\tanswered_correctly\tprior_question_elapsed_time\tprior_question_had_explanation\tbundle_count\n672\t12948517301\t13134\t868\tFalse\t54\t0\t1\t12000.0\t1.0\t1\n673\t12948521851\t13134\t868\tFalse\t53\t3\t0\t7000.0\t1.0\t1\n674\t12948556717\t13134\t1136False\t55\t0\t1\t23000.0\n``` \n12948521851-12948517301\t-23000",
      "votes": null
    },
    {
      "id": "1134461",
      "postDate": "01/01/2021 10:10:17",
      "content": "<p>I am currently using that feature personnaly, and it gives a boost to my model.<br>\nMy personnal formula is :</p>\n<blockquote>\n  <p>TS[-2] - TS[-3] - LAST_QUESTION_REACTION_TIME[-1]*batch_size[-1]</p>\n</blockquote>\n<p>I store those values and compute rolling averages.</p>\n<p>edit: changed sign as per authman comment</p>",
      "rawMarkdown": "I am currently using that feature personnaly, and it gives a boost to my model.\nMy personnal formula is :\n> TS[-2] - TS[-3] - LAST_QUESTION_REACTION_TIME[-1]*batch_size[-1]\n\nI store those values and compute rolling averages.\n\nedit: changed sign as per authman comment",
      "votes": null
    },
    {
      "id": "1134478",
      "postDate": "01/01/2021 10:28:17",
      "content": "<blockquote>\n  <p>I am currently using that feature personnaly, and it gives a boost to my model.<br>\n  My personnal formula is :</p>\n  <blockquote>\n    <p>TS[-2] - TS[-3] + LAST_QUESTION_REACTION_TIME[-1]*batch_size[-1]</p>\n  </blockquote>\n  <p>I store those values and compute rolling averages.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/bowaka\" target=\"_blank\">@bowaka</a> <br>\nits confusing a bit why u have taken TS[-3],TS[-2] while calculating lag time for current q.</p>\n<p>With your formula wat would be values while computing lag time for this user &amp; task container id 53..</p>\n<pre><code>timestamp   user_id content_id  content_type_id task_container_id   user_answer answered_correctly  prior_question_elapsed_time prior_question_had_explanation  bundle_count\n672    12948517301 13134   868 False   54  0   1   12000.0 1.0 1\n673    12948521851 13134   868 False   53  3   0   7000.0  1.0 1\n674    12948556717 13134   1136False   55  0   1   23000.0\n</code></pre>",
      "rawMarkdown": "> I am currently using that feature personnaly, and it gives a boost to my model.\n> My personnal formula is :\n> > TS[-2] - TS[-3] + LAST_QUESTION_REACTION_TIME[-1]*batch_size[-1]\n> \n> I store those values and compute rolling averages.\n\n@bowaka \nits confusing a bit why u have taken TS[-3],TS[-2] while calculating lag time for current q.\n\nWith your formula wat would be values while computing lag time for this user & task container id 53..\n ```\ntimestamp   user_id content_id  content_type_id task_container_id   user_answer answered_correctly  prior_question_elapsed_time prior_question_had_explanation  bundle_count\n672    12948517301 13134   868 False   54  0   1   12000.0 1.0 1\n673    12948521851 13134   868 False   53  3   0   7000.0  1.0 1\n674    12948556717 13134   1136False   55  0   1   23000.0\n```",
      "votes": null
    },
    {
      "id": "1134487",
      "postDate": "01/01/2021 10:48:07",
      "content": "<p>You need record 671 to do it.</p>",
      "rawMarkdown": "You need record 671 to do it.",
      "votes": null
    },
    {
      "id": "1134633",
      "postDate": "01/01/2021 12:52:21",
      "content": "<p>Did you get any gain out of this feature guys ? I didn't get any so far.</p>",
      "rawMarkdown": "Did you get any gain out of this feature guys ? I didn't get any so far.",
      "votes": null
    },
    {
      "id": "1134660",
      "postDate": "01/01/2021 13:20:20",
      "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> in the sequence  below  during the inference  we can compute lag time for Q3 container<br>\nusing Q3  ts, Q2 ts,and Q3 elapsed time given in Test4 . <br>\nlike User id U1=   Q1,Q2,Q3 Test4</p>\n<p>Q3's task container lag time= Q3 Ts-Q2Ts-elt in Q4</p>",
      "rawMarkdown": "authman in the sequence  below  during the inference  we can compute lag time for Q3 container\nusing Q3  ts, Q2 ts,and Q3 elapsed time given in Test4 . \nlike User id U1=   Q1,Q2,Q3 Test4\n\nQ3's task container lag time= Q3 Ts-Q2Ts-elt in Q4",
      "votes": null
    },
    {
      "id": "1134671",
      "postDate": "01/01/2021 13:32:32",
      "content": "<p>Sorry, I am not calculating for current q, this information is not accessible with the data. I calculate lag time between q-3 and q-2 (assuming current row is q-1)</p>\n<p>My turn to try a drawing of the situation 😄:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2Ff1090aaf2daccaad9845760f6b4800e9%2FSans%20titre.png?generation=1609514059944910&amp;alt=media\" alt=\"\"></p>\n<p>In the illustration above, BQ correspond to a same batch of question. They are served together through the API and timestamps are similar for all questions Qx in a batch BQy</p>\n<p>Note: as authman is saying, for your example, you need more rows as you didn't complete a full batch of questions yet.</p>\n<p>edit: change sign '+' in '-' as per authman comment below</p>",
      "rawMarkdown": "Sorry, I am not calculating for current q, this information is not accessible with the data. I calculate lag time between q-3 and q-2 (assuming current row is q-1)\n\nMy turn to try a drawing of the situation 😄:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2Ff1090aaf2daccaad9845760f6b4800e9%2FSans%20titre.png?generation=1609514059944910&alt=media)\n\nIn the illustration above, BQ correspond to a same batch of question. They are served together through the API and timestamps are similar for all questions Qx in a batch BQy\n\nNote: as authman is saying, for your example, you need more rows as you didn't complete a full batch of questions yet.\n\nedit: change sign '+' in '-' as per authman comment below",
      "votes": null
    },
    {
      "id": "1134769",
      "postDate": "01/01/2021 15:10:07",
      "content": "<p>Your equation is identical to mine <a href=\"https://www.kaggle.com/bowaka\" target=\"_blank\">@bowaka</a> except for two things.</p>\n<ol>\n<li>I haven't tried EMA features, yet 😉.</li>\n<li>Check your equation and illustration, the <code>+</code> should be a <code>-</code>. I hope this change results in a boost for you.</li>\n</ol>\n<p>Your illustration is a lot better suited than mine.</p>",
      "rawMarkdown": "Your equation is identical to mine @bowaka except for two things.\n\n1. I haven't tried EMA features, yet 😉.\n1. Check your equation and illustration, the `+` should be a `-`. I hope this change results in a boost for you.\n\nYour illustration is a lot better suited than mine.",
      "votes": null
    },
    {
      "id": "1134775",
      "postDate": "01/01/2021 15:14:00",
      "content": "<p>whoops ! you are totaly right, there was a mistake in the sign ! I'll update the figure</p>",
      "rawMarkdown": "whoops ! you are totaly right, there was a mistake in the sign ! I'll update the figure",
      "votes": null
    },
    {
      "id": "1134900",
      "postDate": "01/01/2021 17:29:18",
      "content": "<p>It has not helped me so far, as mentioned in the opening message. It brings my score down by 0.005, which is significant. If it went in the opposite direction, I'd be writing my inference kernel already 😅.</p>\n<p>I have a few negatives; but it's on the order of 10s of thousands, versus many tens of millions of regular positive records. <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> in that example, it's also interesting that 868 content_id is attempted twice sequentially.</p>",
      "rawMarkdown": "It has not helped me so far, as mentioned in the opening message. It brings my score down by 0.005, which is significant. If it went in the opposite direction, I'd be writing my inference kernel already 😅.\n\nI have a few negatives; but it's on the order of 10s of thousands, versus many tens of millions of regular positive records. @jaideepvalani in that example, it's also interesting that 868 content_id is attempted twice sequentially.",
      "votes": null
    },
    {
      "id": "1134970",
      "postDate": "01/01/2021 18:48:08",
      "content": "<p>When I trained using that feature, it was classified using feature importance gain as the last in my 52 LGBM model features, even worse than the random noise feature that I added.</p>",
      "rawMarkdown": "When I trained using that feature, it was classified using feature importance gain as the last in my 52 LGBM model features, even worse than the random noise feature that I added.",
      "votes": null
    },
    {
      "id": "1135435",
      "postDate": "01/02/2021 08:50:22",
      "content": "<p>Actually I feel it shouldn't matter much even if we do merely tst2 -ts1 <br>\nModel should learn the proportionality here ,if  lag  was more then ts2 will be far and hence ts2 -ts1 will be more and vice-versa </p>",
      "rawMarkdown": "Actually I feel it shouldn't matter much even if we do merely tst2 -ts1 \nModel should learn the proportionality here ,if  lag  was more then ts2 will be far and hence ts2 -ts1 will be more and vice-versa",
      "votes": null
    },
    {
      "id": "1136407",
      "postDate": "01/03/2021 03:42:01",
      "content": "<p>for LGB prior seq based features are of least importance…i suppose.</p>",
      "rawMarkdown": "for LGB prior seq based features are of least importance...i suppose.",
      "votes": null
    },
    {
      "id": "1141103",
      "postDate": "01/06/2021 13:59:01",
      "content": "<p>It looks like I made a serious mistake with my data processing.</p>\n<p>Rather than abstracting and reusing the same code, in my stupidity, I had my algo basically built out twice: one for train and one for val. Unfortunately, I forgot to update the loop header which resulted in <code>prior_question_elapsed_time</code> not being updated in the val set:</p>\n<pre><code>    for idx, (_, user_id, task_container_id, content_type_id, part_id, timestamp, prior_question_elapsed_time, answered_correctly) in enumerate(tqdm(df_train[[\n        'user_id', 'task_container_id', 'content_type_id', 'part_id', 'timestamp', 'prior_question_elapsed_time', 'answered_correctly'\n    ]].itertuples())):\n</code></pre>\n<p>vs</p>\n<pre><code>    for idx, (_, user_id, task_container_id, content_type_id, timestamp, answered_correctly) in enumerate(tqdm(df_valid[[\n        'user_id', 'task_container_id', 'content_type_id', 'timestamp', 'answered_correctly'\n    ]].itertuples())):\n</code></pre>\n<p>Resulting in the last <code>prior_question_elapsed_time</code> value from the train set being used for all val records. I don't use prior_question_elapsed_time for any other engineered features so nothing else in the pipeline should have been affected except the lag feature computation. Going to retry running my lag feature tests. If there's signal, then I encourage others to also double and triple check their code.</p>",
      "rawMarkdown": "It looks like I made a serious mistake with my data processing.\n\nRather than abstracting and reusing the same code, in my stupidity, I had my algo basically built out twice: one for train and one for val. Unfortunately, I forgot to update the loop header which resulted in `prior_question_elapsed_time` not being updated in the val set:\n\n```\n    for idx, (_, user_id, task_container_id, content_type_id, part_id, timestamp, prior_question_elapsed_time, answered_correctly) in enumerate(tqdm(df_train[[\n        'user_id', 'task_container_id', 'content_type_id', 'part_id', 'timestamp', 'prior_question_elapsed_time', 'answered_correctly'\n    ]].itertuples())):\n```\n\nvs\n\n```\n    for idx, (_, user_id, task_container_id, content_type_id, timestamp, answered_correctly) in enumerate(tqdm(df_valid[[\n        'user_id', 'task_container_id', 'content_type_id', 'timestamp', 'answered_correctly'\n    ]].itertuples())):\n```\n\nResulting in the last `prior_question_elapsed_time` value from the train set being used for all val records. I don't use prior_question_elapsed_time for any other engineered features so nothing else in the pipeline should have been affected except the lag feature computation. Going to retry running my lag feature tests. If there's signal, then I encourage others to also double and triple check their code.",
      "votes": null
    },
    {
      "id": "1141112",
      "postDate": "01/06/2021 14:07:10",
      "content": "<p>After verifying my code yesterday, I also found a mistake in the logic of my code. After fixing it, the feature did improve my LGBM model from 0.794 -&gt; 0.795x <br>\nTried it also with SAINT but it made it worse.</p>\n<p>I am curious if you'd experience more improvement out of it though, because I think that my implementation is still buggy, however I don't have the time to re-train the model.</p>",
      "rawMarkdown": "After verifying my code yesterday, I also found a mistake in the logic of my code. After fixing it, the feature did improve my LGBM model from 0.794 -> 0.795x \nTried it also with SAINT but it made it worse.\n\nI am curious if you'd experience more improvement out of it though, because I think that my implementation is still buggy, however I don't have the time to re-train the model.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1132093,
      "author_name": "manikanthr5",
      "author_url": "",
      "post_date": "12/30/2020 06:58:39",
      "content": "<p>I also tried to calculate the lag_time using the way you mentioned. When I saw the negative values, I dropped it and started only using timestamp(t) - timestamp(t-1) as lag_time/response_time. I am yet to debug my SAINT+ model since it's not training. If it works, I will share my results. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1132101,
          "author_name": "manikanthr5",
          "author_url": "",
          "post_date": "12/30/2020 07:02:52",
          "content": "<p>claverru has discussed about the <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632#1107391\" target=\"_blank\">same here</a>. He is only using timestamp(t) - timestamp(t-1), capping it to 1 day and scaling it to [0, 1]</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1132293,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "12/30/2020 09:40:33",
          "content": "<p>After changing that approach I had a boost in my score. Tip: train with negatives.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1132393,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/30/2020 11:26:59",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>   getting confused did  first approach gaveu more lb score or second one i </p>\n<p>secondly.. <br>\nhow many class minutes should one consider for categorizing the lag time durations if in minutes</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1132455,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "12/30/2020 12:22:11",
          "content": "<p>If you are getting negatives in your timestamp diff, it is because you are not handleing the fact that you have some simultaneous timestamps due to questions in the same bundle. </p>\n<p>If after solving the prior bug, you try to compute the lag, you will still have negatives. This is something counterintuitive that has been dicussed in other posts. I suggest you to think about it. </p>\n<p>You have a good position in the LB, maybe playing around this leads you to gold or higher.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1132473,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/30/2020 12:35:32",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>   thanks.. may be i remember one post about the calculation that talked how we should handle repeat Tss for q in bundle. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1132634,
          "author_name": "authman",
          "author_url": "",
          "post_date": "12/30/2020 14:56:17",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> are you embedding this lag feature categorically or with a dense? If categorically, are you simply windsorizing the long tails?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1132785,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/30/2020 17:27:13",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> </p>\n<p><code>The we can calculate for the current container a feature: prior_container_lag_time = prior_container_ts_diff - prior_container_elapsed_time * prior_container_num_items</code></p>\n<p>I again get confused looking at paper time line diagram</p>\n<p>it says<br>\nE1- start time T1 - We dont know<br>\nR1 (Response submitted time) - TIme-T2 - We know it as TS1</p>\n<p>Lagtime Lt1-  To calculate</p>\n<p>E2-Start TIme T3- We dont know </p>\n<p>R2-Stop Time T4 - We know TS2 </p>\n<p>APplying your formula</p>\n<p>Ts2-Ts1 or T4-T2 = Will give T2 to T4 .. in between we dont know T3</p>\n<p>How your formula gives  LT correctly with one known factor</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1132839,
          "author_name": "authman",
          "author_url": "",
          "post_date": "12/30/2020 18:00:51",
          "content": "<p>No worries. Give me a second I'll draw it on paper.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1132843,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/30/2020 18:05:01",
          "content": "<p>Sure  taking this info also </p>\n<p><code>The timestamp column shows when an activity is finished, not when it started. A user could have spent an arbitrary amount of time working on the first problem before the first timestamp value would get logged.</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1132858,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "12/30/2020 18:22:02",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> Continuous time features</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1132888,
          "author_name": "authman",
          "author_url": "",
          "post_date": "12/30/2020 19:00:51",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2Fe0602f275397f69e833cd9c65825680e%2FIMG_6084.jpg?generation=1609354808793957&amp;alt=media\" alt=\"\"></p>\n<p>This is the case when the previous task container has just item. If they have multiple items, just multiply previous_elapsed_time by the number of items in the bundle.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1133188,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/31/2020 02:34:10",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a>  thanks elapsed time which you subtract  is then  of task c id 3 ?</p>\n<p>So from you paper draw  lag of last task container<br>\n= last time stamp  - 2nd last tstamp - last elapsed time <br>\n? </p>\n<p>If this is final case then we need a info  one time step back , and one time step after ( elt of current)  current time step to know current tc lag time . <br>\n<a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>  During inference how we manage if above is  method to get right lagtime for test q  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1133200,
          "author_name": "shujun717",
          "author_url": "",
          "post_date": "12/31/2020 02:46:11",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> That's a great tip. Looking at my score vs gold zone solutions I cant help but feel im doing something very wrong.. But after what you metioned I think I'm beginning to understand what i am doing wrong</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1133221,
          "author_name": "authman",
          "author_url": "",
          "post_date": "12/31/2020 03:13:26",
          "content": "<blockquote>\n  <p>= last time stamp - 2nd last tstamp - last elapsed time </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> correct. And elapsed time should be of the previous container, and luckily, they actually give us exactly that at time \"we are here!\".</p>\n<blockquote>\n  <p>If this is final case then we need a info one time step back , and one time step after ( elt of current) current time step to know current tc lag time . </p>\n</blockquote>\n<p>We will never know the current lag time for a bundle at inference time, similar to how we will never know the current start time of a bundle at inference. Only when we move to the next bundle is it possible for us to rectify previous bundle: start time, lag time, elapsed time. And that shouldn't <em>really</em> be an issue anyway since saint records are shifted due to the starter tag. But even if they weren't, it just means gotta use what you got.</p>",
          "votes": null,
          "replies": [
            {
              "id": 1134282,
              "author_name": "jaideepvalani",
              "author_url": "",
              "post_date": "01/01/2021 05:16:34",
              "content": "<blockquote>\n  <blockquote>\n    <p>= last time stamp - 2nd last tstamp - last elapsed time </p>\n  </blockquote>\n  <p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> correct. And elapsed time should be of the previous container, and luckily, they actually give us exactly that at time \"we are here!\".</p>\n  <blockquote>\n    <p>If this is final case then we need a info one time step back , and one time step after ( elt of current) current time step to know current tc lag time . </p>\n  </blockquote>\n  <p>We will never know the current lag time for a bundle at inference time, similar to how we will never know the current start time of a bundle at inference. Only when we move to the next bundle is it possible for us to rectify previous bundle: start time, lag time, elapsed time. And that shouldn't <em>really</em> be an issue anyway since saint records are shifted due to the starter tag. But even if they weren't, it just means gotta use what you got.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> <br>\n<code>train_df[(train_df.user_id==13134) &amp; (train_df.timestamp&gt;12948521851-10000)]</code></p>\n<p>Will we get negative lag time <br>\nthis user id &amp; task container id 53., applying your formula</p>\n<pre><code>    timestamp   user_id content_id  content_type_id task_container_id   user_answer answered_correctly  prior_question_elapsed_time prior_question_had_explanation  bundle_count\n672    12948517301 13134   868 False   54  0   1   12000.0 1.0 1\n673    12948521851 13134   868 False   53  3   0   7000.0  1.0 1\n674    12948556717 13134   1136False   55  0   1   23000.0\n</code></pre>\n<p><br>\n12948521851-12948517301    -23000</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1133260,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/31/2020 04:09:25",
          "content": "<p>yes but while predictin ans for test q, it will give us elt of last q of users train data series.. we will have to dynamically calculate lag time for  last q, for model to see lag time while predicting ans correctness for test q at tail</p>\n<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>  how many negative values approx one has for user id/tc id</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1134633,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "01/01/2021 12:52:21",
          "content": "<p>Did you get any gain out of this feature guys ? I didn't get any so far.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1134900,
          "author_name": "authman",
          "author_url": "",
          "post_date": "01/01/2021 17:29:18",
          "content": "<p>It has not helped me so far, as mentioned in the opening message. It brings my score down by 0.005, which is significant. If it went in the opposite direction, I'd be writing my inference kernel already 😅.</p>\n<p>I have a few negatives; but it's on the order of 10s of thousands, versus many tens of millions of regular positive records. <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> in that example, it's also interesting that 868 content_id is attempted twice sequentially.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1134970,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "01/01/2021 18:48:08",
          "content": "<p>When I trained using that feature, it was classified using feature importance gain as the last in my 52 LGBM model features, even worse than the random noise feature that I added.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1136407,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/03/2021 03:42:01",
          "content": "<p>for LGB prior seq based features are of least importance…i suppose.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1132407,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "12/30/2020 11:43:50",
      "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> nice observation . Check our your distriution with previous best how do they look like. You should  get some insight.<br>\nSecondly one thing i still understand  if we are given only stop time stamp or completion time stamp. Why first task container solved by student is having Time stamp as 0  ,rather then its actual completion timestamp. Ours lag time approach is also simple so far but gets decent  saint plus score of 79.3</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1134461,
      "author_name": "bowaka",
      "author_url": "",
      "post_date": "01/01/2021 10:10:17",
      "content": "<p>I am currently using that feature personnaly, and it gives a boost to my model.<br>\nMy personnal formula is :</p>\n<blockquote>\n  <p>TS[-2] - TS[-3] - LAST_QUESTION_REACTION_TIME[-1]*batch_size[-1]</p>\n</blockquote>\n<p>I store those values and compute rolling averages.</p>\n<p>edit: changed sign as per authman comment</p>",
      "votes": null,
      "replies": [
        {
          "id": 1134478,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/01/2021 10:28:17",
          "content": "<blockquote>\n  <p>I am currently using that feature personnaly, and it gives a boost to my model.<br>\n  My personnal formula is :</p>\n  <blockquote>\n    <p>TS[-2] - TS[-3] + LAST_QUESTION_REACTION_TIME[-1]*batch_size[-1]</p>\n  </blockquote>\n  <p>I store those values and compute rolling averages.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/bowaka\" target=\"_blank\">@bowaka</a> <br>\nits confusing a bit why u have taken TS[-3],TS[-2] while calculating lag time for current q.</p>\n<p>With your formula wat would be values while computing lag time for this user &amp; task container id 53..</p>\n<pre><code>timestamp   user_id content_id  content_type_id task_container_id   user_answer answered_correctly  prior_question_elapsed_time prior_question_had_explanation  bundle_count\n672    12948517301 13134   868 False   54  0   1   12000.0 1.0 1\n673    12948521851 13134   868 False   53  3   0   7000.0  1.0 1\n674    12948556717 13134   1136False   55  0   1   23000.0\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1134487,
          "author_name": "authman",
          "author_url": "",
          "post_date": "01/01/2021 10:48:07",
          "content": "<p>You need record 671 to do it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1134660,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/01/2021 13:20:20",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> in the sequence  below  during the inference  we can compute lag time for Q3 container<br>\nusing Q3  ts, Q2 ts,and Q3 elapsed time given in Test4 . <br>\nlike User id U1=   Q1,Q2,Q3 Test4</p>\n<p>Q3's task container lag time= Q3 Ts-Q2Ts-elt in Q4</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1134671,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "01/01/2021 13:32:32",
          "content": "<p>Sorry, I am not calculating for current q, this information is not accessible with the data. I calculate lag time between q-3 and q-2 (assuming current row is q-1)</p>\n<p>My turn to try a drawing of the situation 😄:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2Ff1090aaf2daccaad9845760f6b4800e9%2FSans%20titre.png?generation=1609514059944910&amp;alt=media\" alt=\"\"></p>\n<p>In the illustration above, BQ correspond to a same batch of question. They are served together through the API and timestamps are similar for all questions Qx in a batch BQy</p>\n<p>Note: as authman is saying, for your example, you need more rows as you didn't complete a full batch of questions yet.</p>\n<p>edit: change sign '+' in '-' as per authman comment below</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1134769,
          "author_name": "authman",
          "author_url": "",
          "post_date": "01/01/2021 15:10:07",
          "content": "<p>Your equation is identical to mine <a href=\"https://www.kaggle.com/bowaka\" target=\"_blank\">@bowaka</a> except for two things.</p>\n<ol>\n<li>I haven't tried EMA features, yet 😉.</li>\n<li>Check your equation and illustration, the <code>+</code> should be a <code>-</code>. I hope this change results in a boost for you.</li>\n</ol>\n<p>Your illustration is a lot better suited than mine.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1134775,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "01/01/2021 15:14:00",
          "content": "<p>whoops ! you are totaly right, there was a mistake in the sign ! I'll update the figure</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1135435,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/02/2021 08:50:22",
          "content": "<p>Actually I feel it shouldn't matter much even if we do merely tst2 -ts1 <br>\nModel should learn the proportionality here ,if  lag  was more then ts2 will be far and hence ts2 -ts1 will be more and vice-versa </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1141103,
      "author_name": "authman",
      "author_url": "",
      "post_date": "01/06/2021 13:59:01",
      "content": "<p>It looks like I made a serious mistake with my data processing.</p>\n<p>Rather than abstracting and reusing the same code, in my stupidity, I had my algo basically built out twice: one for train and one for val. Unfortunately, I forgot to update the loop header which resulted in <code>prior_question_elapsed_time</code> not being updated in the val set:</p>\n<pre><code>    for idx, (_, user_id, task_container_id, content_type_id, part_id, timestamp, prior_question_elapsed_time, answered_correctly) in enumerate(tqdm(df_train[[\n        'user_id', 'task_container_id', 'content_type_id', 'part_id', 'timestamp', 'prior_question_elapsed_time', 'answered_correctly'\n    ]].itertuples())):\n</code></pre>\n<p>vs</p>\n<pre><code>    for idx, (_, user_id, task_container_id, content_type_id, timestamp, answered_correctly) in enumerate(tqdm(df_valid[[\n        'user_id', 'task_container_id', 'content_type_id', 'timestamp', 'answered_correctly'\n    ]].itertuples())):\n</code></pre>\n<p>Resulting in the last <code>prior_question_elapsed_time</code> value from the train set being used for all val records. I don't use prior_question_elapsed_time for any other engineered features so nothing else in the pipeline should have been affected except the lag feature computation. Going to retry running my lag feature tests. If there's signal, then I encourage others to also double and triple check their code.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1141112,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "01/06/2021 14:07:10",
          "content": "<p>After verifying my code yesterday, I also found a mistake in the logic of my code. After fixing it, the feature did improve my LGBM model from 0.794 -&gt; 0.795x <br>\nTried it also with SAINT but it made it worse.</p>\n<p>I am curious if you'd experience more improvement out of it though, because I think that my implementation is still buggy, however I don't have the time to re-train the model.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1131395": "Very early on in this competition @gunesevitan posted this thread: [Questions about timestamp and prior_question_elapsed_time](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189351), which has been on my mind non-stop. Earlier, I made [this posting](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/205515) where I argued it wasn't possible to compute `lag_time` as defined in the SAINT+ paper. But after calmly investing more time into it, I've discovered that it actually is quite possible to calculate `lag_time`, at least at the `{user_id, task_container_id}` resolution. it's still a bit noisy for a few records, but the number of discrepancies in the final features is still on the order of 1-2% of 100M.\n\nThe way that it works is this:\n\n```\ns e    s e     s e    s e\n```\n\nImagine you have a series of start-end pairs, corresponding to when the exercise was first opened by the student and when the student completed their submission.\n\nWe know that `e` = `timestamp`, as defined in the dataset, the time when the interaction concluded.\n\nWe are also given prior_question_elapsed_time (which is actually average prior_task_container_elapsed_time, per-user), which is equivalent to `e-s` for any given single question task_container_id. In the case there are multiple items, just multiply `prior_question_elapsed_time` by the number of items in the user's task container id.\n\nThe value we are trying to compute is `s-e`. We have all the pieces.\n\nAssuming we maintain a feature ts_diff = ts - ts_previous, which is corrected s.t. all values in the bundle share the first ts_diff, then:\n\nThe we can calculate for the **current** container a feature: `prior_container_lag_time = prior_container_ts_diff - prior_container_elapsed_time * prior_container_num_items`\n\nVisually, that would be (e_x - e_(x-1)) - (e_x - s_x). Just make sure all units are ms, or s.\n\nTo my great surprise, adding this feature did not result in a score boost. Rather, it leads to over-fitting. Not leakage, but overfitting. How can having the lag time of the previous container/bundle result in overfitting of the train set? This is what the distribution looks like on 80M rows after being log transformed and standard scaled. There is a faux peak at 0 because I'm filling some nans. I've tried nan filling with min, max and mean (0).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2F4e76492a9e3aab0603ff19fee1e91cbb%2Findex.png?generation=1609264028033112&alt=media)\n\nAny ideas?",
    "1132093": "I also tried to calculate the lag_time using the way you mentioned. When I saw the negative values, I dropped it and started only using timestamp(t) - timestamp(t-1) as lag_time/response_time. I am yet to debug my SAINT+ model since it's not training. If it works, I will share my results.",
    "1132101": "claverru has discussed about the [same here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632#1107391). He is only using timestamp(t) - timestamp(t-1), capping it to 1 day and scaling it to [0, 1]",
    "1132293": "After changing that approach I had a boost in my score. Tip: train with negatives.",
    "1132393": "claverru   getting confused did  first approach gaveu more lb score or second one i \n \n\nsecondly.. \nhow many class minutes should one consider for categorizing the lag time durations if in minutes",
    "1132407": "authman nice observation . Check our your distriution with previous best how do they look like. You should  get some insight.\nSecondly one thing i still understand  if we are given only stop time stamp or completion time stamp. Why first task container solved by student is having Time stamp as 0  ,rather then its actual completion timestamp. Ours lag time approach is also simple so far but gets decent  saint plus score of 79.3",
    "1132455": "If you are getting negatives in your timestamp diff, it is because you are not handleing the fact that you have some simultaneous timestamps due to questions in the same bundle. \n\nIf after solving the prior bug, you try to compute the lag, you will still have negatives. This is something counterintuitive that has been dicussed in other posts. I suggest you to think about it. \n\nYou have a good position in the LB, maybe playing around this leads you to gold or higher.",
    "1132473": "claverru   thanks.. may be i remember one post about the calculation that talked how we should handle repeat Tss for q in bundle.",
    "1132634": "claverru are you embedding this lag feature categorically or with a dense? If categorically, are you simply windsorizing the long tails?",
    "1132785": "authman \n\n`The we can calculate for the current container a feature: prior_container_lag_time = prior_container_ts_diff - prior_container_elapsed_time * prior_container_num_items`\n\nI again get confused looking at paper time line diagram\n\nit says\nE1- start time T1 - We dont know\nR1 (Response submitted time) - TIme-T2 - We know it as TS1\n\nLagtime Lt1-  To calculate\n\nE2-Start TIme T3- We dont know \n\nR2-Stop Time T4 - We know TS2 \n\nAPplying your formula\n\nTs2-Ts1 or T4-T2 = Will give T2 to T4 .. in between we dont know T3\n\nHow your formula gives  LT correctly with one known factor",
    "1132839": "No worries. Give me a second I'll draw it on paper.",
    "1132843": "Sure  taking this info also \n\n`The timestamp column shows when an activity is finished, not when it started. A user could have spent an arbitrary amount of time working on the first problem before the first timestamp value would get logged.`",
    "1132858": "authman Continuous time features",
    "1132888": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2Fe0602f275397f69e833cd9c65825680e%2FIMG_6084.jpg?generation=1609354808793957&alt=media)\n\nThis is the case when the previous task container has just item. If they have multiple items, just multiply previous_elapsed_time by the number of items in the bundle.",
    "1133188": "authman  thanks elapsed time which you subtract  is then  of task c id 3 ?\n\nSo from you paper draw  lag of last task container\n= last time stamp  - 2nd last tstamp - last elapsed time \n? \n\nIf this is final case then we need a info  one time step back , and one time step after ( elt of current)  current time step to know current tc lag time . \n@claverru  During inference how we manage if above is  method to get right lagtime for test q",
    "1133200": "claverru That's a great tip. Looking at my score vs gold zone solutions I cant help but feel im doing something very wrong.. But after what you metioned I think I'm beginning to understand what i am doing wrong",
    "1133221": "> = last time stamp - 2nd last tstamp - last elapsed time \n\n@jaideepvalani correct. And elapsed time should be of the previous container, and luckily, they actually give us exactly that at time \"we are here!\".\n\n> If this is final case then we need a info one time step back , and one time step after ( elt of current) current time step to know current tc lag time . \n\nWe will never know the current lag time for a bundle at inference time, similar to how we will never know the current start time of a bundle at inference. Only when we move to the next bundle is it possible for us to rectify previous bundle: start time, lag time, elapsed time. And that shouldn't *really* be an issue anyway since saint records are shifted due to the starter tag. But even if they weren't, it just means gotta use what you got.",
    "1133260": "yes but while predictin ans for test q, it will give us elt of last q of users train data series.. we will have to dynamically calculate lag time for  last q, for model to see lag time while predicting ans correctness for test q at tail\n\n@claverru  how many negative values approx one has for user id/tc id",
    "1134282": "> > = last time stamp - 2nd last tstamp - last elapsed time \n> \n> @jaideepvalani correct. And elapsed time should be of the previous container, and luckily, they actually give us exactly that at time \"we are here!\".\n> \n> > If this is final case then we need a info one time step back , and one time step after ( elt of current) current time step to know current tc lag time . \n> \n> We will never know the current lag time for a bundle at inference time, similar to how we will never know the current start time of a bundle at inference. Only when we move to the next bundle is it possible for us to rectify previous bundle: start time, lag time, elapsed time. And that shouldn't *really* be an issue anyway since saint records are shifted due to the starter tag. But even if they weren't, it just means gotta use what you got.\n\n@authman \n`train_df[(train_df.user_id==13134) & (train_df.timestamp>12948521851-10000)]`\n\nWill we get negative lag time \nthis user id & task container id 53., applying your formula\n\n\n\n```\n\ttimestamp\tuser_id\tcontent_id\tcontent_type_id\ttask_container_id\tuser_answer\tanswered_correctly\tprior_question_elapsed_time\tprior_question_had_explanation\tbundle_count\n672\t12948517301\t13134\t868\tFalse\t54\t0\t1\t12000.0\t1.0\t1\n673\t12948521851\t13134\t868\tFalse\t53\t3\t0\t7000.0\t1.0\t1\n674\t12948556717\t13134\t1136False\t55\t0\t1\t23000.0\n``` \n12948521851-12948517301\t-23000",
    "1134461": "I am currently using that feature personnaly, and it gives a boost to my model.\nMy personnal formula is :\n> TS[-2] - TS[-3] - LAST_QUESTION_REACTION_TIME[-1]*batch_size[-1]\n\nI store those values and compute rolling averages.\n\nedit: changed sign as per authman comment",
    "1134478": "> I am currently using that feature personnaly, and it gives a boost to my model.\n> My personnal formula is :\n> > TS[-2] - TS[-3] + LAST_QUESTION_REACTION_TIME[-1]*batch_size[-1]\n> \n> I store those values and compute rolling averages.\n\n@bowaka \nits confusing a bit why u have taken TS[-3],TS[-2] while calculating lag time for current q.\n\nWith your formula wat would be values while computing lag time for this user & task container id 53..\n ```\ntimestamp   user_id content_id  content_type_id task_container_id   user_answer answered_correctly  prior_question_elapsed_time prior_question_had_explanation  bundle_count\n672    12948517301 13134   868 False   54  0   1   12000.0 1.0 1\n673    12948521851 13134   868 False   53  3   0   7000.0  1.0 1\n674    12948556717 13134   1136False   55  0   1   23000.0\n```",
    "1134487": "You need record 671 to do it.",
    "1134633": "Did you get any gain out of this feature guys ? I didn't get any so far.",
    "1134660": "authman in the sequence  below  during the inference  we can compute lag time for Q3 container\nusing Q3  ts, Q2 ts,and Q3 elapsed time given in Test4 . \nlike User id U1=   Q1,Q2,Q3 Test4\n\nQ3's task container lag time= Q3 Ts-Q2Ts-elt in Q4",
    "1134671": "Sorry, I am not calculating for current q, this information is not accessible with the data. I calculate lag time between q-3 and q-2 (assuming current row is q-1)\n\nMy turn to try a drawing of the situation 😄:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2Ff1090aaf2daccaad9845760f6b4800e9%2FSans%20titre.png?generation=1609514059944910&alt=media)\n\nIn the illustration above, BQ correspond to a same batch of question. They are served together through the API and timestamps are similar for all questions Qx in a batch BQy\n\nNote: as authman is saying, for your example, you need more rows as you didn't complete a full batch of questions yet.\n\nedit: change sign '+' in '-' as per authman comment below",
    "1134769": "Your equation is identical to mine @bowaka except for two things.\n\n1. I haven't tried EMA features, yet 😉.\n1. Check your equation and illustration, the `+` should be a `-`. I hope this change results in a boost for you.\n\nYour illustration is a lot better suited than mine.",
    "1134775": "whoops ! you are totaly right, there was a mistake in the sign ! I'll update the figure",
    "1134900": "It has not helped me so far, as mentioned in the opening message. It brings my score down by 0.005, which is significant. If it went in the opposite direction, I'd be writing my inference kernel already 😅.\n\nI have a few negatives; but it's on the order of 10s of thousands, versus many tens of millions of regular positive records. @jaideepvalani in that example, it's also interesting that 868 content_id is attempted twice sequentially.",
    "1134970": "When I trained using that feature, it was classified using feature importance gain as the last in my 52 LGBM model features, even worse than the random noise feature that I added.",
    "1135435": "Actually I feel it shouldn't matter much even if we do merely tst2 -ts1 \nModel should learn the proportionality here ,if  lag  was more then ts2 will be far and hence ts2 -ts1 will be more and vice-versa",
    "1136407": "for LGB prior seq based features are of least importance...i suppose.",
    "1141103": "It looks like I made a serious mistake with my data processing.\n\nRather than abstracting and reusing the same code, in my stupidity, I had my algo basically built out twice: one for train and one for val. Unfortunately, I forgot to update the loop header which resulted in `prior_question_elapsed_time` not being updated in the val set:\n\n```\n    for idx, (_, user_id, task_container_id, content_type_id, part_id, timestamp, prior_question_elapsed_time, answered_correctly) in enumerate(tqdm(df_train[[\n        'user_id', 'task_container_id', 'content_type_id', 'part_id', 'timestamp', 'prior_question_elapsed_time', 'answered_correctly'\n    ]].itertuples())):\n```\n\nvs\n\n```\n    for idx, (_, user_id, task_container_id, content_type_id, timestamp, answered_correctly) in enumerate(tqdm(df_valid[[\n        'user_id', 'task_container_id', 'content_type_id', 'timestamp', 'answered_correctly'\n    ]].itertuples())):\n```\n\nResulting in the last `prior_question_elapsed_time` value from the train set being used for all val records. I don't use prior_question_elapsed_time for any other engineered features so nothing else in the pipeline should have been affected except the lag feature computation. Going to retry running my lag feature tests. If there's signal, then I encourage others to also double and triple check their code.",
    "1141112": "After verifying my code yesterday, I also found a mistake in the logic of my code. After fixing it, the feature did improve my LGBM model from 0.794 -> 0.795x \nTried it also with SAINT but it made it worse.\n\nI am curious if you'd experience more improvement out of it though, because I think that my implementation is still buggy, however I don't have the time to re-train the model."
  },
  "source": "meta"
}